cp: -r or -R?

4 days ago (movq.de)

This blog post is odd because it keeps on hinting about a difference between `-r' and `-R' and links to the source code but never actually says what it is. I'll quote the OpenBSD manual that the post mentions but does not link to for some reason:

> Historic versions of the cp utility had an -r option. This implementation supports that option; however, its use is strongly discouraged, as it does not correctly copy special files, symbolic links or FIFOs.

https://man.openbsd.org/cp

  • You can see it in the code, lstat vs stat and copy_special, copy_fifo vs copy_file. They're all in utils.c.

    from the lstat man page:

         The lstat() function is identical to stat() except when the named file is
         a symbolic link, in which case lstat() returns information about the link
         itself, not the file the link references.
    

    symbolic link example:

        $ tree test1
        test1
        |-- file.txt
        `-- link.txt -> ../link.txt
    
        0 directory, 2 files
    
        $ cp -r test1 test2
        $ tree test2
        test2
        |-- file.txt
        `-- link.txt
    
        0 directory, 2 files
    
        $ cp -R test1 test3
        $ tree test3
        test3
        |-- file.txt
        `-- link.txt -> ../link.txt
    
        0 directory, 2 files
    

    For posterity, -R with added -L

        $ cp -RL test1 test4
        $ tree test4/
        test4/
        |-- file.txt
        `-- link.txt
    
        0 directory, 2 files

  • > links to the source code but never actually says what it is.

    The snippet of code makes it very clear what the difference is, no?

    flag_copy_as_regular = 1 VS flag_copy_as_regular = 0, where regular would be a "regular" and not "special" file.

    • Only clear to those with a heavy UNIX background-- I understood and correctly interpreted it on first read, but the concept of a "regular file" is a bit in the wonk territory for most users, though the implication on symbolic links is a very practical one that surely lead to the removal (in the implementations where it was removed) of this option, since its very surprising that copying a directory heirarchy would cause symbolic links to become copies of what they linked to, instead of remaining symbolic links in the resulting directory.

> nobody™ still runs coreutils from 24 years ago.

  Sprite 2.077 pc386
  
          Welcome to Sprite
  
  root@cherimoya [1] # cp --version
  GNU fileutils 3.9
  root@cherimoya [2] # cp --help | grep recurs
    -r                           copy recursively, non-directories as files
    -R, --recursive              copy directories recursively
  root@cherimoya [3] #

I'll take the honorary title of nobody ;-)

  • Hey, I perfectly understand if not. But do you have the sources that were used to build those binaries?

    The oldest version on the GNU ftp server is fileutils-3.13 [1]. I vaguely remember having some links to older versions, probably somewhere in my archived mail. But I don't remember if it was fileutils-3.9 or earlier.

    I co-maintain GNU coreutils, so I am interested in reading them. If you have them, you can email me privately or on the public mailing list. Both are listed on the homepage [2].

    [1] https://ftp.gnu.org/old-gnu/fileutils/ [2] https://www.gnu.org/software/coreutils/

    • You can find ancient distros on archive.org and some of those have src.rpm's Slackware also had a source dvd with iirc regular tarballs.

      2 replies →

Kind of off topic, but why in the world in 2026 do POSIX utilities still rely on global variables to pass around state? Is it too complicated to pass around a struct address in C or something?

  • I think there is a class of programs where globals are not so terrible. In particular, one-shot programs that do a task and then exit (like `ls`, `cp`, etc).

    You may also notice that some of these utilities will allocate memory, then never free it. They just let the OS handle that on program shutdown. That is also generally not recommended in arbitrary code. But again, it's fine for a one-and-done sort of program.

    For these cases, the whole program is essentially one function call, and the global namespace is essentially the "body" of that function (hand waving a bit).

    It's all a bit subjective, of course (:

    • Reminds me of the story of the memory leak in a missile guidance computer. It was determined that in the worst case scenario not enough memory could leak fast enough for it to be a problem between missile launch and impact. So they just let the explosion handle the garbage collection

-a

Not sure why you wouldn't want to preserve timestamps, links, etc. by default.

  • This is rather missing the point. The headlined article isn't really about how to achieve a goal, but about the weird history and evolution of a tool that leads us to the rather odd situation that we are in today. And it's far from being the only tool that has a weird history, that looks rather nutty if one looks at it from the point of view of a novice having to learn this stuff.

    It's also not even completely covering the weird case of -r and -R for the cp command. On HP-UX, for example, the twain were different, but not in the way that they were in old GNU Core Utilities. That would be too easy. (-:

    The AIX manual for cp explains its difference between -r and -R:

    * https://ibm.com/docs/en/aix/7.1.0?topic=c-cp-command

    Illumos also treats the two differently, but in a subtly different way:

    * https://illumos.org/man/1/cp

    • That's why, just as fork(2) is a primitive for the process creation, copy(2) should've been the primitive for the file creation — creates an exact copy of the file under a new name, but with the exact same content and all of the metadata (except for the name, obviously), including its kind, permissions, timestamps, etc. And no, it wouldn't be prohibitively expensive because all filesystems can quite easily support CoW; after all, most of the created files will be truncate(2)d almost immediately, so there is no point to eagerly duplicate the file contents.

      The "metadata is atomically copied" part would support very nicely the usual text editor's idiom of rename(2)ing a temporary file over the source after fully writing it out — you still need to accurately replicate the permissions and extended attributes. And just as shells are important enough programs to have fork(2) almost exactly suited for them, it would make sense to have copy(2), suited for the text editors.

      3 replies →

  • old cp didn't have `-a`

    Anyway, just use rsync.

    • > just use rsync

      You still need to specify --archive (or --times for the individual option) to preserve mtime in the target copy.

      But yeah, I tend to rsync more than I cp.

-h perhaps the only useful comment here. -R just looks baroque, but ls(1) might think it fine. I once read an article about the inconsistencies in *nix CLI commands, but the picture's much better than hot-key and shortcut conflicts- <which get silently absorbed by whatever's running, no way to tell what they might've done or where they went..> I hestitate to consult documentation, because of course there is non anymore. Why not Google it?

I'm more inclined to use the uppercase -R as it's standardized by POSIX and will generally behave the same on any POSIX compliant system.

  • It's also the only option shown in the -h output and in man pages for some versions of cp.

    I didn't even know -r was a thing until today.

It always seemed like the recursive flag of cp was an implementation detail leaking into the UI. Like, I get that copying a file requires creating more than one inode, but...so? Eventually, graphical OSes agree with me—copy/paste works the same on folders as it does on files.

I really wish there was a way to know if LLMs hallucinate these switches incorrectly, like I do.

Feels like this would be exactly the kind of thing they would get wrong. Fur exactly, the training set isn't trained to know the context of execution (FreeBSD vs macos vs Linux), right?

  • its trained to read both tekst and code which is enough to know the difference.

    appearently i am not :') never knew there was -r

This always get me. I instinctively -r, until chown which of course doesn’t take it.

On a somewhat related note, I really hate that in scp -r and -R mean entirely different things.

  • The worst is when things behave different when you give them `~/somedir` vs `~/somedir/`. I think it's rsync that does that

    • I really like this feature of rsync (trailing "/" means copy the directory contents to the dest, no trailing "/" means copy the directory itself). Other tools, like cp, don't have any way at all to say copy the directory contents to the dest, and for those tools the result depends on whether or not the destination already exists and is a directory (you might end up with a duplicate nested directory). Rsync produces the same result whether the destination already exists or not.

    • rsync, or at least the version I had would behave differently for 'rsync a b' vs 'rsync a/ b/', even though I added the slash for both sides.

  • How about port that is lowercase in ssh and uppercase in scp?

    • scp took -p from rcp/cp, where it already meant preserve times. So port got -P.

would you be safe in using --recursive always? (e.g. shell scripts)

  • Double dash long options are basically a GNU extension. BSD utilities generally don't support them. Apparently macOS does not either (since it's based off of FreeBSD)

    • macOS was based off NeXTSTEP, not FreeBSD.

      And the received wisdom about long options in the BSDs is a quarter of a century out of date. When the BSDs gained a getopt_long() in their C libraries thanks to Klausner and Baron, long options quietly started appearing. This process has been gradually and quietly on-going for the whole of the 21st century.

      4 replies →

Use rsync instead

  • In my minimal attempts to use rsync, I always find examples where they always use a ton of flags alongside the locations. I'm not gonna learn those flags if cp can do it intuitively and with minimal extra commands. Maybe it's just me.

    • Not just you.

      Also, I always have this vague fear that I'll rsync in the wrong direction, or accidentally blow away unrelated files in rsync's efforts to fully synchronize two directories (can't remember if this is a valid concern).

      I'm sure these concerns would go away if I used it regularly, but I just don't. ‘cp' or ’scp’ almost always meet my needs.

      Kinda like the way people are probably right that I should learn to use ’awk’, but I just can't muster the motivation.

    • Do it once, make it your muscle memory - it's really hard for me to learn things by heart, but even I could do it :), and then you can forget about scp as well

there there are antisocial surprises, like YC's 'login' required after writing a post - which led me to go log into another machine, check my list, then come back here and login. Looming is the question whether a login failure would've erased what I wrote. There are of course fantasticaly more egregious UX gaffes.. imagining now [cp] and [ls] buttons .. perhaps better than 'intelligent' locator select/copy/paste madddenlingly split between kb keys and screen touches, highlighting not exactly what precisely you specify, instead crazily jumpping the selection about phrenically, while necessity of scrolling offscreen up and down further complicates - often 'select all' means 'select only whats onscreen' <surprise later, only after a paste> micro displays, big thumbs, filly 5/8ths of my present Android screen obscured by an on-screen keyboaard, even though I'm typing on a BT kb. As you might infer from this rambling paragraph I'm an OOtB ADHD wannabe-Dev type, enough at least to have heard about 'DNRY' which in sum, I'd say I hope Ai somehow obliterates <after the fashion of *nix apt's, perhaps?> with 'you can you it anyway you like - there are a dozen ways to do anything, and they all work all the time - modulo of course in reality that anti-DNRY paridigm taken beyond the UI foments chaos and confusion.