cp: -r or -R?

44 points by zdw · 2 days ago · 56 comments · movq.de ↗
Loading article...

Comments (56)

This blog post is odd because it keeps on hinting about a difference between `-r' and `-R' and links to the source code but never actually says what it is. I'll quote the OpenBSD manual that the post mentions but does not link to for some reason:

> Historic versions of the cp utility had an -r option. This implementation supports that option; however, its use is strongly discouraged, as it does not correctly copy special files, symbolic links or FIFOs.

https://man.openbsd.org/cp

> links to the source code but never actually says what it is.

The snippet of code makes it very clear what the difference is, no?

flag_copy_as_regular = 1 VS flag_copy_as_regular = 0, where regular would be a "regular" and not "special" file.

> nobody™ still runs coreutils from 24 years ago.

  Sprite 2.077 pc386
  
          Welcome to Sprite
  
  root@cherimoya [1] # cp --version
  GNU fileutils 3.9
  root@cherimoya [2] # cp --help | grep recurs
    -r                           copy recursively, non-directories as files
    -R, --recursive              copy directories recursively
  root@cherimoya [3] #
I'll take the honorary title of nobody ;-)
You should have run that command as the `nobody` user.
-a

Not sure why you wouldn't want to preserve timestamps, links, etc. by default.

This is rather missing the point. The headlined article isn't really about how to achieve a goal, but about the weird history and evolution of a tool that leads us to the rather odd situation that we are in today. And it's far from being the only tool that has a weird history, that looks rather nutty if one looks at it from the point of view of a novice having to learn this stuff.

It's also not even completely covering the weird case of -r and -R for the cp command. On HP-UX, for example, the twain were different, but not in the way that they were in old GNU Core Utilities. That would be too easy. (-:

The AIX manual for cp explains its difference between -r and -R:

* https://ibm.com/docs/en/aix/7.1.0?topic=c-cp-command

Illumos also treats the two differently, but in a subtly different way:

* https://illumos.org/man/1/cp

That's why, just as fork(2) is a primitive for the process creation, copy(2) should've been the primitive for the file creation — creates an exact copy of the file under a new name, but with the exact same content and all of the metadata (except for the name, obviously), including its kind, permissions, timestamps, etc. And no, it wouldn't be prohibitively expensive because all filesystems can quite easily support CoW; after all, most of the created files will be truncate(2)d almost immediately, so there is no point to eagerly duplicate the file contents.

The "metadata is atomically copied" part would support very nicely the usual text editor's idiom of rename(2)ing a temporary file over the source after fully writing it out — you still need to accurately replicate the permissions and extended attributes. And just as shells are important enough programs to have fork(2) almost exactly suited for them, it would make sense to have copy(2), suited for the text editors.

When cp was invented, 'most filesystems' did not have the first clue about copy on write.

However, I should note that possibly the first company to invent what you describe was Microsoft.

Novell Netware 386 had an NCOPY command which invoked a Netware extension to the DOS API that told the server to perform the entire copy on the server.

But even earlier, OS/2 1.x had a proper DosCopy() system call. Since it could be passed down to the installable filesystem drivers for intra-volume copies, something like the Netware client for OS/2 could in theory turn it into the same protocol call that did server-side copies. There was a NET COPY command in LAN Manager (and LAN Server, if memory serves) that did the same optimization.

* https://www.edm2.com/index.php/DosCopy_(OS/2_1.x)

* https://www.edm2.com/index.php/FS_COPY

> When cp was invented, 'most filesystems' did not have the first clue about copy on write.

Eh, when fork was invented, most (virtual) memory systems did not have the first clue about copy-on-write either. And honestly, it's really not that difficult to support — it's essentially hard links, just with slightly different semantics.

old cp didn't have `-a`

Anyway, just use rsync.

> just use rsync

You still need to specify --archive (or --times for the individual option) to preserve mtime in the target copy.

But yeah, I tend to rsync more than I cp.

I really wish there was a way to know if LLMs hallucinate these switches incorrectly, like I do.

Feels like this would be exactly the kind of thing they would get wrong. Fur exactly, the training set isn't trained to know the context of execution (FreeBSD vs macos vs Linux), right?

its trained to read both tekst and code which is enough to know the difference.

appearently i am not :') never knew there was -r

It's "ditto", not "dito"
As in the Pokémon, not the Philippines telecom company.
Much easier to press a key twice than to hunt for é for most people. I get your point though
Apparently Pokémon is in my phone’s word book.
I never noticed but indeed in France the official name is "Pokémon" and not "Pokemon". So that the pronunciation in French is correct.

Is it the case in another country to have a localized name for that?

Pokémon is the trademark (see e.g. https://en.wikipedia.org/wiki/Pok%C3%A9mon). Pokemon is common spelling in countries where keyboards don't have é and/or where people aren't used to inputting é.

It is common for brands to localize their names. E.g. Axe (the deodorant) in some countries is branded as Lynx.

Huh, TIL. In German, "dito" is the correct spelling, so I always figured it would be the same in English as well.
It always seemed like the recursive flag of cp was an implementation detail leaking into the UI. Like, I get that copying a file requires creating more than one inode, but...so? Eventually, graphical OSes agree with me—copy/paste works the same on folders as it does on files.
Thirteen years ago I asked the same question:

https://unix.stackexchange.com/questions/82485/when-wouldnt-...

It seems that recursive by default would have been much more intuitive.

It's rather sad that none of the answers were that the cp command simply did not gain a recursive option until the 1980s, well into the 1980s if you were on one side of the Unix wars.

Yes, seriously. When you read about the supposed evils of cat -v from the Unix nostalgia people, remember that it was the same people who gave cat its -v option who also gave cp its -r option, in 4.2BSD.

It took over half a decade to percolate out of the BSD world, too. AT&T Unix System 5 did not have an -r option to cp. Here's Brandon S. Allbery explaining in 1987 how one copies directories on AT&T Unix System 5 Releases 2/3 by combining find and cpio -p:

* https://groups.google.com/g/comp.unix.questions/c/XiumTgkcYR...

Originally we read directories as raw byte streams and liked it, you know. (-:

ditto is an option on macOS for copying files and directories [1].

[1]: https://keith.github.io/xcode-man-pages/ditto.1.html

Use rsync instead
In my minimal attempts to use rsync, I always find examples where they always use a ton of flags alongside the locations. I'm not gonna learn those flags if cp can do it intuitively and with minimal extra commands. Maybe it's just me.
Not just you.

Also, I always have this vague fear that I'll rsync in the wrong direction, or accidentally blow away unrelated files in rsync's efforts to fully synchronize two directories (can't remember if this is a valid concern).

I'm sure these concerns would go away if I used it regularly, but I just don't. ‘cp' or ’scp’ almost always meet my needs.

Kinda like the way people are probably right that I should learn to use ’awk’, but I just can't muster the motivation.

Do it once, make it your muscle memory - it's really hard for me to learn things by heart, but even I could do it :), and then you can forget about scp as well
rsync unfortunately doesn't do relinking at all, so I can't use it as a generic replacement for cp.
I assume they typoed reflinking, a way of doing CoW-based copying of file contents.
I was under the impression it could do that, I think I build a snapshotting mechanism using this.
rsync does not have the ability to copy a file using the filesystem's copy-on-write feature. This is unlike cp, which has the --reflink flag. Here's the bug report on GitHub which the developers closed as "won't fix for now":

https://github.com/RsyncProject/rsync/issues/119#issuecommen...

This always get me. I instinctively -r, until chown which of course doesn’t take it.
I'm more inclined to use the uppercase -R as it's standardized by POSIX and will generally behave the same on any POSIX compliant system.
would you be safe in using --recursive always? (e.g. shell scripts)
There are implementations of cp out in the wild that do not recognise the --recursive flag. OpenBSD was mentioned in the article and there’s also busybox cp https://busybox.net/downloads/BusyBox.html
Double dash long options are basically a GNU extension. BSD utilities generally don't support them. Apparently macOS does not either (since it's based off of FreeBSD)
macOS was based off NeXTSTEP, not FreeBSD.

And the received wisdom about long options in the BSDs is a quarter of a century out of date. When the BSDs gained a getopt_long() in their C libraries thanks to Klausner and Baron, long options quietly started appearing. This process has been gradually and quietly on-going for the whole of the 21st century.

> macOS was based off NextBSD, not FreeBSD.

"NextBSD" was first released in 2015:

* https://en.wikipedia.org/wiki/NextBSD

Over a decade after macOS/Mac OS X was initially released:

> macOS (previously OS X and originally Mac OS X) is a proprietary Unix[7][8] operating system, derived from OPENSTEP for Mach and FreeBSD, which has been marketed and developed by Apple since 2001.

* https://en.wikipedia.org/wiki/MacOS

> Darwin is the core Unix-like operating system of macOS, iOS, watchOS, tvOS, iPadOS, audioOS, visionOS, and bridgeOS. It previously existed as an independent open-source operating system, first released by Apple in 2000. It is composed of code derived from NeXTSTEP, FreeBSD[3] and other BSD operating systems,[7] Mach, and […]

* https://en.wikipedia.org/wiki/Darwin_(operating_system)

I remember reading release notes for FreeBSD in the '00s and seeing the exact same lines in the release notes for earlier versions of OS X.

> Darwin is the core Unix-like operating system of macOS

macOS isn't "Unix-like"; it's an Open Group certified UNIX™ [1].

[1]: https://www.opengroup.org/openbrand/register/brand3725.htm

Bah! Thought NeXTSTEP. Typed NextBSD. Fixed. I've been typing lots of names ending in 'BSD' today. (-:
I believe that -R is the safe works-as-expected-everywhere option.
On a somewhat related note, I really hate that in scp -r and -R mean entirely different things.
The worst is when things behave different when you give them `~/somedir` vs `~/somedir/`. I think it's rsync that does that
I really like this feature of rsync (trailing "/" means copy the directory contents to the dest, no trailing "/" means copy the directory itself). Other tools, like cp, don't have any way at all to say copy the directory contents to the dest, and for those tools the result depends on whether or not the destination already exists and is a directory (you might end up with a duplicate nested directory). Rsync produces the same result whether the destination already exists or not.
rsync, or at least the version I had would behave differently for 'rsync a b' vs 'rsync a/ b/', even though I added the slash for both sides.
Yes, into versus onto. Luckily we have AI to write our command lines.
How about port that is lowercase in ssh and uppercase in scp?
scp took -p from rcp/cp, where it already meant preserve times. So port got -P.
tar cf - . | tar xf - -C <dest>
with a sandwiched `| pv |` for fun stats
And `| ssh <target> ` for a remote copy.

(or ssh <target> prepended rather than sandwiched, to copy from remote)

At some point it starts very much looking like zfs send | pv | ssh <target> zfs receive