Median per-attachment archive time from the service logs. Same engine, same document, three environments. The difference was the mount configuration.
Archiving one attachment takes around 270 ms in test and 810 to 840 ms in production. A mail with nine or ten attachments takes more than eight seconds to file.
CPU, memory and the JVM on the archive tier are near idle. The application team says it is infrastructure. The infrastructure team says it is the application. Both are looking at their own layer and finding nothing wrong.
The engine was fast. Service log timing pairs showed 18 ms to open and 78 ms to save at the median, which is what a healthy Content Server does.
The storage path was not. The NFS volumes were mounted with 64 KB read and write sizes where 1 MB belonged, which means sixteen round trips to write a megabyte instead of one. Access-time updates were switched on for every read. A buffer volume that had been added to help sat on the same backing storage and made saves 16 percent worse.
Set 1 MB read and write sizes and noatime at the storage class, then restart. A restart, not a project: the lowest-risk change on the board.
Keep the timing pairs in the service logs so application time and infrastructure time stay provable before and after every change. Where inbound files can be pre-staged asynchronously, do that too, so the user request links rather than uploads: one to two seconds perceived, even on the slow estate.
16×the round trips to write one megabyte on a 64 KB mount, against one on a 1 MB mount
What the logs said
timings (1 day) open 18 ms save 78 ms storage 610 ms ← engine fast, storage slow system report NFS rsize/wsize 64 KB relatime on buffer volume on same backing store verdict failure mode ONE: mount configuration, not capacity fix 1 MB rsize/wsize, noatime, remove buffer volume, restart, retest with the same timings
