Building Relay

Part 4 · Chapter 4.15

What a thumbnail costs

You will produce: A thumbnail for every uploaded image, and the answer to what one costs — which is not the number anybody reaches for first. Across a real corpus the ratio to the parent spans 775x while the thumbnail's own size spans 4x, so the ratio is a fact about the parent and publishing it as the cost publishes the wrong variable. Below the bound a thumbnail is 97.3% of the parent and the same pixels, which decides a behaviour rather than a figure. Plus the clause's last five words, which turn out to be the hard part: a derived object is named by no message, so one shipped clause refuses to serve it and another reaps it after a day, both working exactly as written — and nothing in the schema could say that two rows were related. The dependency that had to be argued for without borrowing the previous chapter's argument, the fourth round trip nothing had needed before, and a transaction the plan said was there · about 45 minutes including the exercise

Source: SRS — Software Requirements Specification · SAD — Software Architecture Document · docs/12-part-4-structure.md

FR-MED-05 is one sentence and it reads like two easy halves and a footnote:

The system shall generate a thumbnail for images and a poster frame for videos during processing, stored as derived objects sharing the parent's lifecycle.

The footnote is the chapter. Stored as derived objects sharing the parent's lifecycle sounds like a note about where files go. It is the only part of that sentence that needed a migration, and it is the part two clauses this platform already shipped would otherwise destroy.

The clause's last five words

A thumbnail is referenced by no message. The message names the parent; the thumbnail is something the platform made afterwards. That single fact runs into two rules that are already live and already correct.

FR-MED-08 refuses to serve it. Delivery authorises an object by asking which channels hold a message referencing it, then asking whether the caller may read any of those channels. For a thumbnail that list is empty, so nobody may read it — and that is not a bug, it is the clause's own stated rule: an object with no referencing message is readable by nobody, including the credential that uploaded it.

FR-MED-10 would reap it. Unreferenced objects are hard-deleted after 24 hours. A thumbnail is unreferenced by construction, so it would be destroyed a day after it was made.

Both clauses are working exactly as written. What is missing is the platform's ability to say that two rows are related at all — media_objects had no self-reference, no parent column, and nothing distinguishing an object a client uploaded from one the platform produced.

flowchart TB
    m["a message"] -->|names| p["the parent object"]
    p -.->|"named by NO message"| t["the thumbnail"]
    subgraph consequences["what two shipped clauses do about that, both correctly"]
      c1["FR-MED-08<br/>authorised through referencing channels<br/>-> no message, so NOBODY may read it"]
      c2["FR-MED-10<br/>unreferenced objects reaped after 24h<br/>-> DELETED a day after it is made"]
    end
    t --> c1
    t --> c2
    fix["migration 0020<br/>parent_id + composite FK + ON DELETE CASCADE<br/>'its reachability IS its parent's'"]
    fix -.->|answers| c1
    fix -.->|answers| c2
Two correct clauses, and the thing neither of them can see.

The cost is not a ratio

The chapter is called what a thumbnail costs, so here is the measurement that matters, taken over the images this workspace actually contains:

parentparent bytesthumbnailratio
an animated GIF, 748×386, 777 frames3,175,0632,5580.08%
an animated GIF, 488×383, 300 frames4,631,8866,0700.13%
a diagram, 969×58868,5274,1186.01%
a diagram, 933×794107,24810,2589.56%
an icon, 460×46032,3707,10421.95%
a logo, 620×9611,0476,84861.99%

The ratio spans 775×. The thumbnail's own size spans 4× — every one lands between 2,558 and 10,258 bytes, because the output's size is a fact about the 320 px bound and not about whatever the tenant uploaded.

So "a thumbnail costs 2% of the original" is not a sentence this platform can publish. The denominator is a property of somebody else's file. A thumbnail costs about 7 kB. The ratio is a fact about the parent, and publishing it as the cost publishes the wrong variable.

flowchart LR
    subgraph ratio["the ratio to the parent: 775x spread"]
      r1["3,175,063 B gif -> 2,558 B<br/>0.08%"]
      r2["68,527 B png -> 4,118 B<br/>6.01%"]
      r3["11,047 B png -> 6,848 B<br/>61.99%"]
    end
    subgraph size["the thumbnail itself: 4x spread"]
      s1["2,558 B"]
      s2["4,118 B"]
      s3["6,848 B"]
      s4["10,258 B"]
    end
    bound["the 320 px bound decides the output,<br/>the parent decides only the ratio"]
    bound --> size
Two spreads over the same six images. Only one of them is about the thumbnail.

Where the crossover is, and what it decides

The real corpus has no size series, so this one is synthetic: pixel noise encoded as JPEG, which does not compress and is therefore the most favourable case for the ratio. Said plainly rather than passed off as photographs.

parentparent bytesthumbnaildimsp50ratio
120×12010,2009,920120×1202.7 ms97.3%
200×20027,40026,662200×2005.9 ms97.3%
320×32068,31266,544320×32013.4 ms97.4%
400×400106,18258,226320×32013.6 ms54.8%
640×480204,28435,178320×24011.1 ms17.2%
1920×10801,376,62213,258320×18015.2 ms1.0%
4000×30007,984,3174,100320×24050.8 ms0.1%

At or below the bound the thumbnail is 97.3% of the parent and the same pixels. The tenant would store the image twice to save 2.6%.

That is a behaviour, not a figure. An image already inside the bound gets no rendition at all, and its delivered form simply carries no thumbnail key. The rule came out of the table; nobody chose it.

One more thing in that table is worth a second look: the time is not monotone in the parent's pixels. 320×320 costs 13.4 ms and 640×480 costs 11.1. Below the bound the encoder is writing a full-size output, so the work is in the encode rather than the resize. It is the kind of curve a reader assumes is flat.

A fourth round trip, and why the worker had to learn to write

The media worker reads objects. Before this chapter it made three calls to the store and held none of them:

A thumbnail needs the bytes. A stream that ClamAV has already drained cannot be read again, and a prefix is not an image. So a rendition needs a whole-object GET that nothing in this service had ever made — and the alternatives were worse. Teeing the scan stream into a buffer costs no extra request and buffers every object, including the 100 MB video cap, because the scan runs before anything knows the type. One buffered GET up front has the same problem for the same reason.

flowchart LR
    h["signed HEAD<br/>2.793 ms"] --> g["full GET, 7,984,317 B<br/>8.2 ms"]
    g --> r["resize to 320<br/>52.0 ms"]
    r --> tot["60.2 ms total"]
    rss["peak RSS moved 0.4 MB<br/>arithmetic for the bitmap said 36 MB<br/>libvips works in strips"]
    r -.-> rss
The whole bill for one 8 MB image, measured rather than estimated.

The fetch turned out to be the cheap part: 8.2 ms against a 52.0 ms resize, 14% of a 60.2 ms total. And the memory worry did not survive contact either. The arithmetic for a 4000×3000 RGB bitmap is 36 MB, and peak RSS moved 0.4 MB — libvips processes in strips and never holds the whole thing. That figure was written down as an upper bound from arithmetic and explicitly not published as a measurement until something measured it, which is the only reason the 36 MB did not end up in this table.

The gate has three states, and the third one is an ordinary photograph

Deciding whether to spend that fetch looks like two questions — is it an image, and is it bigger than the bound. The probe answers both from a 64 KiB prefix.

Except when it does not. A JPEG stores its dimensions in an SOF0 marker somewhere after the start, and everything before it is segments the decoder walks past. One of those segments is APP1, which holds EXIF — camera, lens, timestamp, GPS, and frequently an embedded preview image. A maximal APP1 is 65,535 bytes, and the probe window is 65,536.

Measured: a JPEG carrying one maximal APP1 — 72,215 bytes, which is what an ordinary camera file looks like — reports no dimensions at all from the prefix, while the same bytes read whole decode as 1200×900.

A two-state gate skips those. It would deny thumbnails to exactly the files most likely to want one, and it would do it silently, because "we could not read the dimensions" and "the image is small" arrive at the same branch. The third state fetches and lets the decoder answer, which costs nothing extra: by then the bytes are in hand.

flowchart TB
    probe["the probe: sniffed type + a 64 KiB prefix"]
    probe --> q1{"an image the decoder reads?"}
    q1 -->|no| none["no rendition, no reason<br/>audio and video leave here"]
    q1 -->|yes| q2{"dimensions?"}
    q2 -->|"known, <= 320"| skip["no fetch, no rendition<br/>R2: the output would be 97.3% of the parent"]
    q2 -->|"known, > 320"| get["fetch the whole object"]
    q2 -->|"UNKNOWN"| get
    get --> made["resize, PUT, record"]
    note["UNKNOWN is an ordinary camera JPEG:<br/>one maximal APP1 segment (72,215 B) pushes<br/>SOF0 past the 64 KiB probe window"]
    note -.-> q2
Two states would have been wrong for a category of file, not an edge case.

What it cost to choose the dependency

Nothing in this platform could decode an image. dimensions.ts reads headers — a PNG IHDR, a GIF screen descriptor, a WebP VP8, a JPEG SOF0 — and touches no pixels. Four options, priced against the worker's own base image:

optionaddedcovers the four allowed typesp50 for 1920×1080
sharp, in-process30,380,799 Byes15.2 ms
ImageMagick, one subprocess per object28,936,284 Byes35.8 ms
ffmpeg, subprocess+113,994,336 Byes, badly—
a sixth containera whole servicedepends—
pure TypeScript0 BPNG only—

The size did not decide it. 30.4 MB against 28.9 MB is 5% on a choice between two fundamentally different relationships, and anyone re-running this should expect them to stay close. What decided it is 2.4× and the shape: a subprocess pays its extra 20 ms as process spawn, on every object, in a service whose entire job is a sweep over a backlog.

Pure TypeScript is out on coverage rather than effort. node:zlib gives inflate, which reaches PNG; JPEG needs a DCT decoder, GIF needs LZW and WebP needs VP8. One of four is not a feature.

The video half is not built, and the reason is in the same table. A poster frame needs a video decoder, and the tool for it costs 114 MB — 3.75× the image half — for the harder half of a clause whose easier half this platform already declined. Chapter 4.13 recorded FR-MED-04 as partly met because duration needs four container parsers and MP3 variable bitrate has no header answer. A frame is strictly harder than a duration. It is recorded as unmet by decision, with the reversal condition: when video attachments are a measured fraction of stored objects, or when the platform already carries ffmpeg for something else, 114 MB buys two things and the question is re-opened.

Three instruments that lied while measuring this

The dependency table above took four attempts, and every wrong number looked exactly like a right one.

apk add imagemagick cannot decode a JPEG. The first size measurement returned 27,524,848 B for a build that answers no decode delegate for this image format. Alpine ships the format delegates as separate packages. A dependency's install size is not its usable install size, and the number that looked right was for an ImageMagick that reads none of the four types this platform allows.

The repair returned a zero that read as "already present". Adding four delegate packages in one apk add reported a delta of 0 B. Two of the four do not exist; apk is all-or-nothing, so the transaction failed whole and installed nothing — and with stderr redirected, the failure presented as no change needed.

A grep with no positive control published a confident zero. Asking magick -list format which of the four types were supported printed nothing, which read as no support. The list writes JPEG* with a trailing asterisk, and the pattern could not match. Re-run with a control — JPEG must appear, because a conversion had just succeeded — the real answer is GIF* JPEG* JPG* PNG* WEBP*.

What the lane cannot tell you

One honest limit, stated where a reader will see it rather than buried in a record.

This platform's development lane holds 6,535 media objects, 5,996 of them images. Four are above the thumbnail bound. Every figure in this chapter that came from real files came from six images of convenience — two icons, two diagrams and two animated GIFs, none of them a photograph — and every figure with a size series came from synthetic noise.

So the ratios demonstrate that the ratio is unstable. They are not an estimate of what a real tenant would see, and the crossover is a property of the encoder and the bound rather than of anybody's traffic. The measurements that are transferable are the ones about this platform: the fetch is 14% of the cost, libvips does not hold the bitmap, and a thumbnail is about 7 kB.