Part 4 · Chapter 4.16
You will produce: A tenant's stored bytes, metered daily and counted by kind, and the weekly comparison the clause asks for against what the object store actually holds. The comparison is the chapter: run against real data it says the meter charges 4,436 MB while the bucket holds 30.7 MB, and 99.9% of the difference is slots reserved and never uploaded to — a reservation measured against a delivery, which is FR-MED-01 working as designed rather than a drift. So the verdict has a third outcome, and the term that explains it cannot be computed from either store alone. Plus a column whose name reads as this chapter's subject and counts something else entirely, a materialised view that was counting zeros under a key it had no business creating, and a coverage number that was right about a service running in the wrong process · about 50 minutes including the exercise
Source: SRS — Software Requirements Specification · SAD — Software Architecture Document · docs/12-part-4-structure.md
Two clauses ask for this chapter, and one of them is a single row in a table of design requirements:
DR-17. Stored-bytes-per-tenant shall be maintained as a daily rollup summing
media_eventsdeltas (uploaded/deleted), reconciled weekly against an object-storage inventory listing — the media analogue of FR-ANL-06.
Everything in that sentence is buildable. A table, a materialised view, a listing, a comparison. What it does not tell you is what the comparison will say, and that turns out to be the chapter — because the first time it ran against real data it reported that the meter charges 4,436 MB while the bucket holds 30.7 MB, and neither number is wrong.
The quota charges a tenant the moment it hands out an upload slot. That is chapter 4.10's
decision and it is the only safe one: a slot issued without charging is a tenant reserving
without limit, so reserveMediaSlot sums declared_bytes over every non-rejected row of the
environment and refuses when the sum would exceed the cap. A pending object is already
charged.
The object store holds bytes somebody actually uploaded.
On the development lane those two populations barely overlap:
| rows | declared | |
|---|---|---|
pending — a slot taken, nothing uploaded | 6,580 | 4,436 MB |
ready — verified | 2,082 | 3,666 kB |
rejected — bytes destroyed, row kept | 237 | 233 kB |
| the bucket | 1,775 objects | 30.7 MB |
99.9% of what the quota charges is slots that were taken and never uploaded to. Nothing here is broken. It is what a development lane looks like when most of its media objects were created by tests that asked for a slot and then asserted something about the response.
But it means does the meter match the store is comparing a reservation against a delivery. A comparison with two outcomes has to call that either agreement — which over a 4.4 GB gap would be true and unbelievable — or drift, which would be wrong.
flowchart TB
subgraph meter["what the meter charges — 8,662 rows, ~4,440 MB"]
m1["PENDING 6,580 rows · 4,436 MB<br/>slots reserved, never uploaded to"]
m2["READY 2,082 rows · 3,666 kB"]
end
subgraph store["what the bucket holds — 1,775 objects, 30.7 MB"]
s1["objects with a media_objects row: 1,328"]
s2["keys with NO row at all: 447 · 24.3 MB"]
end
meter -->|"the comparison DR-17 asks for"| verdict
store --> verdict
verdict["a reservation against a delivery<br/>99.9% of the gap is outstanding slots<br/><br/>so the verdict has THREE outcomes:<br/>agree · reservations-only · a direction"]So the verdict has a third value. reservations-only says the two sides differ and the
difference is entirely accounted for. Only the part that outstanding reservations cannot
explain is a drift, and a store holding more than the meter knows about is the other
direction and is always wrong.
export type StorageVerdict =
| "agree"
| "meter-high"
| "meter-low"
| "reservations-only"
| "not-comparable"
| "no-data";This is chapter 4.7's shape arrived at from the other end. That chapter found a comparison whose two sides could not both exist for any tenant, and its product was a clause amendment rather than a green number. Here the two sides exist and measure different quantities, and the product is a third verdict. Both are the same lesson: a comparison that can only say agree or differ will be made to say one of them about a situation that is neither.
reservations-only needs a number: how many of the charged bytes belong to slots the store has
no reason to hold yet.
The analytical side cannot answer it. media_events records a reserved event when the slot is
issued and there is no uploaded event at all — the client PUTs straight to the store
(ADR-13), so the only two parties who know an upload finished are the client and the store, and
chapter 4.13 built a sweep precisely because nothing tells the platform. From the rollup's point
of view a reserved object and an uploaded one are the same record.
So the term comes from Postgres: the declared bytes of every pending row. Which is nearly
right, and wrong by a measurable amount.
267 of the lane's 6,580 pending rows name a key the bucket actually holds — 6,533,174 bytes,
21% of everything in it. Chapter 4.13's sweep is why that is ordinary rather than exotic: an
object is uploaded and stays pending until the sweep issues a HEAD and moves it. At any
instant some charged-and-uploaded objects are in both totals, and subtracting them as
outstanding removes bytes from both sides at once and invents a meter-low.
Which reservations are still outstanding is a question only the inventory can answer.
flowchart LR
ch["ClickHouse<br/>daily_usage_billing<br/>sum(stored_bytes_delta)"] --> f
mo["the object store<br/>one signed listing<br/>1,775 keys and their sizes"] --> f
pg["Postgres<br/>media_objects WHERE state = 'pending'<br/>object_key + declared_bytes"] --> f
f["reconcileStorage"]
f --> out["metered · inStore · reserved"]
note["the analytical side records 'reserved'<br/>and NO 'uploaded' event —<br/>so it cannot tell an outstanding slot<br/>from a delivered object"]
note -.->|"which is why there are three"| f
note2["and which reservations are STILL outstanding<br/>only the inventory can answer:<br/>267 of 6,580 pending rows<br/>name a key the bucket holds"]
note2 -.-> fChapter 4.7's reconciler reads two stores and the architecture document records that as the one component crossing constitution III's fence. This one reads three, and the third is not a convenience — it is what makes the second's number mean anything.
inventory: 1,775 objects, 30,756,595 bytes, 2 page(s)
tenants examined: 1,868 — rollup 84, bucket 297, outstanding reservations 1,657,
holding media in Postgres 1,678
keys belonging to no tenant: 83, 10,341,075 bytes
verdicts: agree 47 · meter-high 10 · meter-low 1 · not-comparable 1,784 · reservations-only 26
The headline is the fourth verdict. 1,784 of 1,868 tenants — 95.5% — cannot be compared at all, because the rollup knows 84 environments and Postgres knows 1,678 with media rows.
That is chapter 4.6's finding at its widest. A materialised view is a trigger on future inserts, so a rollup created late is permanently short — and here the producer ships in this chapter. The meter's history begins now; the bucket's does not. Every object that existed before this code was written is invisible to the analytical side and always will be.
Chapter 4.7 refused to collapse that into "missing data" and this chapter refuses again.
not-comparable is a verdict, it is counted, and it is by far the commonest answer the job
gives today. A report that folded it into agreement would describe a platform that agrees with
itself about 84 tenants and say nothing about the other 1,784.
The one meter-low is real and is the direction nothing explains: metered 0, bucket 6,852,575
bytes, outstanding reservations 16,266. Bytes the meter has never charged for. On a development
lane that is debris; in production it is a leak, and it is exactly what DR-17 exists to find.
DR-17 says summing deltas, and that is a choice with a price the clause does not state.
A sampled level is self-correcting. Ask the database how many bytes a tenant is storing, write the answer down, and any error you made yesterday is gone today. A summed level is not: it is yesterday's answer plus today's changes, so an error made once is carried for ever.
What is one lost record worth? The honest answer is a distribution, not a number.
flowchart LR
subgraph dist["one lost record, over 8,941 chargeable objects"]
a["min<br/>1 B"]
b["p50<br/>1,024 B"]
c["mean<br/>521,671 B"]
d["p99 and max<br/>26,214,400 B"]
end
b -->|"509x"| c
subgraph share["as a share of the tenant's own level, 925 tenants"]
e["the median tenant's<br/>median object: 14.89%"]
f["its WORST object: 50.00%"]
g["the worst case anywhere: 99.96%"]
end
dist --> perm["and it never self-corrects:<br/>a sum of deltas is re-sampled by nothing"]
share --> permThe mean is 509 times the median. A chapter that published "a lost record costs about half a megabyte" would be stating something true of no object in particular: half the objects on this lane are a kilobyte or smaller, and the largest are the per-kind cap exactly. As a share of the tenant's own level, over the 925 tenants holding more than one object, the median tenant's worst object is 50.00% of everything it has, and somewhere on the lane there is a tenant whose single largest object is 99.96% of its level.
And a record can be lost. The publish happens after the commit and outside the transaction,
which webhooks/analytics.ts argued three movements ago and this chapter inherits:
void publishStorageDelta(this.analytics, this.logger, {
environmentId: this.repo.environment,
mediaId: id,
cause: "reserved",
kind,
bytesDelta: input.bytes,
occurredAt: new Date(),
});A crash between the commit and the publish loses the record. The alternative is worse in two directions: a blocked outcome transaction is a customer's uploads stopping because a metering pipeline is unwell, which constitution III names as a design failure; and a publish inside a transaction that rolls back emits a delta for a slot that was never created — a permanent overcount, of the kind this whole section is about.
flowchart TB
subgraph delta["a summed delta — what DR-17 asks for"]
d1["fires on INSERT, needs no runner"]
d2["one lost record is wrong FOR EVER"]
d3["bounded by the rollup's 25-month TTL"]
end
subgraph sample["a sampled level — the alternative"]
s1["self-correcting: the next sample is the truth"]
s2["needs something to run daily"]
s3["this platform has NO scheduler<br/>zero schedule: triggers, ADR-28"]
end
delta --> chosen["chosen, and the clause chose it first"]
sample --> refused["unbuildable here, not merely unchosen"]
chosen --> repair["so the inventory is the re-base:<br/>the store holds the level directly,<br/>not as a sum"]There is a second cost, and no clause had stated it. The rollup carries a 25-month TTL — DR-09's "daily aggregates for 25 months", applied in chapter 4.6 — so a level summed from the beginning of time is short by exactly the deltas that expired, permanently, with nothing in the system able to notice. Revision 1.23 of the specification now says so, and names the inventory as the re-base: the object store holds the level directly rather than as a sum, so the reconciliation can restate a truncated figure where no amount of re-reading the rollup can.
The per-kind counts are one expression in one materialised view, and the first version of it was wrong in a way that only a test could show.
sumMap(map(kind, toUInt64(if(event = 'reserved', 1, 0))))Every event contributes its kind as a key. A rejection of an audio file contributes
{'audio': 0}. So a day holding one image reservation, one audio rejection and one image
rendition answers {'audio':0,'image':1} — an audio key in a column called uploads_by_kind,
on a day nobody uploaded any audio.
The values are right. What is wrong is the key set: it stops meaning the kinds this tenant uploaded and starts meaning the kinds that had any media event, and a caller iterating the map reports a kind belonging to a different question.
sumMapIf(map(kind, toUInt64(1)), event = 'reserved')That emits no key for a row that does not match, and {} when no row in the group does — which
is the rule the ingester's other rollup read already states: a day with no activity is a missing
row, never a row of zeros.
The zeros were also stable. SummingMergeTree drops a row whose summed columns are all zero; it
does not drop a zero-valued key inside a map. Asked of the server with OPTIMIZE … FINAL:
{'audio':0,'image':1,'video':0} before the merge and the same after. A wrong answer that a
merge would have cleaned up is a nuisance; one that survives is the kind that gets published.
DR-17 asks for an inventory listing and nothing in this platform could list a bucket. It turned
out to be nearly free: the signer written in chapter 4.10 already documents the key-less case,
and GET /{bucket} with no key is a listing.
What was not free is that one response carries a thousand keys and says IsTruncated: true.
Chapter 4.13 shipped a sweep that read one page, and the consequence was an object nobody
uploaded to staying pending for ever. A reconciliation that reads one page reports agreement
about every tenant it never looked at, which is the same defect inside the instrument written to
catch it. The listing pages on marker, and it reports the page cap rather than swallowing it:
a truncated listing makes every tenant's inStore null, and the script exits 2 without printing
a verdict.
It also cost the signer a small change with a sharp reason. Appending &marker=… to an
already-signed URL answers SignatureDoesNotMatch — every parameter must be inside the
signature — so presign had to take extra parameters, and the moment a caller could add one,
the five hand-written entries stopped being safely ordered and became a sort.
The keys carry the tenant because chapter 4.15 kept the platform's layout: an object's key is
${environment_id}/${id}, so a listing can be split per tenant with no database join at all.
Heading each object instead would give the same answer for 73 times the cost — 1,972 ms
against 27 ms for 1,775 objects, serial; 650 ms at eight at a time, which is the figure worth
stating because the ratio depends entirely on a number the earlier research never named.
And 83 keys belong to no tenant — probes and test debris under four prefixes, none of them on the listing's first page, which is how a one-page measurement earlier in this feature concluded every key was tenant-prefixed. They are reported as their own category rather than skipped, because a bucket full of debris should not report clean.
The percentages in this chapter's distribution are facts about a lane whose median tenant holds a handful of objects. On a tenant with ten thousand photographs, one lost record is not 50% of anything. The absolute figures — a kilobyte at the median, twenty-five megabytes at the cap — are facts about the size distribution and the per-kind caps, and those transfer.
The rollup-versus-raw measurement needed a corpus the lane does not have: 240,000 events over 90 days, planted for one tenant. It reads 276 rows and 5,456 bytes where accumulating the same answer from the raw table reads 89,823 rows and 2,874,336 bytes — 527 times the bytes for twice the wall clock, because 2.8 MB is nothing to this engine. DR-17's argument is about bytes scanned and the clock barely shows it, which is worth saying plainly rather than publishing the ratio that flatters it.
And the reconciliation cannot yet fail in one of the two directions it exists to watch. Nothing
deletes a media_objects row — the rejection path subtracts the bytes and keeps the row on
purpose — so no deleted event is ever published, and the failure a summed level is most
exposed to, a lost deleted leaving the meter permanently over, cannot arise. A run of green
verdicts is evidence about three causes and not four.