Measuring a Spreadsheet Instead of Trusting It
Data this note rests on: The measurement boundary stated as a design decision: 195 entries and 2,806 images were counted and every source link can be read as a string, while link availability, parcel transit and seller reliability were not measured and are not estimated.
Test design: sample, method, thresholds
This site can measure two things about a spreadsheet and refuses to estimate the rest. The two are the shape of a link string, read without sending any request, and the structure of a 195-entry pool carrying 2,806 images, read by counting fields. Link availability, parcel transit time and seller reliability are not measured here, not estimated here, and not implied by anything published here. That boundary is the design, not a gap in it, and this article exists to state it precisely enough that a reader can check it.
The sample is fixed and small enough to describe in one sentence. It contains 195 entries extracted on 2026-09-29 from the eight category views of one source marketplace, spread across 8 categories and 20 brands, carrying 2,806 images between them. Every entry has a recorded source link, a recorded price, a recorded image group, a category and a check week. Nothing in the sample was observed twice, and nothing was purchased, opened, or followed.
- Fix the sample before counting. The pool, the date and the source views are stated first, so a figure can never be produced by widening the sample until it looks right.
- Separate the two measurable objects. A link is a string, and the pool is a set of fields. They support different questions and they fail in different ways, so they are never combined into a single score.
- Publish the thresholds before the count. The index threshold is an image group of five or more, the status derivation uses two fields and that threshold, and the review week is uniform across the pool.
- State the exclusions in advance. Availability, transit and seller reliability are named as out of scope at the top, which prevents a later section from quietly answering them with a proxy.
What is out of scope, in one list: whether any link resolves, how long any parcel takes, whether any seller is dependable, whether any item arrived, and whether any price is still current. Each of those would need requests, purchases, or outcomes followed over time. None of them is measured on this site, and none is replaced with a correlated field dressed up as a proxy.
What was measured and when
Both passes ran against the pool of 2026-09-29, and the review week attached to every row is 2026-W40. There is no earlier pass and no later one, so every figure in the results section is a single observation of a single pool. The date matters more than the volume here, because a dated census can be reproduced and an undated one cannot.
The first pass reads link form. It detects the host and maps it to the platform the link belongs to, extracts the identifier parameter the platform uses and falls back to a trailing numeric segment when the parameter is absent, flags tracking parameters from the families that sharers add, and judges the overall form of the string. The whole pass is string analysis. It makes no network request, which is exactly why it can say something about a link without risking a false statement about a listing.
The second pass counts pool structure. It counts entries, categories, brands and distinct source platforms, orders the recorded prices to produce a minimum, a median and a maximum, counts image groups per entry and totals them, and applies the image threshold to partition the pool. Every one of those operations is a count or an ordering over fields that were captured on the extraction date, and each is reproducible from the same pool with the same rules.
| Question | Method | Sample | Status |
|---|---|---|---|
| Which platform does this link belong to | Host string analysis, no request made | The source link recorded with each of the 195 entries | Measured |
| Which identifier does this link carry | Identifier parameter extraction, with a numeric fallback | The source link recorded with each of the 195 entries | Measured |
| Does the link carry tracking parameters | Detection across known sharer parameter families | The source link recorded with each of the 195 entries | Measured |
| Does the link still resolve | Would need a request to the marketplace | No sample | Not measured, by design |
| How long does a parcel take | Would need parcels followed across their transit | No sample | Not measured |
| Is this seller dependable | Would need order outcomes attributed to sellers | No sample | Not measured |
| How large are the image groups | Counting one field per entry | 195 entries, 2,806 images in total | Measured |
| What does the pool cost | Ordering the recorded price field | 195 recorded prices from $5.17 to $451.10 | Measured |
- Source:
- The design of the two passes described in this article, together with the field rules published on the method page and implemented by the link health instrument.
- Sample:
- n = 195 entries for the pool pass. The link pass operates on the source link recorded with each entry, which is present for every row in this pool.
- Recorded:
- 2026-W40
- Known gap:
- The link pass runs per link on demand rather than once over the pool, so no frozen pool-wide tally of link states is published. The four questions marked as not measured are outside the reach of both passes and are left unanswered rather than estimated.
One consequence of the first pass deserves stating plainly, because it is the reason no link statistics appear in the results below. The link reading is a function rather than a dataset: it takes a string and returns a classification, and it produces the same answer for the same input every time. It is exposed as an instrument at /instrument/link-health/, where a reader can paste a link and read the result. What it does not produce is a frozen count of how many links in the pool sit in each state, because no such tally was written down as a dated artifact, and inventing one now from a function would be a reconstruction rather than a measurement.
A deterministic method is not an accurate one about the world. Reading a link string is exact about the string and silent about the listing. Saying that a link has a clean form is a statement a reader can verify by looking at the same characters, and it must not be read as a statement that anything behind those characters works.
Results
The results are three tables and one sentence. The tables report composition, price and image coverage for the pool of 2026-09-29. The sentence is that no link-availability, transit or seller result exists, because those questions were excluded before the work started rather than left unfinished.
| Field | Result | Note |
|---|---|---|
| Entries | 195 | Deduplicated by item identifier across the eight category views. |
| Categories | 8 | Counts per category run from 24 to 26 entries. |
| Brands | 20 distinct values | Distinct non-empty values in the recorded brand field. |
| Source platforms | Weidian for all 195 rows | One platform, with no second platform represented anywhere in the pool. |
| Extraction dates | 1, on 2026-09-29 | No second observation of any field exists. |
- Source:
- Field counts taken directly from the entry pool extracted on 2026-09-29.
- Sample:
- n = 195 entries across 8 categories and 20 brands.
- Recorded:
- 2026-W40
- Known gap:
- Composition describes one pool on one date. A single-platform pool cannot support a statement about spreadsheets in general, and none is made from it here.
| Measure | Result | Note |
|---|---|---|
| Lowest recorded price | $5.17 | In the Accessories category. |
| Median recorded price | $31.76 | Half the pool sits above this figure and half below it. |
| Highest recorded price | $451.10 | Also in the Accessories category, which is why that category resists a single number. |
| Category medians | $16.22 to $68.68 | From Headwear at the low end to Shoes at the high end, a factor of roughly four. |
- Source:
- The recorded price field of each entry in the pool extracted on 2026-09-29.
- Sample:
- n = 195 recorded prices, one per entry, with category medians computed over 24 to 26 entries each.
- Recorded:
- 2026-W40
- Known gap:
- One price per entry at one date. These figures describe spread between items and cannot describe movement of any item, which would require a second observation.
| Measure | Result | Note |
|---|---|---|
| Smallest image group | 1 image | One entry carries a single image. |
| Median image group | 14 images | The middle value across the pool. |
| Largest image group | 46 images | The largest gallery recorded on the extraction date. |
| Total images | 2,806 | The full volume carried by the 195 entries. |
| At or above the five-image threshold | 158 entries, 81% | These rows reach the index layer. |
| Below the five-image threshold | 37 entries | These rows stay out of the index and carry the unverified status. |
- Source:
- Image group counts per entry in the pool extracted on 2026-09-29, read against the index threshold published on the method page.
- Sample:
- n = 195 entries and 2,806 images; the smallest, median and largest groups are 1, 14 and 46 images.
- Recorded:
- 2026-W40
- Known gap:
- The table reports summary measures rather than a full distribution, because those are the figures the frozen record contains. No histogram is reconstructed from them.
The 37 unverified rows are the one result that could easily be misread, so it is worth being explicit: that count is produced by the threshold. The rule excludes any row whose image group holds fewer than five images, and 37 rows meet that condition. The unverified count is therefore not an independent quality finding that happens to equal 37; it is the threshold applied to one field, and a different threshold would produce a different number from the same pool without anything about the items changing.
What these results support: statements about the composition, price spread and image coverage of one pool on one date, and statements about the form of a link string when a reader supplies one. What they do not support: any claim about how many links work, how long anything takes, or how dependable any seller is.
Bias and limits
The limits here are not incidental, and listing them takes less space than defending against them would. Each one below describes something the measurement cannot see, and each one has a specific remedy that would require work this site has not done.
- Selection bias: the pool is one platform, eight category views and one date. It is what was visible in those views on that day, so any share computed from it describes this pool and not a market.
- Threshold bias: the five-image rule manufactures the unverified count. Reporting 37 unverified rows without the rule attached would turn a field comparison into an apparent quality judgement.
- Survivorship: an entry delisted before the extraction date never enters the pool and is not recorded anywhere. A pool built from what is currently visible systematically misses what has already gone.
- No outcome data: nothing here observes arrival, condition, cost at the door, or dispute results. A spreadsheet can be measured on its fields, and its fields do not contain outcomes.
- Determinism is not validity: the link pass returns the same answer every time, which makes it reliable and not thereby meaningful. Reliable readings of the wrong thing remain the wrong thing.
- Summary choices: a median is one decision about how to summarise a distribution, and the Accessories category shows how much a single number can hide. The range is reported alongside every median here for that reason.
Each limit has a remedy, and none of the remedies is a small edit. Selection bias needs a second platform and a second date. Threshold bias needs the rule published beside the count, which this site does. Survivorship needs a record of rows that were present earlier and are absent now, which requires two extractions. Outcome data needs a sample of parcels followed from order to delivery with a stated sampling rule, which is the largest piece of work in the list and the one this site is least likely to attempt. Determinism and summary choices are handled by stating them, which is what this section is for.
Any inference from this pool to the quality of a spreadsheet in general is unsupported. The measurements cover form and structure: which characters a link contains, how many images an entry carries, what the recorded prices look like as a set. They do not cover the three questions a buyer actually asks, and no amount of arithmetic on these fields will produce them.
The value of writing the boundary down is that it can be checked. A reader who wants to know what this site measures can read the rules at /method/, run a link through the checker at /instrument/link-health/, and compare the output with the description above. If any published figure turns out to come from a method outside these two passes, it is an error in this article rather than a change in policy, and it should be reported as one.