Comparison Tables and Matrices
Comparison tables and matrices are among the most useful—and most easily misused—tools in collector research. They can reveal relationships that are hard to see in prose: which characteristics travel together, where a specimen breaks an expected pattern, whether several records describe the same physical object, or which evidence supports competing identifications.
Their value does not come from arranging facts neatly. It comes from imposing discipline on the comparison: the examples must be comparable for the question being asked, the criteria must be defined, observation must remain separate from interpretation, missing information must stay visible, and every important entry must remain traceable to its source. A strong matrix is therefore not merely a presentation device. It is a research instrument.
A table can look stronger than the evidence inside it
A collector is testing whether a boxed set belongs to a suspected early issue. They create a grid comparing three reference copies against their own. The rows list box colour, product code, page count, tray colour and leaflet. Most cells appear to match, so the conclusion feels obvious.
But one reference is a seller photograph, another is a catalogue description with no image, and the third is a forum post that may repeat the same catalogue entry. The tray in the collector's copy is also replaceable, and the colour photographs were taken under different lighting. The matrix has made unlike evidence look equal.
The answer is not to abandon the table. It is to redesign it so the collector can see the object, the source, the observation, the interpretation, the confidence and the evidential significance as different things.
What the collector is actually building
Comparison table and comparison matrix are related, but not identical
A comparison table is usually descriptive. It places objects, editions, records or sales against a shared set of attributes so that presence, absence and difference can be inspected side by side. It is often the best format for summarising established distinctions.
A comparison matrix adds an analytical relationship. It may map evidence to competing hypotheses, compare every specimen with every other specimen, show source agreement and dependence, score degrees of similarity, or track how characteristics change through time. The matrix is not automatically more rigorous; it is simply designed to test relationships rather than only display attributes.
The collector should choose the simpler form that answers the question. A matrix that introduces scoring, weighting or pairwise comparison without a genuine analytical need creates complexity without improving judgement.
Begin with the research question, not with empty columns
A vague prompt such as ‘What is different about these?’ produces a table full of visible details but little analytical value. A stronger question defines the population, the candidate groups and the distinction that matters: ‘Which physical and documentary characteristics reliably distinguish the UK-distributed issue from the US issue?’
That sentence controls the whole design. It determines which examples qualify, which rows matter, which measurements need to be normalised and what kind of conclusion the table can support. If a row does not help answer the question, it belongs in a separate note or a different matrix.
- Identity question: which issue, printing, state or configuration best fits the evidence?
- Completeness question: which components are original, optional, regional, replaced or still uncertain?
- Source question: where do references agree, conflict or depend on one another?
- Market question: are the sales genuinely comparable after condition, completeness and venue are controlled?
- Action question: what can be learned safely, and which intervention would risk damaging evidence?
Define the unit of comparison before gathering data
A physical specimen, an edition, a printing, a variant, a configuration, an auction listing and an observation event are not interchangeable units. An auction listing is a record about an object, often incomplete and sometimes wrong. A photograph is an observation route, not the object itself. A named edition may combine several production runs.
A defensible matrix identifies exactly what each row or column represents. This prevents one physical object from being counted several times through repeat sales, and it prevents a source record from receiving the same authority as direct inspection.
Keep the analytical levels separate
The most reliable data model distinguishes the object, the source describing it, the observation made from that source and the interpretation built from that observation. These levels can be linked, but they should not be flattened into one unqualified cell.
Object level
Physical specimen
One identifiable object in a collection, archive, auction or sale. Assign a persistent reference so later listings are not mistaken for new examples.
Classification level
Edition, printing, variant or state
A proposed or established production category. These labels are interpretations and should not replace the raw specimen observations used to justify them.
Assembly level
Configuration
A particular combination of packaging and contents. It may be factory-original, transitional, region-specific, assembled later or still unexplained.
Documentary level
Source record
A catalogue entry, listing, book, database record or collector report. It may describe a specimen well, badly or derivatively.
Evidence level
Observation event
An examination made on a specific date, by a named observer, under stated conditions and possibly using particular equipment.
Reasoning level
Interpretive claim
A conclusion drawn from one or more observations, such as ‘possible replacement’, ‘consistent with a later printing’ or ‘likely same object as sale B’.
Observation, meaning and collector risk
A matrix becomes auditable when the reader can see the route from what was observed to what it may mean. The following examples show why a verdict should not be hidden inside an observation field.
Tray colour
Evidence
The tray is translucent grey and leaves a 3 mm gap at the right edge.
Meaning
The tray differs from four directly inspected complete references, all of which use opaque black trays with a closer fit.
Collector risk
Calling it an ‘incorrect replacement’ converts a useful observation into a verdict before production variation, shrinkage and later substitution have been tested.
Copyright wording
Evidence
The line ends with ‘Inc.’ and is clearly visible in a perpendicular high-resolution photograph.
Meaning
The wording is compatible with documented later examples and differs from the earlier reference wording.
Collector risk
One matching line may support a classification, but it should not overrule contradictory construction or provenance evidence.
Absent insert
Evidence
No insert is visible in the seller's photographs; the internal area is partly obscured.
Meaning
The insert is not assessable from the available record.
Collector risk
Recording ‘absent’ would transform missing evidence into evidence of missingness and may distort both completeness and rarity claims.
Compare like with like—or state exactly why the comparison is imperfect
Comparability is not a general property of two objects. It is a property of the objects in relation to the specific judgement being made. A restored copy may still be useful for typography, but weak for original finish. A seller photograph may support a product-code observation, but not a precise colour comparison.
Appropriate
Comparable for this question
- One printing against another printing of the same product.
- Regional versions where territory, language and distributor are recorded separately.
- Known complete examples of the same boxed configuration.
- Several auction appearances believed to show one physical object.
- Competing catalogue descriptions of the same proposed edition.
Conditional
Comparable only after context is controlled
- A sealed example and an incomplete loose copy.
- A restored object and an untouched object.
- Photographs taken under substantially different lighting or perspective.
- A prototype and a retail production object.
- Objects separated by major material, format or manufacturing changes.
Source warning
Not independent evidence
- Five websites repeating one catalogue statement.
- Dealer descriptions copied from an earlier auction listing.
- Different crops of the same stock photograph.
- Several reports ultimately based on one unverified reference copy.
- Repeat sales of one object counted as separate surviving specimens.
Choose comparison dimensions that can change the judgement
Rows are editorial decisions. They reveal what the researcher believes matters. The collector should select criteria because they help distinguish competing explanations, not because they are easy to photograph or because they favour a preferred answer.
Identity and bibliographic data
- Publisher, maker, artist, product code, catalogue number, ISBN, copyright line, language, territory, distributor, printer, factory and registration marks.
Physical construction
- Dimensions, weight, material, thickness, mould profile, seams, fasteners, binding, stitching, adhesive, surface finish, coating, embossing, component fit and tolerances.
Graphic and typographic details
- Logo form, font, spacing, line breaks, crop, border width, printer marks, punctuation, trademark symbols, corrections, halftone and print registration.
Contents and completeness
- Component counts, titles, page counts, inserts, maps, cards, dice, trays, bags, certificates, promotional material, errata and regional notices.
Condition and intervention
- Structural integrity, wear, colour retention, staining, corrosion, brittleness, warping, active deterioration, restoration, replacement and presentation.
Provenance, market and source quality
- Ownership history, acquisition, seller, auction, exhibition, publication, restoration history, sale date, realised price, venue, direct inspection, institutional record and hearsay.
Diagnostic value depends on stability
A descriptive characteristic records what an object is like. A diagnostic characteristic helps distinguish between candidate identities. A feature becomes diagnostic only when it is consistently associated with a group, sufficiently stable, reliably observable and tested against plausible alternative explanations.
Often stronger
Relatively stable attributes
Engraved mould marks, printed product codes, page signatures, plate changes, fixed text revisions, die-cut profiles and documented component counts are often more dependable because they arise from controlled production decisions.
Use tolerances
Production-variable attributes
Ink density, trim position, staple alignment, plastic shade, weight, glue spread, registration and minor dimensions may vary within one run. Record ranges and method before treating them as diagnostic.
High caution
Post-production-variable attributes
Colour, gloss, flatness, odour, flexibility, completeness, component pairing, packaging seals and surface texture may change through ageing, storage, handling, restoration or later assembly.
Controlled terminology and visible missing data
Equivalent concepts recorded under inconsistent names fragment the evidence. ‘First issue’, ‘first edition’, ‘first printing’ and ‘version one’ may not mean the same thing. ‘UK edition’ may describe manufacture, distribution, language, seller location or nothing more than a collector convention.
Every important field should have a definition, allowed values, a rule for uncertainty and notes for edge cases. Blank cells should be prohibited because a blank has no stable meaning.
| Code | Meaning | Use |
|---|---|---|
| P | Present | The feature was positively observed. |
| A | Confirmed absent | Inspection was sufficient to establish genuine absence. |
| NV | Not visible | The available image or view does not expose the feature. |
| NR | Not reported | The source does not mention the feature. |
| NA | Not applicable | The field does not apply to this unit of comparison. |
| NT | Not tested | The feature could be assessed, but no examination has been made. |
| U | Uncertain | A tentative observation exists but cannot be resolved confidently. |
Do not confuse unknown with low confidence
Unknown means the evidence is not available. Low confidence means a tentative judgement has been made from limited evidence. Keeping those states separate prevents missing information from being converted into a weak assertion.
Report the denominator as well as the percentage: ‘Black tray observed in 8 of 10 assessable examples; 6 additional examples could not be assessed from available photographs’ is far more informative than ‘80% had black trays.’
Select the matrix type that matches the research task
Different matrices support different decisions. Combining identity, condition, provenance, value and intervention in one wide grid usually reduces clarity. Build several focused views from the same underlying evidence where necessary.
Edition-difference table
Best for: Summarising established differences between versions or printings.
Must show: Stable distinguishing criteria, sample size, exceptions and dated evidence.
Caution: Do not turn one unusual specimen into a formal edition rule.
Specimen observation matrix
Best for: Comparing multiple physical copies before assigning them to groups.
Must show: Raw observations, source route, missingness and persistent specimen IDs.
Caution: Keep proposed classification outside the observation fields.
Component-completeness matrix
Best for: Games, boxed toys, kits, portfolios and mixed-media sets.
Must show: Manufacturer evidence, sealed examples, period reviews, physical fit and replacement risk.
Caution: Frequency among survivors does not prove original inclusion.
Evidence-to-hypothesis matrix
Best for: Authentication, attribution and disputed identification.
Must show: Several plausible explanations and evidence that supports, weakens or contradicts each.
Caution: Do not convert qualitative judgements into unsupported probabilities.
Source-agreement matrix
Best for: Reconciling books, catalogues, auction records and collector claims.
Must show: Source date, type, access to the object, evidence shown and likely dependency.
Caution: Repeated claims are not independent confirmations.
Chronology matrix
Best for: Logo changes, packaging revisions, printings and production sequences.
Must show: Dated advertisements, invoices, catalogues, provenance or other external anchors.
Caution: A visually neat sequence is not automatically the historical sequence.
Condition-axis matrix
Best for: Explaining why similar headline grades conceal different risks.
Must show: Structure, surface, colour, completeness, restoration, active deterioration and presentation separately.
Caution: Condition, originality, completeness and visual appeal are different judgements.
Market-comparable matrix
Best for: Valuation research where sales require normalisation.
Must show: Issue, date, realised versus asking price, condition, completeness, venue, premium and provenance.
Caution: Raw price rows without context create false comparability.
Intervention decision matrix
Best for: Choosing research actions where inspection, opening or sampling could alter evidence.
Must show: Research value, damage risk, cost, reversibility and specialist need.
Caution: The most informative action may still be disproportionate or destructive.
Research-gap matrix
Best for: Turning uncertainty into a focused next-step plan.
Must show: Evidence held, evidence missing, best next source and likely value of obtaining it.
Caution: Do not gather more data merely because a cell is empty; prioritise decision-changing gaps.
The matrix reveals patterns; it does not explain them automatically
Missingness
Binary fields erase important states
‘No’ may mean absent, not visible, not reported, not tested, not applicable or uncertain. Use controlled codes and prohibit ambiguous blanks.
Pattern
Correlation is not causation
If blue boxes always contain revised booklets, the matrix reveals association. It does not prove the changes were introduced together. Later pairing, regional distribution or sample dependency may produce the same pattern.
Classification
A cluster is not yet a named variant
Traits that repeatedly co-occur may define a coherent group, but transitional manufacture, substituted contents, dealer assembly and undocumented overlap remain alternative explanations.
Sampling
Frequency is not production rarity
Collector datasets overrepresent high-condition, expensive, complete and unusual objects. Two documented examples may be rare, poorly recorded or merely concentrated outside the sampled community.
False completeness
A full table can be a warning sign
Research data normally contain gaps. A matrix with no unknowns may indicate that uncertainty has been forced into categories, assumptions have been copied into cells or blanks have silently become ‘no’.
False precision
Numbers can create false authority
Scores such as ‘92% match’ or ‘83% likely first printing’ are structured opinions unless the model, weights, independence assumptions and reference population have been validated.
Transitional examples deserve provisional language
Factories use old packaging with new contents, exhaust component stocks gradually, run several suppliers, correct material mid-run and distribute different configurations simultaneously. A specimen that combines traits from two proposed groups should not immediately be labelled an error, hybrid, counterfeit or rare variant.
‘Mixed configuration of uncertain cause’ is often the more accurate description until repeated, independently documented examples establish a production pattern.
Build the matrix as a reproducible research process
The strongest matrix prevents the researcher from changing the comparison rules after seeing the desired result. The workflow below keeps criteria, sources, missingness and alternatives visible from the beginning.
Frame the question
Write one sentence stating exactly what the comparison must establish.
Quality test: Could a reader tell which conclusion the table is allowed to support?
Define the population
Set inclusion and exclusion rules for specimens, records, sales or observations.
Quality test: Are obvious counterexamples being excluded merely because they complicate the pattern?
Establish identities
Assign persistent IDs and reconcile repeat listings before counting examples.
Quality test: Could one object appear more than once in the dataset?
Choose criteria
Select only characteristics relevant to the research question and classify their likely stability.
Quality test: Would any row change the conclusion, or is it only descriptive decoration?
Create a data dictionary
Define each field, allowed value, unit, uncertainty code and edge-case rule.
Quality test: Would two researchers record the same observation in the same way?
Gather raw observations
Record what can be seen or measured before assigning group labels.
Quality test: Has interpretation entered the observation field?
Attach sources
Make every significant entry traceable to inspection, image, document or correspondence.
Quality test: Can a later reader recover the evidential route behind the cell?
Normalise and expose gaps
Standardise units, dates, terms and photography, while coding missing information explicitly.
Quality test: Are unlike measurements or ‘unknown’ states being treated as direct matches?
Test the pattern
Look for clusters, outliers, contradictions, transitional combinations and source dependency.
Quality test: What alternative explanation could produce the same arrangement?
State and version the finding
Use proportionate confidence, record limitations and preserve revision history.
Quality test: Does the conclusion distinguish established differences from working hypotheses?
Measurements and images need their own controls
Measurements need method and tolerance
A value such as ‘210 mm’ looks precise but says little without the quantity measured, instrument, orientation, object state, packaging inclusion, precision and estimated uncertainty. A better entry records ‘cover width 210.4 mm, measured closed at the centre edge with digital callipers; estimated uncertainty ±0.3 mm.’
The evidential importance of a difference depends on manufacturing context. One millimetre may be diagnostic in a precision-machined part and meaningless in hand-trimmed paper packaging. Compare ranges and tolerances, not isolated numbers.
Photographs must be normalised before visual differences are trusted
Image matrices are valuable for logos, marks, seams, signatures, typography, tool marks, wear patterns and component shapes. They are also vulnerable to angle, scale, crop, lighting, white balance, glare, lens distortion, compression and digital processing.
A strong image comparison preserves full-object context, a standardised close-up, scale where appropriate, source and date, and a warning wherever direct comparison is unreliable. Apparent colour or size differences caused by photography should never be promoted into a variant claim.
Preserve data provenance, not only the polished view
A wide comparison table is readable for humans, but it becomes difficult to maintain once every cell needs a source, observer, date and confidence. The underlying research is often better preserved in long form: object ID, attribute, value, source, observer, date and confidence as separate fields.
Focused wide matrices can then be generated for a particular question without losing the source behind each entry. This also allows one corrected observation to update several analytical views without rewriting the evidence manually.
Documentation checklist
- Research question and intended use of the matrix.
- Population, inclusion rules and excluded examples.
- Persistent object and source identifiers.
- Definitions, controlled terms and uncertainty codes.
- Observation and measurement methods.
- Source list, access dates and likely source dependencies.
- Treatment of missing values and unassessable examples.
- Any weighting or scoring rules and their limitations.
- Contradictory evidence and alternative explanations.
- Sample size, sample bias and known coverage gaps.
- Confidence attached to individual claims, not only the final conclusion.
- Revision date, version history and reasons for changed classifications.
Know when a table is the wrong tool
Use prose or responsive labelled cards when each difference requires substantial qualification, the contexts are not equivalent, the comparison concerns only a few complex cases, chronology is essential or every cell would need a footnote. A prose-heavy horizontal table encourages superficial scanning and hides the very distinctions the research needs to preserve.
A matrix is strongest when the data are compact, consistently structured and genuinely comparable. The subject determines the format; the existence of spreadsheet software does not.
Domain boundaries
Research informs; authentication concludes
Authentication boundary
A matrix can organise consistency and contradiction against verified references. It does not create authenticity from superficial resemblance, and public tables should not disclose every anti-counterfeit diagnostic where that would assist forgery.
Do not collapse separate axes
Grading and condition boundary
Condition comparison should keep structural integrity, surface wear, colour retention, completeness, restoration, active deterioration and presentation separate. A single global grade cannot explain every collector risk.
Price is not identity evidence
Valuation boundary
Market matrices belong primarily to valuation practice. Keep asking prices separate from realised sales and control for issue, condition, completeness, provenance, venue, premium, tax and date.
Research actions can alter the object
Preservation and restoration boundary
Opening packaging, removing labels, cleaning, sampling or disassembly may destroy evidence. Use an intervention decision matrix and escalate before any action that is irreversible, material-dependent or potentially damaging.
The data needs provenance too
Provenance boundary
Object provenance and data provenance are related but distinct. The matrix should record who observed what, from which source, on what date and with what confidence, so later researchers can audit the chain behind each claim.
Specialist thresholds
Escalate when the reference population is uncertain
One accepted example may itself contain replacements or undocumented variation. Seek specialist review where the matrix is being used to establish a new edition, attribute a maker or support a high-value authenticity claim.
Escalate before invasive or irreversible examination
Sampling, opening sealed packaging, removing adhesives, exposing hidden surfaces or testing materials may alter value, condition or evidence. A conservator or appropriate material specialist should define the least damaging route.
Escalate when measurement or imaging exceeds routine observation
Scientific imaging, spectroscopy, substrate analysis, ink comparison and precision metrology require suitable equipment, reference standards and interpretation. A collector table can frame the question but should not overstate an amateur result.
Escalate when the conclusion will travel
A matrix supporting a published correction, expert submission, insurance dispute, catalogue attribution or major sale should be reproducible, source-stable and independently reviewed.
What a defensible conclusion sounds like
A matrix can support statements such as: ‘the examined examples form two consistent physical groups’; ‘one attribute appears diagnostic within the current sample’; ‘the supposed variant is more plausibly condition-related’; ‘several records probably describe the same object’; or ‘the current evidence does not establish a chronological sequence.’
It cannot independently prove exact production quantities, intentionality behind a change, authenticity from superficial resemblance, rarity from low observed frequency, originality of every component or a manufacturing chronology unsupported by dated evidence. The conclusion should therefore state the sample, exceptions, missing evidence and plausible alternatives alongside the finding.
Key takeaways
- Build the table to answer one defined research question, not to display everything known about the object.
- Keep the object, source, observation, interpretation, confidence and evidential significance distinguishable.
- Compare like with like, and record contextual differences when perfect equivalence is impossible.
- Use controlled terminology and explicit missing-data codes; never allow a blank to become an accidental ‘no’.
- Treat clusters, scores, frequencies and chronological patterns as findings to test, not automatic verdicts.
- Preserve the provenance and revision history of the matrix so later readers can audit and update it.
- Use prose or labelled cards when compact tabular structure would flatten necessary qualification.
Continue learning
Absence of Evidence
Return to judging missing records, missing examples and silence in the evidence.
Back to Comparative Analysis
Return to the Comparative Analysis sub-domain and its full topic list.
Research Confidence Through Comparison
Continue to turning comparative evidence into proportionate research confidence.
Related topics
Selecting Comparison Examples
Choose reference examples that are relevant, independent and suitable for the judgement being tested.
Like-for-Like Comparison
Understand why comparability depends on the exact question rather than superficial similarity.
Variant Clusters
Examine how recurring combinations can be held as provisional clusters before becoming named variants.
Photographs and Descriptions Consistency
Connect research records with accurate sale descriptions and consistent photographic disclosure.