Hypotheses and Testing

A hypothesis is a provisional, testable explanation for an uncertainty about a collectible. It may concern identity, authenticity, date, attribution, production sequence, variant status, completeness, provenance, rarity or market behaviour. It is more disciplined than a guess because it predicts what evidence should exist if it is correct, identifies what would weaken it and remains open to revision.

Good hypothesis testing does not ask how to prove the most attractive story. It asks which explanation best accounts for all available evidence, including awkward and contradictory findings, what would change the conclusion and whether another researcher could reconstruct the reasoning. Collectibles rarely provide mathematical proof; the practical goal is a proportionate best current explanation.

Collector scenario

The theory that becomes too neat

Three boxed games share blue endpapers and the same rear label. A tempting theory forms: the blue endpaper copies must be an early issue made before the label was redesigned. That is a useful hypothesis because it creates a research direction.

Yet the same pattern could mark one printer, one distributor, a later batch, a repair source or several copies that passed through the same dealer. The collector should not discard the first idea simply because alternatives exist. The task is to compare what each explanation predicts and discover which one survives the full evidence, not merely the examples that made the story attractive.

Orientation

From research question to testable explanation

A research question identifies what is unknown. A hypothesis proposes an answer that can be tested. The wording determines whether the research can genuinely succeed or fail.

Research question

Was this boxed game issued before or after the documented 1981 revision?

The question defines the uncertainty without deciding the answer.

Testable hypothesis

This copy represents an earlier production state because its contents list, logo treatment and printer marks predate the documented revision.

The claim predicts several observable relationships and can be damaged by a securely dated counterexample.

Weak formulation

I think this is rare.

Rare has not been defined. The relevant population, observation period and evidence source are unknown, so the statement cannot be tested consistently. A workable claim might instead define rarity as the proportion of independently identified examples recorded within named archives during a stated period.

Recognition

The main forms of collectible hypothesis

Different claims require different evidence. Naming the form of hypothesis helps prevent a test for one question being mistaken for an answer to another.

Identity

What is the object?

Tests edition, issue, printing, production state, market or intended purpose through codes, dimensions, typography, packaging, contents and contemporary records.

This is the 1977 UK retail issue rather than the later export version.

Authenticity

Was it made as claimed?

Tests whether the object, signature or component was produced in the claimed period and manner using provenance, materials, manufacture, ageing and documented comparison examples.

The card is an original production example rather than a modern reproduction.

Attribution

Who made or altered it?

Tests maker, designer, printer, sculptor, signer or assembler through marks, construction, records, chronology and comparison. Style alone is rarely decisive.

This miniature was sculpted by Designer X rather than Designer Y.

Dating

When was it made or changed?

Uses date codes, addresses, prices, catalogue appearances, materials, inscriptions and ownership records. The result is often a range rather than a single year.

This copy was produced after the new logo appeared but before the address changed.

Sequence

Which production state came first?

Compares several independent features across multiple examples. No sequence should rest on one logo, label or serial number alone.

The yellow-label issue preceded the white-label issue.

Variant

Is the difference intentional?

Distinguishes a genuine issued variant from normal tolerance, fading, damage, repair, mixed components, later alteration or counterfeit modification.

The blue-backed card is an intentional factory variant.

Provenance

Where has it been?

Tests ownership, custody, movement or use through catalogues, invoices, labels, correspondence, photographs and estate or institutional records.

This manuscript belonged to Collector A before the 1954 sale.

Completeness

What was originally supplied?

Compares the present object with contemporary contents lists, sealed examples, photographs, trays, inventories and replacement-part records.

The set is factory complete, with no later substitutions.

Market

Does a measurable pattern exist?

Tests price, demand, scarcity or buyer behaviour using defined populations, controlled comparisons, achieved prices and duplicate-aware datasets.

Intact inserts command a premium after controlling for condition and venue.

Causal

Why did the pattern occur?

Requires evidence connecting cause and effect, not merely two events occurring together. Internal records and chronology are especially valuable.

The packaging changed because retailers reported display damage.

Domain boundary

Authenticity, originality, completeness and condition are separate questions

An object may be authentic but contain replacement parts; a genuine book may have been rebound; an authentic poster may have restored margins. Test the principal object's authenticity separately from component originality, period compatibility, completeness, repair and present condition.

Detailed conclusions about treatment belong primarily to Restoration; physical stability and safe handling belong to Preservation; formal grade belongs to Grading; ownership history belongs to Provenance. Hypothesis testing connects those evidence streams but does not replace their specialist methods.

Judgement

Test the preferred idea against competing explanations

A hypothesis becomes stronger by surviving alternatives, not by accumulating only agreeable examples.

Suppose an object carries a logo not found on familiar copies. The official-variant explanation is only one possibility. A later chronological revision, regional production, replacement label or full reproduction might explain the same feature. Each should be written down before the evidence is evaluated.

Official variant

The manufacturer intentionally issued the different logo.

Later revision

The logo marks a later production state rather than a simultaneous variant.

Regional issue

A printer, licensee or distributor used locally modified artwork.

Later alteration

The label or marked component was replaced after manufacture.

Reproduction

The entire object or component is modern.

Compact evidence matrix

ObservationSupportsWeakens
Period-correct paperVariant, revision, regional issueModern reproduction
Different adhesiveAlteration or reproductionOriginal untouched label
Regional distributor markRegional issueNo direct effect on revision
Appears in a dated cataloguePeriod issue or revisionLater alteration or modern reproduction

Do not count marks as if every item has equal weight. A securely dated catalogue image may outweigh several seller descriptions. Five websites may still be one source if all copied the same auction entry. Evidence quantity is not evidence independence.

Evidence

Weight evidence by proximity, reliability and independence

No single hierarchy fits every collectible, but collectors should distinguish strong evidence from repeated assertion and understand the limits of each source type.

Direct contemporaneous evidence

Evidence value

Production records, dated invoices, contracts, correspondence, proofs, catalogues, photographs and securely dated receipts are closest to the event being studied.

Collector risk

Contemporary does not mean infallible. Records may be incomplete, promotional, mistaken or applicable only to one market.

Object-based physical evidence

Evidence value

Construction, materials, printing, tool marks, mould features, repairs, component compatibility and laboratory results can test what the object itself permits.

Collector risk

A period-compatible material proves compatibility, not the date or identity of the finished object.

Securely documented comparison objects

Evidence value

Museum examples, reference collections, excavated objects, sealed copies and long-provenanced items establish reliable ranges of genuine variation.

Collector risk

One assumed genuine comparison object can turn an unsupported belief into circular authentication.

Later specialist scholarship

Evidence value

Reference books, specialist databases, academic articles and collector studies can synthesise dispersed evidence and define useful comparison criteria.

Collector risk

Authority depends on citations, transparency, correction practice and whether the underlying sources were actually inspected.

Market and community evidence

Evidence value

Auction descriptions, dealer records, collector testimony, forums, censuses and listings reveal leads, patterns and surviving examples.

Collector risk

Repeated listings may describe the same object or copy the same unsupported claim. Visibility is not production quantity.

Unsupported tradition

Evidence value

Inherited labels, seller folklore and statements that 'everyone knows' may supply a research lead.

Collector risk

Tradition can begin a hypothesis, but it cannot finish the test without independent evidence.

Comparison

Build a reference group, not a single comparison

Comparison is powerful only when the objects and records being compared are appropriate to the question.

Like for like

Match title, edition, condition, completeness, sale period, market and region for price research; match material, dimensions, typography, print pattern, binding and construction for authentication.

Reference groups

Include authenticated positives, known reproductions, later issues, suspected examples and borderline or transitional objects. Genuine variation needs a range, not one ideal specimen.

Blind review

Where feasible, remove labels such as genuine, rare or suspect before asking a specialist to compare examples. This reduces expectation effects.

A single anomaly is not yet a variant

An unusual feature may be a new variant, production error, repair, replacement, counterfeit, post-production modification, degradation, photography problem or data-entry mistake.

  1. Verify the observation and rule out imaging or measurement error.
  2. Inspect the object physically and compare normal production variation.
  3. Search for independently sourced examples with the same feature cluster.
  4. Look for documentary support and classify provisionally.
  5. Publish the criteria and uncertainty, then revise when new evidence appears.

Testing

Choose tests that answer the question without damaging the evidence

Scientific and physical examination can be decisive, but a machine result is never a substitute for a well-framed question.

Non-destructive examination

Depending on the material and specialist access, relevant techniques may include microscopy, ultraviolet or infrared examination, X-ray imaging, X-ray fluorescence, Raman spectroscopy, reflectance methods, digital image analysis, surface topography or computed tomography.

Begin with the least invasive method capable of answering the stated question.

Micro-destructive or destructive testing

Fibre, pigment or ink sampling, radiocarbon dating, dendrochronology, thermoluminescence, metallography and chemical extraction may remove or alter original material.

The possible knowledge gain must justify the effect on preservation, market value, future research and owner rights.

Compatibility is not identity

“The paper is compatible with 1950s manufacture” does not mean “the object was definitely made in the 1950s.”

Old stock can be used later, later materials can be introduced during repair and a genuine component can be assembled into a false whole. Scientific results should be integrated with provenance, manufacture, chronology and comparison evidence.

Before any sampling

  1. State the hypothesis and exact question.
  2. Confirm non-destructive methods are insufficient.
  3. Obtain owner consent and specialist advice.
  4. Document the object before intervention.
  5. Record sample location, size and method.
  6. Preserve residual samples where possible.
  7. Keep the complete report and calibration details.
  8. Disclose the intervention in future records.

Method

A practical hypothesis-testing workflow

The process should remain visible from first observation to revised conclusion.

01

Define the uncertainty

Write the research problem in one neutral sentence. Avoid building value words such as rare, prototype or first issue into the question.

02

Describe before interpreting

Record dimensions, materials, marks, text, construction, components, condition, provenance and photographs before naming what those observations may mean.

03

State the principal hypothesis

Write one provisional explanation that is specific enough to expose itself to evidence.

04

List competing explanations

Include ordinary and commercially inconvenient alternatives: fading, repair, replacement, printer change, mixed components, later alteration or reproduction.

05

Operationalise the terms

Define words such as variant, rare, original, complete, early and population in observable terms. Undefined language cannot be tested consistently.

06

Write predictions

State what should be found if each explanation is correct, including at least one observation that would seriously weaken the preferred hypothesis.

07

Design proportionate tests

Choose documentary, comparative, physical, statistical or scientific methods that answer the question without exceeding the object's value, significance or tolerance for intervention.

08

Set decision criteria

Decide in advance what would count as support, contradiction or an unresolved result. This reduces moving the goalposts after evidence appears.

09

Gather and preserve evidence

Keep original images, measurements, source copies, search logs, laboratory reports, correspondence, exclusions and failed tests.

10

Compare all explanations

Use an evidence matrix so that one dramatic item does not hide several weaker contradictions or dependencies.

11

State the result proportionately

Use language such as contradicted, unsupported, plausible, provisionally supported, strongly supported, best current explanation or unresolved.

12

Seek review and publish the trail

Ask another informed researcher to inspect the evidence, comparison group, omitted alternatives and reasoning. Preserve enough detail for reconstruction and later revision.

Interpretation

Use negative evidence carefully

The absence of an expected record can matter, but only when the record was likely to exist, survive and be found by the documented search.

Weak argument

I cannot find it in a catalogue, so it was never issued.

The search scope, catalogue coverage and survival of records are unknown.

Stronger argument

It is absent from the surviving annual catalogues, price lists and distributor sheets for the period, although comparable products were consistently recorded.

The collector explains why the missing record was expected and what was searched.

Negative evidence becomes stronger when record-keeping was systematic, archival survival is good, comparable items are present and the hypothesis specifically predicts that the missing evidence should exist.

Advanced analysis

Statistics for populations; explicit updating for individual objects

Formal statistical testing and object-level reasoning solve different problems.

Population and market hypotheses

Price premiums, sale frequency, regional distribution, grade patterns and variant proportions require a defined population, sample, variables, coding rules and treatment of missing data, outliers and duplicate objects.

Statistical significance is not automatically practical significance. Asking prices are not achieved prices, visible auctions do not represent every private transaction and the same object may appear repeatedly.

Individual objects and Bayesian reasoning

Unique collectibles rarely fit a simple null-hypothesis test. Begin with the plausibility of each explanation and update relative confidence as independent evidence arrives.

Period-compatible materials, a dated photograph and a former employee's identification may increase confidence; an anachronistic adhesive or digital font may reduce it sharply. The point is to make the update visible, not to invent a precise percentage.

False positives and false negatives have unequal consequences

A false positive may create a fictitious variant, inflate price and contaminate catalogues. A false negative may hide a genuine short-run issue, lose historical information or cause undervaluation.

The acceptable threshold should reflect the consequence. In authentication, falsely declaring a counterfeit genuine may be more damaging than cautiously leaving a genuine object unresolved.

Quality control

Make the test repeatable and open to replication

A conclusion gains credibility when another researcher can understand and, where possible, repeat the method.

Repeatability

Could the same researcher repeat the examination using the recorded equipment, lighting, measurements, image settings, comparison points and coding rules?

Replicability

Could another researcher independently apply the method to the retained evidence or a comparable sample and obtain a compatible result?

“The colour looked wrong” is not reproducible. A stronger record identifies the illuminant, calibrated target, camera settings, measured values, comparison sample and permitted range, while still acknowledging fading, storage and production variation.

Failure modes

Common ways collectors accidentally protect a preferred conclusion

Commercially loaded starting point

The test is designed to justify rare, prototype or first issue. Rewrite the question neutrally before considering value.

Circular authentication

Objects are called genuine because they resemble examples authenticated by resemblance to them. Anchor comparison to independently documented material.

Seller testimony as proof

A confident story is treated as a documented chain. Separate the claim, its source and its corroboration.

One-feature testing

A correct font is allowed to authenticate paper, printing, binding and provenance. Test multiple independent features.

Ignoring normal variation

Hand finishing, supplier changes and print tolerances are mistaken for forgery or variant status. Establish the genuine range first.

Laboratory result as verdict

A compatible material is treated as proof of identity. Integrate science with history, provenance and construction.

Dependent sources counted repeatedly

Auction and web descriptions copy one another. Trace claims to their earliest identifiable source.

Data dredging

Categories and exclusions change until a significant result appears. Define variables and criteria before examining outcomes.

Survivorship and market bias

Only complete, attractive, sold or publicly visible examples are studied. Include ordinary, damaged, unsold and poorly documented material where relevant.

Premature certainty

A plausible explanation is published as established fact. Use confidence language and state what remains unresolved.

Worked example

A supposed undocumented first printing

A strong test can reject the collector's original label while still producing a significant discovery.

Research question

Is this rulebook an undocumented first printing?

Principal hypothesis

The rulebook predates the recognised first retail printing.

Competing explanations

  • Later regional printing
  • Printer's or publisher's proof
  • Modern facsimile
  • Normal copy with replacement covers

Predictions and tests

  • Collate text, pagination, signatures and imposition.
  • Compare paper, ink, staples, binding and printer marks.
  • Search registrations, publisher files and period catalogues.
  • Trace provenance and compare known proofs and regional issues.
  • Look for anachronistic digital artefacts or later materials.

Proportionate conclusion

The copy is provisionally identified as a pre-publication proof rather than a first printing. Its text predates the retail edition, but the untrimmed format, hand assembly and lack of distribution evidence fit proof production better.

Action hierarchy

Match the depth of testing to the consequence of the claim

Not every personal catalogue note requires laboratory analysis. The effort should rise with financial, historical, reputational and preservation risk.

Low consequence

Labelling a low-value personal item

  • Basic catalogue and reference comparison
  • Clear observation notes and photographs
  • Provisional wording that preserves uncertainty

Moderate consequence

Publishing a new variant or paying a substantial premium

  • Multiple independently sourced examples
  • Provenance and source-dependency review
  • Comparison matrix and specialist opinion

Do not let a value premium become evidence for the classification that created it.

High consequence

Declaring a major rarity genuine or accusing an object of being counterfeit

  • Direct physical examination
  • Documented chain of evidence
  • Qualified specialist review
  • Scientific analysis where relevant
  • Explicit legal, ethical and uncertainty review

High-consequence claims require unusually careful language, disclosure and preservation of the complete record.

Documentation

Preserve the hypothesis record, not only the conclusion

Hypothesis record

  • Stable hypothesis ID
  • Wording and date proposed
  • Research question
  • Researcher
  • Competing hypotheses

Object record

  • Stable object identifier
  • Ownership or repository
  • Dimensions and materials
  • Condition and components
  • Marks, provenance and photographs

Test plan

  • Predictions
  • Methods
  • Comparison sample
  • Acceptance and rejection criteria
  • Risks, permissions and limitations

Results

  • Raw observations and measurements
  • Original images and reports
  • Uncertainty and anomalies
  • Exclusions and reasons
  • Failed or inconclusive tests

Evaluation

  • Evidence for and against
  • Independence and reliability
  • Alternative explanations
  • Confidence level
  • Reviewer comments

Revision history

  • Original conclusion
  • New evidence
  • Changed interpretation
  • Date and author of revision
  • Effect on catalogue wording

Reporting

Use controlled confidence language

Confidence labels should summarise the reasoning, not replace it.

Confirmed

Direct, reliable and independently verifiable evidence leaves no realistic competing explanation.

Strongly supported

Several independent evidence streams favour the hypothesis and remaining alternatives are weak.

Provisionally supported

Current evidence favours the hypothesis, but important gaps or tests remain.

Plausible

The explanation fits known evidence but lacks decisive testing.

Possible

The explanation cannot be ruled out, but evidence is limited.

Unresolved

Available evidence does not adequately distinguish the competing explanations.

Unlikely

The balance of available evidence weighs against the hypothesis.

Contradicted

Reliable evidence is incompatible with the hypothesis as stated.

Specialist threshold

When the collector should stop and escalate

Seek a qualified specialist when the object cannot be examined safely, the test could remove original material, the claim affects a high-value purchase or public accusation, specialised equipment or calibration is required, legal or cultural due diligence is involved, or the collector lacks authenticated comparison material.

Some hypotheses remain untestable because archives are lost, provenance is too incomplete, comparison objects are unavailable or the necessary intervention would be disproportionate. Record such claims as tradition, attribution, possibility, unverified claim or research lead. Repetition does not convert them into fact.

Key takeaways

  • Describe the object before interpreting it.
  • Turn claims into predictions that can support, weaken or contradict them.
  • Test competing explanations, including ordinary and commercially inconvenient ones.
  • Trace repeated claims to their underlying source and assess evidence independence.
  • Use well-documented comparison groups and establish normal variation first.
  • Separate authenticity, originality, completeness, condition and value.
  • Treat scientific results as one evidence stream, not an automatic verdict.
  • Record negative, failed and inconclusive results as part of the evidence trail.
  • Use confidence language proportionate to the evidence and consequences.
  • Prefer the best current explanation to unsupported certainty, and revise it when better evidence appears.

Final collector principle

A hypothesis is not a label attached to an object. It is a structured explanation exposed to evidence, uncertainty, alternative accounts and future correction.

Continue learning

Related topics