The parent method
What does this Observatory compare, and what may it conclude?
The Observatory asks one set of questions across seven domains of work. This page sets out how a claim is evaluated: what kind of record counts as evidence, what a conclusion drawn from it may say, and where it has to stop. It compares questions rather than results.
Two readings, never one
How do you tell what a record is from what it establishes?
How far a domain's calibration has got, what kind of record the evidence is, and what a claim is licensed to say are three separate readings. None substitutes for another and no operation combines them.
Evidence basis
What kind of thing a record is. A property of the record, and the reading most often collapsed: a demonstration and an agent acting on real work both read as "it works".
-
Announced
A claim about what an agent will or can do.
Does not establishThat it has been built, shown, or run anywhere.
-
Demonstrated
Shown working, under conditions the source controlled.
Does not establishThat anyone has deployed it, or that it runs on real work.
-
Deployed
Present in a real environment.
Does not establishThat it acted, or that any effect of its acting was recorded.
-
Observed in operation
Seen acting on real work, with the effect recorded.
Does not establishHow often it happens, how well it performs, or what it changed.
Claim standing
What the evidence licenses you to say. A property of the claim.
-
Observation
Direct evidence establishes it.
-
Pattern
Repeated observations support it.
-
Interpretation or hypothesis
Evidence supports a reasoned explanation.
-
Conditional
The outcome depends on stated conditions.
-
Not established
The present evidence cannot support the conclusion.
Conditional describes the form of a claim, not the strength of its evidence. A well-evidenced conditional claim and a thin one are both conditional.
Not established is a finding, not a gap in the interface. It says the evidence we checked cannot support the conclusion, and it is rendered with the reason attached.
What the method refuses to infer
Why does each level of evidence stop where it does?
Four inferences look reasonable and are not supported. Each one is a level of the framework read as though it were the level above it.
-
An announcement does not establish a deployment
A claim about what an agent will do is a claim about intent. Nothing about it establishes that the thing was built, shown, or run anywhere.
-
Workflow coordination does not establish firm change
Is the organisation changing around that coordination? Roles, decision rights, systems or governance change. Until that is shown, it remains workflow coordination.
-
A firm changing does not establish ecosystem change
Are counterparties, platforms or institutions adapting? Relationships, rules or strategic positions change. Until that is shown, it remains change inside one firm.
-
An ecosystem adapting does not establish a reshuffle
Have control, strategic position, bargaining power or value moved? Their distribution has demonstrably changed. Until that is shown, it remains ecosystem adaptation.
Reshuffle is the cross-level conclusion, not a fourth operational level. Agents can become more capable, more active and more deeply embedded without producing one.
Does the agent merely complete work, or determine how interdependent work is organised? Task automation: Performs defined work inside existing rules, dependencies and decision rights. Agentic coordination: Organises interdependent work across changing conditions, steps, systems or actors.
How the domains are compared
How do you compare seven domains without turning a hypothesis into a finding?
- A research design is not a finding. A domain's fields describe what will be examined and why that system is structurally distinctive. None of them reports an observation, and the registry refuses to build if one starts to.
- Claim standing is stated in words, never by colour alone. Every claim carries a mark and a written label, so it survives greyscale and a screen reader.
- The comparison compares questions, not answers. The cross-domain matrix says what each question would require the research to observe in that domain. Until two domains hold evidence there is nothing to compare, so the instrument publishes the comparison it intends to make rather than its result.
- A domain's own vocabulary stays in that domain. Procurement measures how far into the work an effect reaches, and how far a constructed configuration sits from its records. Neither is a claim standing, and neither is shown cross-domain.
The domain vocabularies that are deliberately not mapped
-
Participation depth
How far into the work a documented effect reaches.
Advises · Judges or completes · Executes
A depth reading says what an agent did in the workflow. It says nothing about what kind of record the evidence is, or what a claim may assert.
-
Distance from the evidence
How much of a constructed configuration the records already show.
anchored · emerging · conditional · frontier
It describes a construction's relationship to the corpus, not a claim's standing. A configuration nothing evidences can still fit a form of buying perfectly.
-
Enforcement depth
Whether a documented effect advises, executes, or makes an outcome binding.
Advises · Executes · Binds the outcome
An authority reading, not an evidence reading. It is the subject of the Control question rather than an input to it.
The method, applied
Where is this method applied to real evidence?
Agentic Procurement applies these questions to a corpus. Its own evidence surface owns what this page does not: which sources were admitted, how each cell was coded, what every count denominates, and what the corpus cannot settle.
Its record uses its own words for claim standing, and they map to the vocabulary above without being rewritten inside the frozen files:
- Observed Observation
- Pattern Pattern
- Hypothesis Interpretation or hypothesis
- Not established Not established