The OBDS research surface publishes the evidence and the boundary of the evidence: deterministic cross-language result hashes, a governed communications benchmark, and a written statement of what OBDS does not prove.
Can a system retrieve the right evidence and still decide wrong?
That is the question this research exists to answer. The short version is yes, and the two are different problems. Retrieval asks what looks relevant. Governance has to ask which truth is authoritative, which scope applies, whether it is valid now, whether required truth is missing, whether a conflict is relevant, and whether execution may proceed.
OBDS Governed Result Hash — deterministic interoperability evidence
The strongest thing OBDS can currently demonstrate: the same governed inputs, run through independent implementations, produce the same governed decision and the same result hash. And the mirror of it, which matters more in practice: the same governed inputs that do not support the request produce the same refusal. Published with 59 of 59 cross-language vectors byte-identical and a runnable verify.py.
OBDS Governed Communications Benchmark
Fourteen reproducible cases, each shipping its own manifest, build plan and expected governed decision. Every wrong claim in the benchmark is a downstream transformation of a correct statement: a revenue share becomes a product share, an eligibility figure becomes an alignment figure, a process comparison becomes a whole-product claim, a fiscal-year fact becomes a present-tense one. That is the failure mode retrieval cannot see, because the retrieved passage is correct and the claim built from it is not.
Case 14 is adversarial: it was written to defeat OBDS, and it does. It is published as a failure rather than removed, because an under-declared build target builds successfully with valid hashes and no governance layer can see it.
Known Limits of OBDS
A governance specification that cannot state its own boundary is a marketing claim. In one line: OBDS can govern declared truth, and it cannot prove undeclared truth does not exist. It cannot detect an approved value that is false in the world, and it cannot detect a build target that under-declares what it should have required.
What this is not
The research directory is not a second specification and nothing in it changes a normative artefact. It is not part of the release package. The benchmark cases are neutral and de-identified; they name no company and no sector. Results are stated against the reference implementation only and say nothing about how any particular model behaves in production.