Open Brand Definition
OBDS / 4.1.3 stable
English only
Research / SUPABRAND Companion to OBDS 4.1.3

Governed brand truth,
across two
submissions

Two separately authored submissions reached 39/40 on hidden-case curation and made the same 93/93 Task Facts decisions. Neither completed full governed enforcement: 0 of 20 end to end.

CorpusSUPABRAND, a synthetic and deliberately difficult brand archive.
Round 4Run against OBDS 4.1.2, September 2026.
VerdictMixed.
01 / The problem
In one sentence

The governed semantics aligned strongly. Complete execution did not.

Retrieval can find the right source and still use the wrong truth

A retrieval system ranks passages by relevance. A brand decision needs something else: which approved truth is authoritative here, whether it is still valid, whether it applies to this market and channel, and whether anything may be produced at all.

Retrieved passageRelevantMay it govern this task
A 2024 range claim, approved at the timeYesNo. Its approval period has ended.
A current range claim approved for the United Kingdom onlyYesNo. Wrong market.
A current range claim approved for the EU, with a required footnoteYesYes, with the footnote.

An illustrative case in the style of SUPABRAND, not taken from the benchmark: a post for Germany, published today, about an e-bike.

Only one of the three may govern the post. If that one were missing, the correct outcome is to stop and ask, not to fall back on the closest alternative.

Access / retrieval  ->  possible information
OBDS                ->  governed applicability
Execution           ->  action

OBDS does not replace retrieval, and it does not judge whether a claim is true in the world. It records which approved truth applies to which task, and when execution must not proceed.

02 / The corpus

What SUPABRAND is

SUPABRAND Mobility SE is a fictional mobility manufacturer headquartered in Vienna. Everything in it is invented: no real company, certification, partnership or product performance is asserted.

Architecture
A masterbrand, four endorsed subbrands and a repair and parts programme.
Reach
Twelve example markets, thirteen working locales, editorial material in five languages.
Library
Agency and brand-operations documents from 2022 to 2026, with a handover snapshot dated 7 September 2026.
Sources
25 source documents, including a claims register and an evidence dossier.
Assets
40 assets with a register of their production and rights state.

The corpus behaves like a real brand archive. Documents overlap and some contradict each other. Approvals expire. Rules differ by market, channel, product and date. Some things the brand does not know. Retained artwork is not automatically licensed for a new placement.

Some of the difficulty is deliberate. The evaluator holds 20 hidden cases that participants never see, covering scope, validity periods, conflicting authorities, claims that apply only in certain contexts, and truth that must stay unknown. They are described by category only, because SUPABRAND remains in use as a test.

03 / Setup

How Round 4 ran

Both candidates received the same public corpus at one pinned commit and OBDS 4.1.2 at its release commit. The evaluator verified every corpus file byte for byte, hashed each candidate's complete file tree on receipt, and rechecked at the end. Neither had changed during the evaluation.

Candidate ACandidate B
Built by, as labelled by the operatorClaudeAstra
LanguagePythonJavaScript (Node)
Task Facts evaluatorThe published OBDS Python reference evaluator, called from its own runnerIts own independently written evaluator
DeliverableLibrary code and a draft package, with no final report or sealSealed package with a final report

What was separate, and what was not

Separate: two systems, two languages, two submissions, and no reference found from one candidate's output to the other's.

  • Candidate A had an internal curation guide from earlier rounds and cited it in its code, so the guide-free premise of the round does not hold for A.
  • Candidate B disclosed that it had read the published Task Facts conformance results beforehand.
  • Candidate A's Task Facts evaluator is the published reference implementation, not one that A wrote.
  • Without complete access logs, isolation between the two cannot be certified.
  • The evaluation was run by the project, separately from both candidates. It is not an external audit.

Each system ran once. There were no repeated runs, no equal budgets, and model versions were not recorded. One qualified run per system does not establish general model capability, and this page does not rank the two systems.

How the scoring was protected

  • The 20 hidden cases and their expected outcomes were recorded before any candidate code ran.
  • Curation scores were frozen, and the frozen files hashed, before the two candidates were compared. No material from earlier rounds was read.
  • 54 independent applicability probes were written with expected results fixed before execution, then run against each candidate's unchanged code in temporary copies.
  • Both Task Facts paths ran the official OBDS suite, and the same evaluator-owned inputs were replayed through both.
  • Each package was copied to a new location and rerun to test portability.
  • Every claim a candidate made about itself was audited against evidence.

The role of Task Facts

Task Facts 1.0 is an optional OBDS capability, added in OBDS 4.1.0. It defines how the facts of one specific task, for example the market, the date and the set of claims or assets involved, are represented, hashed and evaluated against declared conditions. It covers sets, item-to-item associations and dates, and keeps missing, empty, unknown and contradictory values apart. A fact being present does not make it true: verification and provenance stay separate. Round 4 was the first SUPABRAND round in which this contract existed.

04 / Results

Round 4: mixed

Three numbers carry the result, and each is narrower than it looks.

39/40 Curation, both candidates Hidden-case curation points. 19 of 20 cases fully correct, one partial in each submission. It measures whether the governed truth was captured, not whether it was executed.
93/93 Bounded replay agreement 58 evaluator-owned inputs replayed through both Task Facts paths: every six-field decision record identical, invalid and unresolved records included. Bounded replay agreement, not universal interoperability.
0/20 Complete governed execution None of the 20 hidden cases completed the required end-to-end governed pipeline for either candidate. This is not 20 incorrect decisions.
         Same public SUPABRAND corpus, one pinned commit
                              |
        +---------------------+---------------------+
        |                                           |
  Candidate A: Claude / Python              Candidate B: Astra / Node
  also had the internal curation guide      public material only
        |                                           |
   39/40 curation                            39/40 curation
   (20 hidden cases)                         (20 hidden cases)
         \                                        /
          +-------------------+--------------------+
                              |
            Task Facts replay, 58 evaluator inputs
               93/93 identical decision records
       A's path: published Python reference evaluator
       B's path: independently written Node evaluator
                              |
                 Complete governed execution
                     0 of 20 demonstrated

One corpus, two submissions, one shared decision layer, no completed execution. The two paths were not operationally identical: only Candidate B wrote its own evaluator.

MeasureCandidate ACandidate B
Hidden-case curation19 of 20 fully correct, 39/40 points19 of 20 fully correct, 39/40 points
Complete governed execution of a hidden case0 of 20 demonstrated0 of 20 demonstrated
Independent applicability probes28/28 passed26/26 passed
Official Task Facts suite66/66 fixtures, 36/36 regressions66/66 fixtures, 36/36 regressions
Shared Task Facts replay93/93 records identical, one shared result across both paths
PortabilityPartially portablePortable

Some hidden cases had a bounded executable counterpart in each submission. None of them amounted to a complete governed run, and the delivered coverage differs between the two, so the two are not comparable as a score.

The evaluator called Candidate B the stronger and more reviewable package and the preferred basis for a real brand package. It remains a draft that needs further curation, occurrence coverage, trusted verification and governed integration. Neither candidate is ready for production.

05 / What worked

What worked

Curation reached a high level in both submissions

Working from the same sources with OBDS 4.1.2, both captured 39 of 40 points on cases they had never seen. Candidate A also had the internal curation guide, so only Candidate B demonstrates this from public material alone.

Two Task Facts implementations produced identical decisions

93 of 93 six-field records matched across an independently written JavaScript evaluator and the published Python reference evaluator, invalid and unresolved cases included. Both paths also pass the official suite: 66 fixture decisions and 36 regression records.

The applicability layer behaved as specified on the tested probes

All 54 independent probes passed. They covered several claims or assets in one task, partner and campaign combinations, market against locale, missing, empty, unknown and contradictory inputs, misuse of a single value where a set is required, and ambiguous times. Neither submission needed a private predicate language for the conditions tested.

One package was reproducible elsewhere

Candidate B was copied to a fresh location and environment and passed all six of its steps with its documented, pinned dependencies, without the original machine or paths. The first attempt failed on a missing Python package, which the documented lockfile resolved.

The approval boundary held

Candidate B's package stayed a draft, and its production build refused that draft and produced no compiled context for a model. This demonstrates the approval gate, nothing beyond it.

Bounded claims were supported

Candidate B stated draft status, partial coverage, no brand approval and no production readiness. The evaluator's audit found those statements supported, and Candidate B kept the brand's genuine unknowns explicit as unknown rather than resolving them.

06 / What failed

What failed

0 / 20

Neither candidate demonstrated a complete governed execution for any of the 20 hidden cases. Neither delivered a path from curated sources, through a Task Facts decision, into a governed production result. This is the central negative result of Round 4. It does not mean that 20 decisions were wrong.

Where the execution boundary sits

  • Neither submission delivered a trusted host that carries every relevant occurrence, source scope and validity from the curated sources through Task Facts decisions into execution. The evaluator names this as the most important remaining gap.
  • OBDS 4.1 has no adopted binding from sources through Task Facts to Compiled Runtime, and no conformance test for such a binding.
  • Item-to-item associations were only partly handled, partly through implementer omissions and partly because general joins, relations across several targets and every-occurrence expansion are not in the standard.
  • Neither candidate authenticates where a task fact comes from, so a production gate would rest on unverified input.

Some executed decisions were wrong

Candidate A's private brand adapter was tested on 19 decisions. It matched the evaluator on 11 and differed on 8: two false allows, where content would have been let through that should have been blocked; two false blocks; three that turned an unresolved unknown or escalation into a plain block; and one that lost which item a problem belonged to. The evaluator attributes these to curation and implementation errors, not to missing OBDS guidance.

Associations were only partly handled

Some rules depend on which item relates to which: this asset with that product, this partner with that placement, this campaign beside that hero image. Of four such challenges, Candidate A left three unhandled and Candidate B two. OBDS supports occurrence identifiers and exact association pairs, so part of this was an implementer omission. The rest is recorded as a standard gap.

Candidate B's clean error record is narrow

Candidate B showed no false allows or false blocks, but only inside a narrow dependency layer. Many brand conditions were not implemented, and several obligations were not fully curated. A full-brand error rate for B cannot be measured from this round.

The intended experiment was not cleanly realised

Candidate A never reached a verifiable final state: no final report, no seal, and files its own code refers to are missing. It also used the internal curation guide the round was meant to exclude. Candidate B had seen published conformance results in advance. This is therefore not a comparison of two complete, fully separate, guide-free submissions.

Portability was partial for Candidate A

Its submission ran only when the delivered OBDS inputs beside it were copied as well, and it has no reproduction entry point of its own.

Provenance was not established

Neither candidate authenticates where a task fact comes from, and both created their own test contexts. A caller asserting that rights exist does not establish that they do.

07 / For OBDS

What this says about OBDS

Within this bounded experiment, three statements hold.

A shared contract for task-level facts can remove the need for a private condition language at the predicate level

Both submissions used the published Task Facts semantics for the conditions tested, and two implementations agreed exactly. The evaluator's answer to whether OBDS 4.1.2 removed that need is partially: at the predicate layer yes, for complete brand decisions no.

The public specification appears sufficient for faithful curation, and not yet for complete execution

Candidate B supports the first half from public material alone. Candidate A cannot corroborate it, having used the internal guide. The evaluator's answer on authoring sufficiency is also partially.

The next boundary is the host, not the vocabulary

What is missing is a trusted host that carries occurrence, scope and validity from sources through decisions into execution, together with an adopted binding for that path, a conformance test for it, standard occurrence expansion, interval operators and a protocol for authenticated verifiers.

Rounds 1 to 3 ran on OBDS 4.0.4 and led to the Task Facts work. Their scores used different methods and are not compared with Round 4, and no claim is made that changes to OBDS caused an improvement between rounds.

08 / Unproven

What remains unproven

This round does not demonstrate any of the following.

  • Complete automated brand enforcement, on SUPABRAND or any brand.
  • Production readiness of any submission, or of an OBDS-based production pipeline.
  • End-to-end portability of conditional brand logic, as opposed to portability at the Task Facts layer.
  • That the public OBDS material alone is sufficient for independent authoring, since only one candidate supports it.
  • Two independently authored evaluators from the two candidates, because one used the reference implementation.
  • A clean, isolated, guide-free experiment.
  • Any ranking of the systems involved.
  • That changes to OBDS caused score changes between rounds.
  • Authenticated provenance, rights verification, or the completeness of an asset inventory.
  • Extraction of facts from natural-language task requests, or complete publication review.
  • Full-brand rates of false allows and false blocks.
  • Behaviour on real brands, or beyond one synthetic corpus.
  • Usability for external implementers outside this project.
  • Review by an external, third-party evaluator.
09 / Reproduce

Reproduce and inspect

The corpus and the specification are public. The hidden cases are not, because publishing them would turn SUPABRAND into an answer key.

Corpus
The public SUPABRAND corpus, at the commit used in every round: ktdmax/supabrand.xyz @ 2d1651b.
Specification
OBDS 4.1.2, the release Round 4 ran against: OBDS-4.1.2.md, section 35 for Task Facts and section 26.10 for its conformance claims.
Task Facts suite
The published evaluators and their suite: reference/task-facts/1.0/README.md. Anyone can rerun the 66 fixture decisions and 36 regression records.
Not published
The hidden cases, their expected outcomes, the probe and replay inputs, and the candidates' full submissions, whose curated packages contain near-correct answers. The 93-record replay therefore cannot be rerun from public material; the official suite can.
10 / Open question

The question that stays open

Can independent systems reconstruct and apply the same governed brand truth from the same difficult source material?

At the level of decisions, this round says yes. Acting on that truth end to end is still untested.

Round 4 was the first SUPABRAND round with a published contract for task-level decisions, and the first in which two implementations could be compared decision by decision. They agreed on every shared decision, curated conflicting material to the same score, and needed no private condition language for the conditions tested. It also found a boundary: no hidden case completed a governed run, associations were only partly handled, and the step from correct decisions to trusted, governed execution was not delivered.