UK Regulatory Burden — Validation Rubric
Version 1.1 — authoritative (project-lead sign-off read of v2.6 complete)
v1.1 (2026-07-21) — six additions from pilot adjudication: (i) constituted-duties rule; (ii) specifications-vs-embedded-duties test; (iii) IB boundary + named-target discipline (schema gains source_target), incl. defence-provision routing (§4); (iv) compliant-actor test; (v) polarity conventions; (vi) capacity note. No previously-traced counts change.
v1.0.1 (2026-07-18) — exclusion-family taxonomy aligned to the ratified v2 schema (structural promoted to family; amendment_machinery to counted_at_source subclass). Recording-schema alignment only; no counting rules changed.
v1.0 — AUTHORITATIVE. The project lead’s full sign-off read of v2.6 is complete (read date 2026-07-17): definitions signed off (comment 11), the schema-multiplicity prerequisite agreed (comment 10), no structural findings. This is the rubric the labelling pilot runs against. Read-fixes folded (wording, not structure): capacity axis — ownership sub-rule (held outside a trade → personal; embedded in market activity → economic; undifferentiated owner-duties → both/either) and “consumption” reframed as use of goods/services already acquired (acquiring transaction economic, subsequent private use personal); compulsory-insurance example nuanced (both/either where undifferentiated, personal where the scheme splits the private side); §1 intro examples flagged as scope-illustrations, not tags; the frontier paragraph now states the frontier_hook / frontier_target_type tagging inline (comment 6); the passive-voice one-liner softened to the context provided with the candidate (route to review, never guess); §7 multiplicity note updated (representation resolved by the leaf-anchored burden-set schema — label-store build pending, scheduled before hand-labelling; production emission architecture on the stage-2 agenda); polarity note gains an inline surface-verb example; §5 flagged open-ended-by-design; schema gains unapplied_amendment; worked examples added (CCR 2013 reg 10 vs reg 16 split-vs-fold; CRA 2015 s.26 remedies-architecture zero); FSMA one-home proxy note added; aggregate-noun style line added. Nothing here changes a classifier rule. The human project lead owns and revises this document.
Purpose. This document defines how to classify provisions in UK in-force legislation for the regulatory burden measure. The PRODUCTION classifier is a trained non-LLM model (Legal-BERT, fine-tuned) that runs the full corpus. The LLMs (Claude and Gemini) use this rubric to propose training labels and to independently cross-check outputs (human review and validation of this training data follows this); disagreements between them are the signal that surfaces hard cases for human review. This rubric is the conceptual guide for that labelling and verification, and documents the definitions the training data embodies. The human project lead owns and revises it; no classifier changes the rules.
1. What we are measuring
The project counts regulatory burdens imposed on private actors by UK in-force legislation. It is a broad measure of the prescriptive control regulation imposes on private actors across society — not only on economic activity. Obligations on individuals in their personal capacity (for example, road-safety or animal-welfare duties) are in scope, alongside obligations on businesses and organisations. (These examples illustrate SCOPE — that personal-capacity duties are in scope — not fixed capacity tags: the economic/personal tag is decided case by case per the axis below.)
The core definition
Regulatory burden = an obligation or prohibition imposed on a private actor with legal force. This is the unit the project counts. Every burden carries a category (section 2) and a polarity — obligation or prohibition (section 2A). Wherever a rule below refers to a burden, it applies equally to both polarities; “obligation or prohibition” is not restated in each rule — it is inherited from this definition. Style: refer to the aggregate as “burdens”; “obligations and prohibitions” is the definitional expansion, used where the primitive is being defined.
Private actor = a business, individual, or third-sector organisation, acting in the course of its own commercial, personal, or organisational activities. This includes private companies, sole traders, individuals in personal capacity, and charities. It also includes a private firm delivering public services under contract: a contractor is a private actor even when the work it does is public in nature, because its role arises from a commercial contract, not from statute, and it holds no statutory powers. The burden falls on it as a private entity.
Public body = a body created by, or deriving its function from, statute; exercising statutory powers; and accountable to a public regulator — regardless of how it is funded. This covers ministers, government departments, regulators, courts, named officeholders, and statutory bodies such as the FSCS and the Financial Ombudsman Service (which are public for this purpose despite being industry/levy-funded, because their functions and powers are statutory). Obligations on public bodies are regulatory machinery, not regulatory burden, and are excluded by definition — regardless of category or polarity.
Distinguishing the two — the test is the SOURCE of the body’s function: statutory (public) versus contractual/commercial (private). It is NOT whether the activity looks “public”, nor how the body is funded. A private contractor doing public-sector work is private (its function is contractual); a levy-funded statutory body is public (its function is statutory). Genuinely mixed bodies (e.g. a charity exercising some quasi-public statutory function) route to review rather than being forced. The detailed function test is at section 5C.
Note — currency and commencement are delegated to the corpus layer. “With legal force” is enforced upstream by the corpus layer (na_inforce / CLML status filtering); labellers do not re-litigate whether a provision is in force or commenced — the corpus has already filtered to in-force text.
What counts as one burden — the unit
The unit of count is the distinct burden, not the sentence. A single sentence can impose several burdens, and a single burden can span several sentences; the count tracks burdens, not grammar.
- Enumerated burdens — where legislation enumerates distinct things the actor must do or must not do (a chapeau plus “(a)… (b)… (c)…”, each a separate requirement), count each leaf as its own unit.
- Conditions and factors are not separate burdens — where the enumerated items are conditions, factors, or considerations attached to a single duty (e.g. “must have regard to — (a) cost; (b) risk; (c) benefit”), count the single burden once, NOT the items. This is the most common over-counting trap: a list does not automatically mean multiple burdens.
- A burden spread across sentences — where one burden is stated across several sentences (the duty, then elaboration), count it once.
The discriminator — does the leaf, read alone, state something the actor must DO or must NOT DO (→ a separate unit) or something that QUALIFIES how, when, or whether a single burden applies (→ part of one unit)? Where a leaf genuinely cannot be cleanly called a burden or a condition, route it to review rather than forcing the call. (The precise burden/condition boundary is to be tested against real enumerated provisions before this is locked — see section 7.)
Refinement — specifications of a single act fold; independent operative constraints count each. Where a chapeau introduces a list, ask what the leaves ARE relative to the duty. Where they specify the components, contents, or particulars of ONE required act — the information a single record must contain, the elements a single notification must include, the particulars a single statement must set out — they fold into that one burden, however long the list. Where each leaf is an independent operative constraint the actor must separately satisfy — a distinct thing to do or not do, each violable on its own — count each. Worked contrast set: UK GDPR Art 30(1) (maintain a record “containing all of the following” (a)–(g)) → FOLD to one; ERA 1996 s.1 (give a written statement whose particulars are listed (a)–(n)) → FOLD to one; Environmental Permitting reg 24(3) (a notification “must — (a) be made on the form, (b) include specified information, (c) specify the date”) → FOLD to one; BUT Environmental Permitting Sch para 5(2) (the groundwater-activity conditions (a)…(g), each an independent limit — nothing added, temperature caps, siting, discharge distances) → COUNT EACH. The test is single-act-specification (fold) vs independently-violable-constraint (count).
Worked pair — a general duty and its particulars (a further instance of the fold/count-each test). Health and Safety at Work Act 1974 s.2: the s.2(1) general duty is elaborated by s.2(2)(a)–(e), which are NOUN-PHRASE particulars describing the EXTENT of that one duty (“the provision and maintenance of plant and systems of work…”, “arrangements for…”) — they specify what the single duty covers, so they FOLD into it. Factories Act 1961 s.1: the leaves are separate IMPERATIVE requirements (kept clean; and the associated cleaning obligations), each independently violable, so COUNT EACH. Test: imperative-leaf list → count each; noun-phrase particulars of one duty → fold.
Worked pair — split vs fold (from the supplement round). Consumer Contracts Regs 2013 reg 10 (pre-contract information, off-premises): the trader “(a) must give the consumer the information listed in Schedule 2 … and (b) if a right to cancel exists, must give the consumer a cancellation form” → COUNT EACH — (a) and (b) are two independently-violable duties (giving the information; giving the cancellation form). Contrast reg 16 (distance-contract confirmation): the trader “must give the consumer confirmation of the contract … [which] must include all the information referred to in Schedule 2” → FOLD to one — the Schedule-2 items specify the CONTENTS of the single confirmation, not separate acts. And a remedies-architecture zero: CRA 2015 s.26 (instalment deliveries — “the consumer is not bound to accept delivery by instalments unless agreed”, then default rules on defective instalments) imposes no standing operative duty on a private actor → ZERO burdens (a typed exclusion, not a count).
Specifications vs embedded duties — the breach-without-attempting test (v1.1). When a chapeau introduces a list, distinguish SPECIFICATIONS of one act (which fold) from EMBEDDED independent duties (which count each) with one test: could the actor breach this item while never attempting the parent act? If NO — the only way to violate the item is by doing the parent act badly — it specifies the parent and folds. If YES — the item is violable on its own, independently of the parent — it is an embedded duty and counts. Grammatical-shadow note: surface form (imperative vs noun-phrase) is a HINT, not the test — a noun-phrase item can still be independently violable and an imperative can still be mere specification; apply the breach test, not the grammar. Worked triple: a form with prescribed contents (“the notification must (a) be on the form, (b) include X, (c) state the date”) → 1 (you cannot breach (b) without making a notification); “maintain records AND submit a return” → 2 (you can maintain but fail to submit, or the reverse); “obtain an accredited assessment” where accreditation is a separable standing requirement → 2 / route to review. Section-scale pair: UK GDPR Art 13 (disclosure-notice contents fold → count the disclosure duties, not the items) vs Environmental Permitting Sch 2 para 7 burials (each locational/operational condition (a)–(n) is independently violable → count each). This sharpens the single-act-specification vs independently-violable-constraint refinement above with an operational test.
A count, not a weighted measure
This is a count of distinct burdens. It does not weight burdens by cost, magnitude, or importance. A trivial record-keeping duty and a major capital-adequacy requirement each count as one. Magnitude is deliberately out of scope — weighting by cost would import unbounded subjectivity. The measure is a refinement of the restriction-counting tradition (of which QuantGov’s RegData is the best-known example): where a restriction word-count proxy simply tallies prescriptive terms, this measure deduplicates cross-references, counts at source, distinguishes conditions from burdens, and filters non-operative language — a curated count rather than a term-frequency proxy. That it is a count and not a cost-weighted measure is a stated methodological choice, not an omission.
Labeller instruction — do not filter by perceived importance. Every qualifying burden counts equally, regardless of how minor or major it seems. Do not under-label duties that feel trivial.
Parallel jurisdictional instruments count separately. Where the same regulatory control is enacted in parallel devolved instruments (e.g. a GB/E&W regulation and its Scottish, Welsh, or NI counterpart — uksi / ssi / wsi / nisr), the primary count is residence-based: each instrument’s duty is its own burden, exactly as the unit rule and the amendment rule already work. Labellers NEVER make a sameness or duplication judgement — classify what is in front of you. The deduplicated “distinct rules” view (how much distinct regulatory control society faces) is a downstream analysis lens built post-labelling by a family-linker; it does not affect labelling. (Two-lens design; the methodology reports both lenses with the rationale.)
Economic vs personal — a disaggregation axis
Because the scope is broad (all private actors, including individuals in personal capacity), each burden also carries a tag recording the capacity in which it falls. This lets the broad “control on society” total be disaggregated into an economic subset without re-labelling. This axis is separate from, and applied after, the private/public gate above: a burden is first confirmed to fall on a private actor (or it is excluded), and only then tagged economic or personal.
- economic — the duty attaches to VOLUNTARY MARKET ACTIVITY on the burdened actor’s side: exchange, or production for exchange — selling, buying, letting, hiring, employing, working for pay, or producing/supplying for the market. It is economic whether the activity is formal or informal, and on EITHER side of the transaction (supplier or acquirer). A private individual selling a car on an online marketplace is economic; a business buying supplies is economic; a consumer’s contractual duty to pay the trader (CRA 2015 s.25) is economic. The economic role, not the legal form of the actor, is what matters.
- personal / civic — NO market activity on the burdened actor’s side: the duty attaches to ownership, status, conduct in ordinary life, or the use of goods/services already acquired, not to any exchange or production the actor undertakes. Seatbelt rules, firearm registration, and dog microchipping are personal — the duty attaches to owning or doing, not to a market transaction the actor is engaged in.
- both / either — the burden applies regardless of whether the actor is acting economically or personally (much “any person who does X” regulation). Discriminator: use both/either when the burden genuinely applies in either capacity; use ambiguous when you are unsure which capacity it falls in.
- ambiguous — used GENEROUSLY. When the economic/personal call needs more than a moment’s thought, tag ambiguous and move on; do not deliberate. The ambiguous set is quarantined and revisited as a focused later-phase pass (potentially via LLM classification against a refined rule, validated on a held-out sample). This deliberately defers the fuzzy boundary rather than forcing an inconsistent call at labelling time.
Classify by whether the duty attaches to VOLUNTARY MARKET ACTIVITY on the burdened actor’s side. The discriminator is the market/non-market character of the actor’s OWN activity — not the grammatical subject, and not downstream market effects. Ask: in the activity this duty attaches to, is the actor exchanging or producing-for-exchange (economic), or merely owning / being / doing / consuming (personal)? (This redefinition supersedes the earlier activity-profile formulation.)
- Economic — market activity on either side. Selling, buying, letting, hiring, employing, working for pay, or producing/supplying for the market — formal or informal — is economic, and BOTH parties to the exchange are in economic capacity (supplier and acquirer alike). Informality does not demote it: a one-off private online sale is still market activity.
- Compulsion carve-out (side-specific). A transaction the regulatory scheme itself COMPELS does not, by that compulsion, confer economic capacity on the compelled party — classify that party by the underlying activity the duty is bolted to. A dog-microchipping fee and an MOT are PERSONAL for the owner/driver (the compulsion is bolted to ownership / to driving in ordinary life), even though money changes hands. Compulsory motor insurance follows the underlying driving duty: both/either where it binds all drivers undifferentiated, personal where the scheme splits out the private side. The PROVIDER side is acting in its trade and is economic: the vet, the test garage, the insurer. Compulsion carves out only the side the scheme compels.
- personal — no market activity on the actor’s side: the duty attaches to ownership, status, conduct in ordinary life, or the use of goods/services already acquired, with no exchange or production-for-exchange BY THE ACTOR. Two refinements. (a) Ownership: held OUTSIDE any trade or business is personal (a privately-owned car); ownership embedded in market activity — business assets, property held for letting, trading stock — is economic (a company fleet car); an undifferentiated duty on owners as such is both/either. (b) Use of goods/services already acquired: the acquiring transaction is economic (either side, per the CRA 2015 s.25 ruling), but a duty attaching to the subsequent private use is personal — a consumer’s contractual duty to pay is economic, a TV licence (attaching to watching) is personal.
- both / either — an undifferentiated duty that binds market and non-market doers alike under one standard (driving, land ownership, possession, waste handling). Requires the POSITIVE judgement “binds identically in either capacity”; mere uncertainty is ambiguous.
- ambiguous — used GENEROUSLY. If the capacity call needs more than a moment’s thought, or the market/non-market character of the activity cannot be determined, tag ambiguous and move on.
- Splits — two kinds, handled differently. (a) A ROLE-split within one exchange (trader vs consumer) does NOT split capacity: both sides are economic, and the distinction lives in the obligated_party field, not the capacity tag. (b) A MARKET-vs-NON-MARKET split of an activity (driving for hire vs driving privately) DOES decide capacity: the hire duty is economic, the private-driving duty personal, and neither is both/either.
- Contrast cases: a private individual selling a car on eBay → economic (market activity, informal, seller side); a business buying supplies → economic (acquirer side); a dog-owner’s microchipping fee → personal (compulsion bolted to ownership); the vet charging for the chipping → economic (their trade).
Classify by capacity, not by market effect — nearly every regulation has some economic consequence (a seatbelt law affects car sales). That does NOT make it economic. Classify by the capacity in which the burden falls — commercial activity vs ownership/status/daily life — not by whether the rule has downstream market effects.
A floor, not a partition — and what the economic figure includes. Because ambiguous is used generously and quarantined, the phase-1 economic figure is a LOWER BOUND, reported as “at least X burdens fall on voluntary market activity”, not “X% of burdens are economic”. The economic floor = burdens on voluntary market activity, which INCLUDES consumer-side transactors (a consumer’s contractual duty to pay is a burden on market activity). A separate, narrower BUSINESS-BURDEN cut = economic AND a business-side obligated_party — defined and reported as its OWN line, not folded into the floor. Pre-emptive note (answered in advance): the economic figure includes consumers transacting, by design — “economic capacity” means the duty attaches to market activity, not that the actor is a business; the business-only figure is the separate cut.
Capacity — ambiguous means unresolvable, not broad (v1.1). Reserve the ambiguous capacity tag for cases where the economic/personal character genuinely CANNOT be determined — not as a catch-all for breadth. An undifferentiated duty that binds market and non-market doers alike under one standard is both_either (the positive “binds identically in either capacity” judgement), NOT ambiguous. Worked example: the UK GDPR controller/processor duties bind whoever processes personal data, in any capacity, under one standard → both_either. (This does not narrow the generous use of ambiguous for genuinely fuzzy economic-vs-personal calls; it distinguishes undifferentiated-breadth, which is a positive both_either call, from true unresolvability.)
Note — the obligated party is recorded separately for every burden (see the schema mapping and section 5B); the economic/personal tag is a classification of that party’s capacity.
The two-part test for any provision
A provision is a regulatory burden only if both are true:
- Operative — it actually imposes a binding obligation or prohibition (not merely defines, deems, describes, or enables).
- Falls on a private actor — the burden lands on a private actor (per the definition above), not on a public body or no one.
If either test fails, the provision is not a regulatory burden. Most classifier errors come from failing to apply these two tests rigorously — see sections 4 and 5.
Statutory private-law duties are in scope — both sides. A statutory obligation with legal force on a private actor counts whether the parent statute is “regulatory” or private-law, whether it binds the trader side or the consumer side, and whether it is a mandatory term or a displaceable default. The measure counts statutory obligations as defined here, NOT a pre-sorted “regulation” subset. Worked example: Consumer Rights Act 2015 s.25 (on delivery of the wrong quantity, a consumer who accepts the goods “must pay for them at the contract rate”) → 1 burden; category 2 (conditional — operational); polarity obligation; obligated_party the consumer; economic capacity (a duty on market activity, consumer side — see the capacity axis in section 1).
2. The six categories
1. Direct burden
A standing obligation or prohibition that applies at all times, requiring compliance activity continuously.
Example: “Every employer shall insure against liability for injury to employees.”
Licensing: holding and maintaining a licence to operate is a Direct burden — a standing constraint on operating in a regulated market. (The one-off act of applying is the acquisition of the licence; the ongoing requirement to hold one is the burden.)
Cat 1 vs Cat 2 boundary: Cat 1 = the burden persists continuously once the actor is in scope (e.g. “must hold employers’ liability insurance”). Cat 2 = a fresh, discrete compliance act is triggered by each discrete event (e.g. “must report each workplace death within 10 days”).
2. Conditional burden (operational)
A burden activated by an organic operational event within the private actor’s own activities — the actor anticipates or controls the trigger — requiring an immediate, direct compliance action when it occurs. Direct in nature when triggered; lower ongoing cost than an always-on direct burden because the trigger is intermittent.
Anchoring line: the trigger is an event or circumstance arising within the actor’s own activities, whether or not the actor deliberately caused it.
Contrast pair: a workplace death → Cat 2 (operational); a regulator’s inspection demand → Cat 5 (regulator-triggered).
Example: “An employer must report a workplace death to the HSE within 10 days.”
3. Implied burden (IB)
A burden revealed through a defence provision rather than stated as a direct command. Passive: the actor can satisfy it by an after-the-fact explanation, without pre-existing operational setup.
Example: Wild Mammals Protection Act — a defence to show an animal was killed as an act of mercy; the actor need only explain the circumstances if challenged.
Count-at-source caveat: where a defence reveals a standard that is ALSO the subject of a primary prohibition elsewhere, count the burden once, at the primary prohibition — do not count both the prohibition and the defence for the same rule (see count-at-source, section 3).
4. Implied burden active (IBA)
As IB, but the defence structurally requires an active, pre-existing compliance programme built and maintained before any challenge arises. Cannot be satisfied by after-the-fact explanation.
Example: Bribery Act s.7 “adequate procedures” — the organisation must have built and maintained an anti-bribery programme in advance.
Dividing line (IB vs IBA): significant pre-existing compliance infrastructure, built and maintained in advance = IBA. After-the-fact explanation suffices = IB.
5. Conditional burden (regulator-triggered)
A burden activated by an official discretionary decision of an EXTERNAL regulatory authority — the regulator (not the actor) controls the trigger, such as a notice, warrant, or inspection — to which the private actor must submit. The actor may never face it, but must be able to comply if they do. (Distinguished from Cat 2 by who controls the trigger: Cat 2 = the actor’s own operational event; Cat 5 = the regulator’s decision.)
Example: “A person must produce records when an inspector demands them.”
6. Ambiguous
Genuinely unclear category classification after careful reading. Flag rather than force. Use sparingly for the CATEGORY call. (Distinct from the economic/personal “ambiguous” tag in section 1, which is used generously; the category-6 label is for genuine uncertainty about which of the five substantive categories applies.)
Cat 6 and polarity: a Category-6 provision still routes to the review queue; polarity is not forced on it. Both the category and the polarity of a Cat-6 provision go to human/LLM adjudication together.
Precedence — defence-revealed expression dominates the category call
Where a burden is revealed through a defence, that expression dominates the category call: the burden is IB or IBA (per the IB/IBA dividing line above) regardless of its underlying trigger profile. This is logged as a deliberate simplification — it collapses the trigger-timing axis (Cats 1/2/5) into the defence-structure axis (Cats 3/4) for defence-revealed burdens. Revisitable if the pilot shows defence-revealed burdens with materially different trigger profiles actually mattering (see section 7).
Excluded: public body obligations
Obligations on public bodies (per the definition in section 1) are not regulatory burden and are excluded. See section 5C for the function test that resolves hard/hybrid cases.
2A. Polarity (an attribute on every burden)
Independently of its category, every burden is tagged for polarity — whether it requires the actor to act or to refrain. This is a separate axis from the six categories: a burden of any category may be an obligation or a prohibition. It is captured because positive obligations and prohibitions carry materially different compliance costs. (This is the one place where the obligation/prohibition distinction is determined; elsewhere the rules refer to “burden” and treat both polarities identically.)
The rule — classify by the operative requirement on the actor:
- obligation — the provision requires the actor to DO something (file, maintain, report, ensure, hold).
- prohibition — the provision requires the actor to REFRAIN from something (must not, shall not, may not, no person shall).
- review — genuinely ambiguous AFTER applying the rule; routed to human/LLM adjudication. Reserved for genuine post-rule uncertainty, NOT a dumping ground for mild doubt.
Edge cases (resolved by the operative-requirement rule, not the surface verb):
- “Must ensure that X does not happen” → obligation — it requires the actor to act (to take steps to prevent), despite the embedded negative.
- “Must not permit X” / “must take steps to prevent X” → obligation — the operative requirement is a positive duty to act, despite the surface “must not”. Do not read these as prohibitions on the strength of the negative verb.
- Offence-as-obligation (“is guilty of an offence if he markets…”) → prohibition — the operative requirement is to refrain, expressed through criminalisation.
- Defence-revealed polarity — a burden revealed through a defence takes the polarity of the underlying standard it reveals, not the grammar of the defence: Wild Mammals Protection Act → prohibition (the underlying standard forbids the act); Bribery Act s.7 → obligation (the underlying standard requires adequate procedures).
- Both / unclear — a provision inseparably imposing BOTH a positive obligation and a prohibition, or where the operative requirement itself is unclear → review.
Note: polarity is a PROPOSED label emitted by the LLM labellers, cross-checked against the other model and adjudicated by the human lead. Surface verb-based detection mishandles the edge cases above — e.g. “must ensure X does not happen” reads as a prohibition on its verb but is an obligation — so polarity is cross-checked, not trusted.
Polarity conventions (v1.1). (a) Refrain-entirely vs when-doing. A duty to refrain from an activity ENTIRELY is a prohibition; a duty that, WHEN you undertake an activity, requires you to perform it in a certain way is an obligation. Contrast pair: a “no person shall do X” ban → prohibition; a “when doing X, do Y / do not commence X unless Y” requirement → obligation (CDM reg 9(1), “a designer must not commence work … unless satisfied that …”, counts as an obligation on the design activity, not a bare prohibition, because it requires positive steps when the activity is undertaken). (b) Per-condition polarity for gatewayed condition lists. Where an ensure-chapeau (“the operator ensures that—”) gateways a list of individuated conditions that are counted as separate burdens, polarity is assigned PER CONDITION (each (a)…(n) is an obligation or prohibition on its own terms), not inherited wholesale from the “ensure” wrapper. Section 2A’s “must ensure X does not happen → obligation” example applies only where the ensure-duty is ITSELF the single counted burden; it does not override per-condition polarity once the conditions are counted separately (burials Sch 2 para 7: (c) and (m) are obligations, the other twelve prohibitions).
Database schema mapping
Numeric category IDs (1–6) are canonical and rename-proof; the category names are presentational only — see category_mapping.md. The data, model, and pipeline key on the integer ID; renaming a category (here or for a publication) never touches data, model, or code.
For pipeline alignment, each labelled burden carries, as separate fields:
- category — one of the six category tags (canonical numeric ID 1–6; string tags per category_mapping.md, covering Cats 2–6). (Category 1 is tagged ‘direct’. Note: ‘private_actor’ is the PARENT class — the provision imposes a private-actor burden — not a category tag.)
- polarity — obligation / prohibition / review.
- obligated_party — who the burden actually falls on (e.g. “employer”, “operator”, “any person”), recorded separately from the grammatical subject of the sentence. Essential for re-attributed rights (section 5B), where the sentence subject is not the burdened party.
- actor_capacity — economic / personal / both / ambiguous (section 1). A coarse classification of the obligated party’s capacity; “ambiguous” used generously.
- introduced_by / introduced_year — the amending instrument (and year) that introduced the burden; the parent instrument itself for original text. NULLABLE, best-effort from markup/metadata — labelling never stalls on lineage. Rationale: residence (which instrument the burden is part of — keys the count) and provenance (which instrument introduced it — keys the growth-over-time analysis) are separate attributes over the same record. Coverage caveat: verified 100% recoverable on the modern, CLML-clean 7-Act test set; corpus-wide coverage TBC (older consolidations, where amendment annotation is thinnest, would plausibly sag). Provenance is resolved at the amendment-fragment grain, not the section (one section can be patched by several instruments).
- family_id — NULLABLE; links duplicate-family burdens across parallel jurisdictional instruments (the “distinct rules” lens). Populated POST-LABELLING by the family-linker; labellers never set it (they never make a sameness judgement — see section 1).
- orphan — true when the burden reached classification through the orphan-triage lane (section 6); lets that slice be audited without being shaped differently.
- frontier_hook / frontier_target_type — frontier_hook is true for a burden counted as a frontier proxy under the deepest-visible-layer rule (section 3): a statutory duty to comply with an out-of-measure instrument. frontier_target_type ∈ {permit_licence, notice, byelaw, regulator_rulebook, other}. Recording them makes the proxy population enumerable, so any future scope expansion (re-classifying proxies to counted_at_source as the layer’s contents are counted) is a query, not archaeology.
- source_target — (v1.1) for a counted_at_source exclusion or a frontier proxy, the provision (or out-of-measure instrument) at which the burden is actually counted. NULLABLE, but required whenever counted_at_source or frontier_hook is set. A source_target naming a provision OUTSIDE the visible candidate is PROVISIONAL pending verification of that target’s text (see the named-target discipline, section 3).
- unapplied_amendment — NULLABLE flag; true where a burden is an enumerated pending insertion — an amendment enacted but not yet applied to its target in the corpus’s consolidated text (the editorial-backlog tail noted in the section-3 amendment rule). Lets that bounded tail be counted or quantified rather than silently missed; set only where the pre-labelling unapplied-amendment enumeration flags the fragment.
Typed exclusions (records that carry no burden). A candidate that is excluded is recorded with a typed exclusion — exclusion_family plus a best-effort exclusion_subclass — not dropped silently; excluded candidates are also worked negative training data. Families: non_operative (section 4), counted_at_source (section 3), public_body_or_no_one (sections 1 / 5C), and structural — operative-but-not-a-burden under the counting rules, a distinct joint from non_operative (which has no operative content). Sub-classes are the matched pattern (e.g. within non_operative: deeming, definitional, machinery_procedural, powers_to_make_secondary, scheme_machinery, list_of_contents; within counted_at_source: cross_reference, compliance_hook, enabling_power, penalty_as_consequence, amendment_machinery (textual amendments — machinery counted once at the consolidated target, section 3), and secondary_offence_reference if distinguished; within structural: bare_permission, scope_eligibility, condition_factor_list, single_act_specification, procedural_right_v_state, liability_attribution, burden_removal). mixed_other is available in any family for dual-pattern cases. Two-tier rule: dual-model agreement is computed on the FAMILY (the count-relevant call); sub-class mismatches within an agreed family are logged, not adjudicated.
Multiplicity (unresolved — see section 7): the unit rule allows one candidate to contain several burdens and one burden to span several candidates. The schema must be able to represent this (a burden-count per candidate, and grouping across candidates) or the unit rule cannot be recorded. This is a prerequisite to resolve before labelling begins.
3. Decision rules (settled)
Count the burden at its source. Count each burden once, at the provision that actually imposes it — not at every provision that references, penalises, or reveals it. A cross-reference (“a requirement under section 5”, “such a payment”, “contravenes section 4”) does not create a new burden; the burden sits at the referenced provision. Offence cross-references and penalty-as-consequence (below), the enabling-power rule (below), the amendment-machinery rule (below), and the defence-revealed IB caveat (section 2) are all instances of this general rule.
Textual amendments are machinery, not burdens. A provision whose operative content is the textual amendment of another instrument (“for X substitute Y”, “after subsection (2) insert—”) is NOT itself a burden — it is machinery. The burden it creates or modifies is counted once, at the amended target in its consolidated in-force form (which is what the corpus holds). Counting the amending instruction AND the consolidated target double-counts. Honest caveat: legislation.gov.uk’s revised text carries an editorial backlog of unapplied effects, so a small tail of amendments is not yet reflected in their targets — a known, bounded under-count, not hidden.
Enabling powers vs burdens imposed by the Act itself. A provision that itself imposes a burden on a private actor is counted in the Act, even where a public body administers or enforces it — a statutory requirement to hold a licence (obligation) and a statutory ban on an activity (prohibition) are both burdens of the Act. BUT a power (or duty) to make secondary legislation imposes no present burden on any private actor and is NOT counted (whether the future regulations would create obligations or prohibitions); the burden, if it ever exists, is counted in the resulting instrument, which is in the corpus. The test: does any private actor bear a burden on the strength of this provision alone (count), or does the provision only authorise a public body to create burdens later (do not count)?
Compliance / contravention hooks. A provision that merely requires compliance with — or prohibits contravention of — regulations made under it (“a person must comply with requirements imposed by regulations under this section”; “must not contravene regulations under this section”) is NOT a separate burden. It carries no content of its own; the burden is whatever those regulations impose, counted there. Do not count the hook.
Frontier proxies — count each burden once, at the deepest layer the measure can see. This generalises count-at-source across the boundary of what is in the measure. A compliance/contravention hook whose target is IN-measure (in-corpus legislation — regulations under this section, another Act) stays EXCLUDED: the target is counted at its own home, and counting the hook as well would double-count. But a statutory duty to comply with an OUT-OF-measure instrument — an administrative instrument (permit, licence, notice), a byelaw, a regulator rulebook, or any binding target not in the corpus — IS the burden, counted once as a frontier proxy: it stands at the visible frontier for the invisible layer of rules behind it, which the measure cannot otherwise see, so if the compliance duty were dropped that entire layer would be captured nowhere. Worked examples: Environmental Permitting reg 38(2) (“It is an offence for a person to fail to comply with or to contravene an environmental permit condition”) — the permit conditions are out of measure, so the duty to comply is counted; “contravention of a byelaw made under this section is an offence” — the byelaw is out of measure, so the compliance duty is counted. Contrast “must comply with requirements imposed by regulations under this section”, which stays excluded (those regulations are in the corpus, counted at their own home). The count is ONE proxy per out-of-measure target — never an attempt to enumerate that target’s internal contents (individual permit conditions, individual byelaw rules), which are outside the statute book. The scope-expansion flip (state it explicitly): counts are NOT additive across scope expansions — if the measure’s scope ever widens to take in an out-of-measure layer (e.g. a regulator’s rulebook), that layer’s frontier proxies re-classify to counted_at_source as the layer’s own contents are counted, so the same rule is never counted in two scopes at once. (The full scope disclosure — what is in vs out of measure, binding vs non-binding — is in the coverage/scope methodology note.) Frontier burdens are tagged frontier_hook + frontier_target_type at labelling time (schema mapping, §2A), so the proxy population is enumerable and any future scope expansion is a query, not archaeology.
Meta-duties — “ensure/verify compliance with X”. A duty to ensure or verify compliance with some body of rules X is a compliance hook, and follows the hook rule’s in/out-of-measure test. Where X is IN-measure legislation, the meta-duty is excluded and the burden is counted at X’s own provisions — e.g. Retained Regulation 1169/2011 Art 8(2) (“shall ensure the presence and accuracy of the food information in accordance with the applicable food information law”) points at Articles 9-onwards of the same instrument, counted there. Where X is OUT-OF-measure (a hypothetical “must ensure compliance with the FCA rules”), the ensure/verify duty counts as the frontier proxy for that layer (frontier proxies, above). Convention: a verify limb FOLDS into the hook unless it imposes distinct, independently-violable content of its own — a duty to keep verification records, or to run checks on other parties — which counts separately under the unit rule. (Pilot watch-item: whether labellers apply the verify-folds convention consistently.)
One statutory home per rule-layer (frontier proxies). A comply-with-rulebook frontier proxy has ONE home per out-of-measure rule-layer and is counted once THERE, never re-counted at each rule-mandating section. For the FCA Handbook that home is the s.137A general-rule-making-and-enforcement complex; when FSMA 2000 is labelled, do not re-count the proxy at every “the FCA must make rules …” provision (there are dozens). Guards the proxy count against inflation.
Offence-as-obligation, and its boundary. A provision that creates a burden by criminalising its breach IS a burden, classified by the obligation or prohibition it implies — BUT only where the offence is the primary statement of the rule itself, not a secondary cross-reference to a burden defined elsewhere. “A person is guilty of an offence if he markets a knife as suitable for combat” = primary prohibition = direct burden. “A person who contravenes section 4 is guilty of an offence” = secondary cross-reference; the burden is section 4, not this provision. (TNA’s dataset misses primary offence-obligations systematically; capturing them is a distinctive contribution.)
Penalty-as-consequence. Penalty terms (“shall be liable to”) are filtered out when a cross-reference marker is present (“guilty of an offence under”, “contravenes”, “fails to comply with”) — the penalty is the consequence of a burden stated elsewhere. Where the penalty provision IS the primary statement of the burden, classify normally.
Chain-ordering — count at the requirement, not the offence that enforces it. Where one provision creates a burden (a power to require, or a requirement) and a separate provision makes failing it an offence, count the burden once, at the requirement; the offence provision is enforcement, not a second burden. Worked example: Environmental Permitting reg 61(1) (an authority “may require [a] person to provide such information … as is specified in the notice”) creates the information duty (a regulator-triggered conditional burden on the person); reg 38(4)(a) (“to fail to comply with a notice under regulation 61(1) … without reasonable excuse” is an offence) merely enforces it — count at 61(1), not at 38(4)(a). This orders the offence-as-obligation and penalty-as-consequence rules: an offence that cross-refers up a chain to a requirement stated elsewhere is enforcement of that requirement, counted there.
Accessory conduct vs attribution of another’s breach. Specific accessory conduct — a defined act the actor itself must not do (knowingly cause, knowingly permit, aid) — is a distinct operative prohibition and counts: Environmental Permitting reg 38(1)(b) (“knowingly cause or knowingly permit the contravention of regulation 12(1)(a)”) is a burden in its own right. BUT a provision that merely attributes an existing breach to another party, stating no new conduct standard, is machinery — do not count: reg 38(6) (“if an offence … is due to the act or default of some other person, that other person is also guilty”) and ERA 1996 s.47B(1B) (“that thing is treated as also done by the worker’s employer”) reallocate liability for a breach counted at its own provision (attribution/deeming, section 4).
Multiple modal verbs in one sentence. A sentence containing multiple prescriptive terms is ONE burden unless it explicitly imposes genuinely distinct, independent operational requirements — in which case each distinct burden is a separate unit (section 1). Judge by the operative intent, not the count of modal verbs.
Definitional sub-clauses. Sub-clauses elaborating or defining part of a single burden are part of that one burden, not separate burdens.
Permission and exemption structures — the three-step resolution. A permission (“a person may …”), exemption, or carve-out (“does not apply”, “is not guilty … unless”, “except under”) is not classified on sight. (1) Detection only — record that a permission/exemption structure is present; its presence neither creates nor excludes a burden by itself. (2) Scope-vs-conduct sort — for each leaf apply the test: can the actor violate this leaf by its own conduct while otherwise remaining in the class? If yes, it is a conduct constraint (a burden candidate); if no, it is scope — it defines who or what is in the class (a gateway state) and folds, carrying no burden. Straddlers route to ambiguous-leaf review. (3) Unit rule on the conduct pile — apply the section-1 unit rule and the single-act-vs-independent-constraint refinement to the conduct leaves only. Worked set: Environmental Permitting Sch para 5(2) → the conduct pile (independent conditions, counted each — seven on this provision); reg 24(3) → one (conduct, but the components specify a single notification, so they fold); UK GDPR Art 6(1) lawfulness bases (a)–(f) → fold (gateway states — they define when processing is lawful, not conduct the controller performs); reg 40(2) and ERA 1996 s.1(5) → scope, fold (they carve out who the offence / the duty reaches, stating no conduct). This replaces the former bare-permissions rule; the earlier point survives within it — “may do X only if Y” still surfaces the embedded conduct condition Y at step 2, classified by the unit rule at step 3.
Constituted duties at a power provision — never zero a power-section on sight (v1.1). A provision framed as a regulator power (“the regulator may serve a notice”) can still CONSTITUTE a standing duty on the private actor at that same provision — where the section builds out a regime (the notice’s contents, the steps it may require, the effect of service) whose operative upshot is a duty the actor must discharge. Recognition test: does the section, taken as a whole, put the actor under an obligation a compliant actor could be required to perform — not merely authorise the regulator to act? If yes, count the constituted duty HERE (typically Cat 5, regulator-triggered), tagging frontier_hook where the operative instrument is the out-of-measure notice. Offence-side mirror: a separate offence of failing that notice is then counted_at_source at this constituting provision, not itself. Do NOT reflexively zero a section because its lead verb is “may”: a power section that constitutes a compliance duty is a burden-bearing section. Worked pair: Factories Act 1961 s.110 (a home-work-in-infectious-premises order constitutes the occupier’s duty) and Environmental Permitting reg 36 (the enforcement-notice regime at leaves 36(1)–(2) and the incident-notice regime at 36(5)–(6) each constitute the operator’s compliance duty; 36(7) withdrawal is machinery; the reg 38 offence counts counted_at_source → reg 36).
IB boundary and named-target discipline (v1.1). A defence-revealed category (implied burden / implied burden active, section 2) applies ONLY where the standard exists nowhere as a primary statement — an orphan standard revealed solely through a defence. Where a primary statement of the rule exists elsewhere, the defence or proof-allocation provision is counted_at_source at that primary provision, not read as a fresh IB. Every counted_at_source (and every frontier) claim NAMES its target in the source_target field (schema mapping, §2A); a claim whose target lies outside the visible context is PROVISIONAL pending verification of that target’s text. Worked pair: ERA 1996 s.98 (the s.98(1) allocation of the burden of proof is litigation machinery pointing upward) vs s.94 (“An employee has the right not to be unfairly dismissed” — the primary statement, where the re-attributed employer prohibition is counted). Process note: verifying a named target’s actual words caught a mis-read in the pilot — name the target and check it, do not count an IB on the strength of a defence when the primary rule is stated elsewhere.
The compliant-actor test — count impositions a compliant actor could bear; exclude breach-conditional liabilities (v1.1). A regulator-triggered or conditional imposition IS a burden where a fully COMPLIANT actor could still be required to bear it (a levy it must pay, an inspection it must submit to, a notice it must comply with). A liability that arises ONLY on breach — a penalty, forfeiture, or cost-recovery / restitutionary charge incurred only by having contravened — is an enforcement consequence, not a burden, and is counted_at_source at the duty it enforces. The line: a liability you bear (only if you breach) is a consequence; conduct you must perform (whether or not you breach) is a burden. Worked quartet: a periodic levy ✓ (a compliant actor pays it); a duty to submit to inspection ✓; a duty to comply with a served notice ✓ (reg 36); a local-authority cost-recovery charge ✗ (Land Reform (Scotland) Act 2003 s.23(4) — incurred only where the owner failed to reinstate; cf. the cost-recovery precedent in the 1875 public-health legislation). Liability-you-bear vs conduct-you-perform is the operative distinction.
Schedules. Provisions in schedules inherit their subject and burden from the parent provision.
4. Non-operative language — EXCLUDE
The single biggest source of false positives. These look prescriptive (“shall”, “must”, “may not”, “is to”) but impose no operative burden on anyone. Exclude.
- Deeming / interpretive: “shall be treated as”, “shall be deemed”, “is/are to be treated as”, “shall be presumed”, “shall be regarded as”, “are to be read as”. (subclass deeming)
Deeming — interpretive vs substance-creating (the one KEEP exception). INTERPRETIVE deeming, which classifies something for a scheme’s purposes, stays excluded (“any waste marked with an asterisk shall be considered as hazardous waste for the purposes of any legislation” — Commission Decision 2000/532 → 0 burdens). SUBSTANCE-CREATING deeming, which brings an obligation between private parties into EXISTENCE — the statutory implied/deemed contract term — IS a burden, re-attributed to the obligated party (section 5B): “every contract to supply goods is to be treated as including a term that…” (Consumer Rights Act 2015 s.13, goods to match sample) → 1 burden on the trader. Test: does the deeming merely label something for a scheme (exclude), or does it create a live duty one private party owes another (count)?
- Definitional: provisions that define a term or set out what something means. (subclass definitional)
- Machinery / procedural: “must be made by statutory instrument”, “shall not take effect until”, “shall continue in force”, “shall be defrayed out of”. (subclass machinery_procedural)
- Powers or duties to make secondary legislation: “the Secretary of State may by regulations…” — and equally “the Secretary of State shall make regulations…”: a duty to legislate is still not a present burden on any private actor (the burden, if any, arises in the resulting instrument, which is in the corpus). A power or duty to legislate, not a present burden (see section 3, enabling powers). (subclass powers_to_make_secondary)
- Scheme-machinery rules: rules about how a register or scheme operates. “An establishment may not be registered more than once” — a uniqueness rule on the register, not a burden on the actor; the burden of registering sits elsewhere. (subclass scheme_machinery)
- List-of-contents of a public instrument: an enumerated list describing what an order/notice must contain. Contrast list-of-requirements on a private actor, which ARE burdens (section 5A). (subclass list_of_contents)
Defence provisions — “it is a defence to prove…” (v1.1). Run the defence-revealed check FIRST: if the provision reveals a private-actor standard that has no other home (an orphan standard revealed only through the defence), it is an implied burden (IB/IBA, section 2) and COUNTS. Otherwise — where the underlying offence or standard is stated primarily elsewhere — the defence merely removes or limits liability and is a typed exclusion, structural/burden_removal, regardless of its subject-matter flavour. Worked pair: Trade Descriptions Act 1968 s.24 (the due-diligence defence reveals an orphan standard → counts, IBA) vs Bribery Act 2010 s.13(1) (a defence to the ss.1/2/6 offences, which are stated primarily → excluded, structural/burden_removal). This routes the defence-revealed IB caveat (section 2) and the burden_removal subclass to one test.
Recording exclusions (typed). An excluded candidate is recorded with a typed exclusion, not dropped silently — exclusion_family = non_operative here, plus the matched exclusion_subclass above; the count-at-source and amendment rules in section 3 record counted_at_source (with amendment_machinery as its subclass for textual amendments, counted once at the consolidated target), operative-but-not-a-burden cases record structural, and the public-body gate records public_body_or_no_one. mixed_other is available where two patterns genuinely apply. See the schema mapping (section 2A) for the two-tier agreement rule. Typed exclusions serve the extraction→count audit and are worked negative training data for Legal-BERT.
5. Hard cases and the precision distinctions
These are the cases a naive phrase-matching filter gets wrong. Each requires judgement, not a keyword. This catalogue is open-ended by design: it grows from pilot adjudications. A case that fits none of the distinctions below routes to review and, once adjudicated, may become a new distinction here; the parked watch-items (section 7) are the known frontier.
A. List-of-requirements vs list-of-contents
- KEEP (burden): each item imposes a requirement on a private actor — “the company must: (a) maintain X, (b) submit Y.” Where the items are independent operative requirements, each is a separate burden under the unit rule (section 1). Where instead they specify the components, contents, or particulars of a single required act (the contents of one record, the elements of one notification, the particulars of one statement), they fold into that one burden — see the single-act-vs-independent-constraint refinement in section 1. KEEP-each turns on whether each item is independently violable, not on list structure alone.
- EXCLUDE (non-operative): the items describe the contents of a public instrument — “the order must specify: (a) the date, (b) the area.”
Conceptual rule: resolve a list item’s subject from its governing parent clause; decide whether the items impose duties on a private actor or describe a public document’s contents. Do not decide on list structure alone. (See also the burden-vs-condition distinction in section 1: a list of factors qualifying one duty is one burden, not several.)
B. Substantive rights vs procedural/remedial rights
- KEEP (burden, re-attributed): a substantive right on a private actor that creates a correlative burden on another private actor. “A worker has the right not to be subjected to detriment” → the employer must not subject them → a genuine employer burden (a prohibition). Capture it, and record the obligated party (the employer) in the obligated_party field — not the grammatical subject (the worker).
- EXCLUDE: a procedural/remedial right against the state — e.g. a right to appeal to a tribunal, where the corresponding duty falls on a public body.
Discriminator: does the right create a corresponding burden on a private actor (keep) or on a public body / no one (exclude)? Do not decide on the grammatical subject alone.
C. Public / private actor classification — the function test
Elaborates the private-actor/public-body definitions in section 1. The discriminating test is the SOURCE of the body’s function — statutory versus contractual — not the public-ness of the work or the source of its funding.
- Public (excluded): created by or deriving its function from statute, exercising statutory powers, and accountable to a public regulator — regardless of funding. The FSCS and the Financial Ombudsman Service are public for burden purposes: their functions and powers are statutory, even though they are levy-funded. An obligation on such a body is regulatory machinery, not private-actor burden.
- Private (counted): a private firm delivering public services under contract is a private actor. Its role arises from a commercial contract, not statute; it holds no statutory powers; the burden falls on it as a private contractor. Contracted delivery of public-sector work does NOT make a body public.
Note — levy duties still count. Excluding a statutory body (e.g. the FSCS) does not exclude the statutory levy obligations that fall on the private FIRMS which fund it — those are burdens on private actors and are counted. Exclude the public body; count the private-actor duties around it.
Genuinely mixed bodies (e.g. a charity exercising some quasi-public statutory function) → route to review rather than forcing the call.
D. Context-dependent terms
Some terms flip classification by context. “Scheme manager” = the FSCS (public) in FSMA; = a private trustee/administrator (private, real burden) in occupational pensions. Resolve by surrounding context and the Act’s definitions, not the bare term. Where unresolved, route to ambiguous for human review.
Two labelling one-liners. (1) Passive-voice duty-bearer — where a duty is stated in the passive with no named actor (“records shall be kept”), resolve the duty-bearer from the context PROVIDED WITH THE CANDIDATE and record it in obligated_party; where the provided context cannot resolve it, route to review (review_reason = context_term) — never guess. (2) Third-party-private triggers — where a private actor’s duty is triggered by another private party’s act (HSWA 1974 s.2(7): the employer’s duty to consult arises on the employees’ request), classify by nearest fit — here category 2 (conditional — operational) — and the pilot logs the frequency of this shape.
6. Using this rubric in the validation pipeline
Both models classify against this exact document. Claude and Gemini receive identical rubric text and classify independently, without seeing each other’s labels.
- Agreements between the two models → high confidence; light human spot-check.
- Disagreements (on category OR polarity) → the high-value review queue; the human lead adjudicates against this rubric.
Human review concentrates on (a) all disagreements, (b) a deliberate slice of the hard categories in section 5, and (c) a random sample of agreements as a backstop.
Agreement is not the goal. The dual-model setup exists to SURFACE disagreement, not minimise it. Two models agreeing is reassuring but not proof — both can share a blind spot (e.g. non-operative modality). Periodic full independent human classification of an Act remains the ultimate backstop.
Orphan candidates. Some candidates arrive without a section tree (n_leaves=0, partial context — “orphan” material). These route to a triage lane BEFORE classification: an LLM first-pass clears clearly-non-operative fragments (recorded as typed exclusions, section 4) and escalates everything else; escalated fragments are enriched with their retrievable surrounding raw text and re-enter THIS same classification path, flagged orphan=true so the slice stays auditable. A labeller treats an escalated, context-enriched orphan exactly like any other candidate — there is no parallel pipeline. (Lane design: see the implementation plan.)
The human lead owns this document. Any rule that proves unclear or wrong is revised here first, then re-applied. If a rule changes materially, previously-labelled Acts may need re-checking.
7. Open items for the project lead
- Schema-multiplicity contract — AGREED (sign-off comment 10). Representation is resolved by the leaf-anchored burden-set schema (a burden-count per candidate, plus grouping/merging across candidates); the label store that persists it is NOT yet built and is scheduled BEFORE hand-labelling begins. The production emission architecture (per-candidate classification vs counting head vs BIO/span tagging vs decompose-then-classify) remains on the stage-2 allocation agenda; the leaf-anchored schema keeps those options open. (Selection criterion retained: the chosen architecture must emit usable confidence outputs, since production review depends on confidence-thresholding replacing model disagreement as the uncertainty signal.)
- DONE — full definition pass and the project-lead sign-off read completed 2026-07-17 (comment 11). This document is now authoritative (v1.0) for the labelling pilot.
- Test the burden-unit rule against real enumerated provisions. Before locking section 1’s unit rule, run the burden-vs-condition discriminator against real chapeau-plus-list provisions from the corpus — especially to confirm it cleanly separates lists of distinct burdens (count each) from lists of factors/conditions qualifying a single duty (count once). Parked, pilot-gated — revisit trigger: the hand-labelling pass + the dry run (the field-test of this section).
- Economic/personal boundary is deliberately deferred. In phase 1 the economic/personal tag uses “ambiguous” generously and the ambiguous set is quarantined; the economic figure is reported as a floor. The capacity axis (section 1) was REDEFINED in v2.6 to turn on voluntary market activity with a side-specific compulsion carve-out (superseding the earlier activity-profile formulation); it is new and pilot-gated. Refine the boundary and resolve the ambiguous pile in a later phase, against real accumulated cases — potentially via LLM classification validated on a held-out sample the model has not been coached on. Parked, pilot-gated — revisit trigger: pilot data, including the both/either base-rate question.
- Defence-revealed precedence is a stated simplification (section 2). Parked, pilot-gated — revisit trigger: a pilot showing defence-revealed burdens with materially different trigger profiles actually mattering (would motivate an orthogonal defence flag rather than collapsing into IB/IBA).
- The over-identification rate from validation is preliminary (small unflagged sample); not to be treated as established until a larger unflagged sample is validated.
- Add worked NEGATIVE examples from your own validated Acts as you go — real sentences that look like burdens but aren’t, with the reasoning. These instruct both models better than definitions alone. (Typed exclusions, section 4, are the structured form of these.)
- Confirm the polarity ‘review’ threshold is being applied with discipline once a batch has run, so the review bucket doesn’t bloat.
Version 1.1 — authoritative. Project-lead sign-off read of v2.6 completed 2026-07-17 (definitions signed off, schema-multiplicity prerequisite agreed, no structural findings); v1.0.1 (2026-07-18) aligned the exclusion-family taxonomy to the ratified v2 schema (recording-schema only); v1.1 (2026-07-21) folded six clarifications from the pilot adjudication session (constituted-duties rule, specifications-vs-embedded-duties test, IB boundary + named-target discipline with source_target, compliant-actor test, polarity conventions, capacity note) — no previously-traced counts change. This is the rubric the labelling pilot runs against. Category IDs are canonical (category_mapping.md); names presentational. Still open (pilot-gated, section 7): the production emission architecture (stage-2 agenda) and the parked watch-items.