CANONICAL EVENT OPERATOR GUIDE · v5

Bind the live environment.
Do not redesign the product.

LogSense reconstructs and measures what happened. HAIEC owns the frozen event governing instance, deterministic Control Test, exact evidence refs and limitations. The older v4 guide is preserved at the bottom only as an archived historical reference.

DISPLAY NAME != EVIDENCE IDENTITYINGEST TIME != EVENT TIMEMULTIPLE RECORDS != MULTIPLE ACTIONSLOGSENSE MEASUREMENT != HAIEC VERDICTLIVE DECISION != POST-RUN CONTROL SATISFACTION
GATE A · NEMOTRON

Organizer requirement is the agent path, not the HAIEC evaluator. Verify actual model identity of the supplied agents. If confirmed Nemotron-backed, that is the cleanest path. If we create/replace an agent, configure it to use Nemotron and retain model/invocation evidence.

GATE B · AWS ↔ AICT

Create the unique HAIEC connector and prove discovery. Required facts: connector identity, AWS account/role/log group bindings, discovered agent sys_ids/native IDs, and what session/trace/model evidence is genuinely visible.

GATE C · TELEMETRY

Prove one real gateway/OTel/CloudWatch record before freeze. If direct HAIEC egress is blocked, use runtime evidence → LogSense → Competition Evidence Bundle → HAIEC.

01

What is frozen

The HAIEC code baseline is not the event policy freeze.

READY / KEEP STABLE

HAIEC evaluator semantics, canonical evidence intake, C7/C9/C16 logic, result vocabulary, evidence binding and judge-query surfaces.

NOT YET FROZEN

Event-specific thresholds, baseline, cap, expected-event scope, source identities, evidence-binding profile and assessed-run selection rule.

FREEZE BEFORE RUN

After discovery + calibration, create the event governing instance. Freeze control ID/version, exact scope, threshold/baseline/cap, evidence requirements and source/binding profile.

RUN BINDING

If the platform generates the run ID at execution time, bind the exact generated ID under the already-frozen governing rules. Never choose a convenient run after seeing results.

Common evidence binding

Native run/session ID, authoritative RUN_START, trace/span IDs, agent execution + model/tool invocation IDs, gateway identity, ServiceNow sys_id/native AWS IDs, eventTime/observedAt/ingestedAt semantics, clock domain/skew, source qualification and assessed-run selection rule.

C7 · Event recording

Expected-event basis/manifest, required zones/enforcement points, stable event identity/order, timing-gap limit and allowed exception rate. Observed events alone cannot prove completeness.

C9 · Drift/performance

KPI name, producer, unit, direction, cadence, comparable-window rule, calibration data, frozen baseline snapshot/digest, drift threshold, violating-window allowance, alert owner/recipient and ServiceNow human-loop path.

C16 · Spend cap

Provider call IDs, input/output/cache token semantics, retry accounting, cost/token basis, hard cap, over-cap allowance if any, enforcement/refusal hook and definition of an executed provider call.

DISCOVER → QUALIFY → CALIBRATE → DECLARE → FREEZE → RUN → EVIDENCE → CONTROL TEST

If scope/threshold/baseline/cap/binding rules change after freeze: NEW POLICY VERSION → NEW RUN. Preserve the prior result.

02

First hour

Resolve interfaces and identity before touching architecture.

0–15 · CONNECT

Workshop/AWS, AgentCore agents, gateway, OTel/CloudWatch, digital twin/KPI, ServiceNow AICT connector path, HAIEC external connectivity and LogSense.

15–30 · IDENTITY

Run/session/RUN_START, trace/span, agent execution, model/tool call IDs, gateway ID, AWS native IDs, ServiceNow sys_ids, connector ID and Nemotron model identity.

30–45 · CONTROL FACTS

C7 expected-event basis; C9 KPI/baseline/human loop; C16 token/retry/cap/refusal accounting. Capture exact schemas/payload examples.

45–60 · QUALIFY

READY / LIMITED / BLOCKED / ONSITE VERIFY. Confirm assumptions with mentor/support before the first event freeze.

03

How the event stack feeds HAIEC

Each source must answer a specific evidence question.

AWS / AgentCore

Need account/region, agent/runtime/deployment IDs, run/session IDs, gateway endpoint and model/tool invocation IDs. Direct HAIEC ingestion is preferred only if sanctioned; otherwise LogSense normalizes/export-binds the evidence.

Agent Gateway

Treat as an enforcement + visibility candidate. Prove that the relevant invocations actually cross it. C16 real-time bonus requires a genuine synchronous pre-invocation/refusal hook; runtime ALLOW/DENY is not the post-run verdict.

OTel / CloudWatch

Primary runtime evidence sources. Capture schema/version, event-time fields, trace/span/call IDs, resource identity, sampling/retention/export method and mirrored-record behavior before counting C7/C16.

ServiceNow AICT

Create a unique HAIEC-named AWS connector. Prove agent discovery and capture connector ID + sys_ids/native IDs. Model/trace discovery may be incomplete; qualify what is actually visible. Use AICT/human-loop facts as evidence, not as HAIEC verdicts.

ODA / Cedar / Guardrails

Use only when approval/effective state, scope, actor and revision are established. A CR or policy file by itself does not automatically establish POLICY_AUTHORIZED; effective permission evidence is needed for EFFECTIVELY_GRANTED.

Digital Twin / KPI

Identify the C9 KPI producer, field, unit/direction and comparable-window semantics. Calibration can inform the baseline; freeze the metric profile + baseline before the assessed run.

Nemotron

Organizer requirement applies to the agent path: a valid solution needs Nemotron in the agents. Verify the supplied agents' actual model identity first. If they are Nemotron-backed, preserve that binding; if we create/replace an agent, configure it to use Nemotron. HAIEC's evaluator remains deterministic.

HAIEC MCP

Own MCP may be brought into the event and connected via AgentCore. Use it for query/operation friction reduction or a Nemotron-backed query assistant. MCP/LLM may resolve structured query parameters or explain persisted results; neither selects thresholds nor computes the authoritative verdict.

04

Five evidence planes

Never force one source to prove a different plane.

REQUESTED

What was asked / declared.

POLICY_AUTHORIZED

What an approved, effective policy/change permitted.

EFFECTIVELY_GRANTED

What the acting identity could actually exercise through IAM/ACL/runtime/gateway state.

CODE_CAPABLE

What exact deployed-bound source/config could cause. If source-to-deployment binding is absent: NOT ESTABLISHED.

OBSERVED

What runtime evidence shows actually happened.

A COPIED/FORKED REPO != DEPLOYED SOURCE.

Static analysis is only CODE_CAPABLE evidence when the analyzed snapshot can be bound to the running asset. Missing planes stay UNKNOWN / NOT ESTABLISHED.

05

Two questions, two tools

Do not collapse forensic reconstruction and control judgment.

WHAT HAPPENED? → LOGSENSE

Ingest, normalize, correlate, reconstruct, arbitrary time-window investigation where supported, KPI/token measurements, mirrored-record handling, source refs and gaps.

NO CONTROL VERDICT
DID THE CONTROL HOLD? → HAIEC

Frozen governing instance → exact run/window → deterministic calculation → SATISFIED / NOT_SATISFIED / NOT_EVALUATED → evidence refs → limitations.

QUERY != EVALUATE
LogSense repository is private.

If a teammate/Devin GitHub account cannot open it, request repository access first; once authorized, the repo root remains the safest onboarding entry because they may not have LogSense installed locally.

06

Control edge semantics

Truthful failure states are part of the demonstration.

C7

Observed events alone cannot prove completeness. Freeze expected event basis + assessed scope. Required evidence missing/unjoinable → visible gap or NOT_EVALUATED, not a green result.

C9

Freeze KPI/baseline/window/threshold/allowance. Human-loop path must exist and be evidenced. Record alert + named recipient/queue + acknowledgement/action when present; documented silence remains an observed governance fact, not fabricated remediation.

C16

Separate EXECUTED WITHIN CAP, OVER-CAP ATTEMPT PREVENTED and EXECUTED OVER-CAP BREACH. A runtime DENY is preventive evidence, not automatically an overspend breach. No atomic/no-overshoot claim unless proven.

07

Judgment Day

Bring reproducible proof, not a green dashboard.

Evidence + thresholds

Exportable evidence file plus dated/versioned threshold/governance document frozen before assessed runs.

Tool + test cases

Working Control Test/query tool plus test cases and named PASS/Breach runs under the same governing version.

Architecture + gaps

One-page execution/enforcement architecture and an honest gap/remediation/retest register.

2-MINUTE PROOF FLOW

Judge names control → select exact assessed run → show frozen governing instance → execute/query persisted Control Test → show arithmetic + verdict → open exact evidence refs → state gaps/non-claims.

08

Support + schedule

Operational details that prevent avoidable event-day failure.

SUPPORT

Stay connected to the team Microsoft Teams room. One SPOC posts in Main Meeting Chat, starting with the team name. Prefix urgent issues with Blocker. Use the onsite help desk in parallel.

TUESDAY 6 OCT

07:00–07:45 check-in · 08:00–10:00 final working time · 10:30 judging kickoff · 10:45–13:30 preliminary judging · 14:00 finalists · 14:30–16:30 finalist judging. Venue: Marriott Dallas Allen / Innovate Americas.

Archived v4 Field Guide — preserved reference, superseded where this canonical guide differs
DO NOT TREAT THIS SECTION AS THE CURRENT OPERATING AUTHORITY. It is preserved unchanged for provenance. Later immersion-session facts supersede earlier v4 assumptions, particularly Nemotron applicability, AWS↔AICT importance, human-loop framing, and live integration details.
Complete Event Field Guide v4.0 · Source-separated · Corrected

Trustworthy AI & Data Hackathon
The Operating Manual

Organizer requirements, our tooling, and HAIEC platform capabilities — kept distinct. Every claim is tagged with its source. The six required steps, the C16 worked-example shape, the C7 multi-point rule, and the ALLOWED block are now first-class sections.

TM Forum · ODA powered Oct 4–7, 2026 Dallas → Allen, TX 3 controls · 1 minimum verdictOwner: HAIEC
SOURCE KEY:
ORGANIZER From the nine official screenshots
OUR STACK LogSense/HAIEC architecture
HAIEC CAP Platform capability (not scored)
ONSITE VERIFY Confirm at venue
0

The six required steps — in order

ORGANIZER The exact organizer sequence. Do these in order. Each depends on the one before.

1
Choose your control(s)
Three to choose from: event recording (7), drift (9), spend cap (16). One completed control is the minimum; a second and a third are bonus.
2
Declare your threshold ranges
Metric, observation limit, exception tolerance, measurement basis, named owner. Versioned and dated before any assessed run — we check the date.
3
Decide where in the timeline you check
Post-run forensics is the floor: a tool that inspects evidence after the fact. Doing it continuously, inline as it runs, is bonus.
4
Build the control
At an enforcement point of your choosing, emitting records bound to the control, the threshold version and the run — not just a timestamp.
5
Build the control test / evaluator
The thing that returns a verdict: recompute the measure, reconcile expected against observed, count exceptions, compare with your rule.
6
Run it both ways, and record the gaps
One passing case, one breaching case. Do not move the threshold to make it pass. Write down what you could not close.
✓

You are done when: someone not on your team can ask “was control 9 satisfied for run FM-B?” and get an answer in two minutes, with the records behind it.

01

Event at a glance

ORGANIZER Identity, dates, venues.

Event
Trustworthy AI & Data Hackathon

Hosted by TM Forum · Powered by ODA

Dates
Sun 4 – Wed 7 Oct 2026

3 hack days + awards day

Days 1–2 venue
AT&T HQ / Discovery District

208 South Akard, Dallas, TX

Days 3–4 venue
Marriott Dallas Allen

777 Watters Creek Blvd, Allen, TX

Build a working solution that can answer whether a declared control was satisfied for a named run or time window — using records an independent judge can inspect. — Immersion Session Report, central outcome
02

What the event wants you to do

ORGANIZER The mission and the central question.

★The mission

Prove that collaborating AI agents resolving a telecom network fault meet declared, dated requirements — and that the proof is independently verifiable by a judge who is not on your team.

evidence-based assurance independent verification honest gaps
?The central question
“Was control X satisfied for run Y / window Z?”

The judge names the control and the run; your tool must answer within two minutes with supporting records an outsider can open and verify.

Axis 1 completion tiers

ORGANIZER Axis 1 is about how many controls you finished end-to-end.

1Minimum — 1 of 3

Complete one control end-to-end, with a passing and a breaching case. You are in the running.

2Ahead — 2 of 3

Complete a second control to the same standard. Ahead of a team with one.

3Full set — 3 of 3

All three controls complete to the same standard. 3/3 is the full set. Continuous compliance, adversarial testing and second enforcement point are separate Axis 3 bonus — not part of "the full set."

⚠

Claim aggregation matters. ORGANIZER The management claim C-FM-01 is only fully answered when all three control branches are covered. A completed subset (1/3 or 2/3) is partial coverage, not a pass.

03

The claim and the three controls

ORGANIZER One management claim; three branch controls; each with a measurable child specification.

CLAIM C-FM-01
“During fault resolution, the collaborating agents meet our declared event-recording, drift and token-spend requirements.”
7  Record the events AIA-LOG-001

Mode: Event-driven · every in-scope event recorded and joined to one run, across all three zones. (Organizer wording: "every in-scope event" — not "every meaningful action".)

Objective: All events recorded — coverage [C%], gap < [G ms], breaches ≤ [B7%].

Risk: Missing events, broken run links, excessive unexplained gaps.

Test: Reconcile expected records; calculate gaps; count breaches.

Evidence: Expected IDs, events, timestamps, run links.

Gap test checks timing, not completeness. Missing event ≠ proof the action never happened.
9  Monitor drift AIA-ARC-006

Mode: Continuous · live measure watched against a frozen baseline, with a named owner.

Objective: Drift ≤ [D%] against baseline, violating windows ≤ [B9%].

Risk: Drift goes unnoticed, or measurements themselves are missing.

Test: Recalculate drift; count breaches; check monitoring and alert response.

Evidence: Baseline, live values, windows, alerts, owner, required response.

A missing measurement is not zero drift. Raising an alert does not erase the breach.
16  Limit token spend ACN-COST-001

Mode: In-transaction · checked mid-action at the shared model gateway.

Objective: All agents plus retries, per run: total spend ≤ [N tokens]. Over-cap run policy declared separately.

Risk: A loop overspends, or calls are missed or charged to the wrong run.

Test: Reconcile usage; sum each run; compare with cap; check stop records.

Evidence: Call IDs, token counts, budget version, stop records.

Missing usage → PARTIAL, never zero. Genuine retries count again; duplicate telemetry counts once.
!C7 enforcement-point requirement ORGANIZER

The organizer worked-example footer states explicitly: Control 7 has to work at every enforcement point at once. It is not a per-point control that you can scope to a single gateway. The expected-event stream must span all three zones and every enforcement point that carries in-scope events for the run.

Implication: your C7 expected-event manifest and your C7 completeness claim must cover the full cross-zone path. A partial-governed-coverage claim loses marks.

!C16 over-cap allowance is a policy field ORGANIZER

The worked example says "zero over-cap runs permitted" — but the underlying requirement is that the team declares whether any over-cap run is permitted at all. This is a governed threshold decision, not example text.

Implication: your C16 policy must carry an explicit overCapAllowance field alongside the cap value N. Zero is a valid declaration. Any non-zero allowance must be justified and dated.

🔗

The claim → control → evidence chain: ORGANIZER a control is the safeguard; a control test checks it. Results flow back up to the claim, and a completed subset is partial coverage — not a pass.

Five ways a control gets checked

ORGANIZER Real controls from the EU AI Act library, including the three you will build.

📄Static

Somebody reads a document, occasionally. e.g. AIA-GOV-001 — AI policy approved by top management.

📅Scheduled

A job runs on a fixed cadence. e.g. AIA-INV-001 — complete AI system inventory maintained.

⚡Event-driven

Something happens and the check fires. e.g. C7 — automatic recording of events.

⏱In-transaction

Checked mid-action, before it completes. e.g. C16 — per-run spend cap.

📈Continuous

A number watched against a limit, always. e.g. C9 — production drift.

Compliance timing — before, while, after

BEFORE IT RUNS

Pre-deployment compliance. “May this go live at all?” Offline evaluation, policy-as-code, signoff, threshold declaration. Output: go / no-go, once.

WHILE IT RUNS

Continuous compliance. “Is this allowed, right now?” Evaluate, enforce, record. Output: allow, deny or alert, per action.

AFTER IT RAN

Post-run compliance. “Did the control hold between T1 and T2?” Recompute from records. Output: a dated pass or fail on the control.

⚠

The two things people confuse. ORGANIZER An Evaluator judges the subject — one action or measurement — in the moment. A Tester / Control Test judges the control — did it hold for everything in the window — after the fact. Different consumers, different output.

!Same frozen threshold for PASS and BREACH ORGANIZER

The organizer states explicitly: run it both ways and do not move the threshold to make it pass. The PASS demonstration and the BREACH demonstration must use the same frozen governing control/threshold version.

Implication: changing the threshold, baseline, cap or exception tolerance between the two runs manufactures the expected result and is a documented deduction. Freeze once, use twice.

04

Worked example — C16 at the shared model gateway

ORGANIZER The organizer's five-stage shape. Your implementation may differ — the shape should not.

“One enforcement point, one control, one threshold range, one question a judge can ask. Yours will differ, the shape should not.”

Declare
10,000 tokens per fault-resolution run All three agents and every retry included. Zero over-cap runs permitted. Frozen as G-FM-v1 and dated before the first assessed run.
Enforce
The shared model gateway is the enforcement point Every model call from every zone crosses it. On each call it asks the evaluator "is run FM-A still under cap?" and acts on the verdict.
Record
The gateway emits one record per call to your ledger Asynchronously, off the request path, bound to the run, the control and the threshold version — not just a timestamp.
Test
controltest 16 --run FM-A Reconciles every call for the run, sums input and output tokens, compares against the cap named in G-FM-v1, and returns a verdict with the call identifiers behind the number.
Show
FM-A 9,600 SATISFIED · FM-B 10,800 NOT SATISFIED One of two runs over cap, none permitted, so objective 16 is not satisfied — and the 800-token overshoot stays visible.
📐

The shape lesson: ORGANIZER C7 works at every enforcement point at once; C9 runs on a schedule against a frozen baseline; C16 runs in-transaction. Same record/binding shape, different execution mode. That separation is the point.

05

Allowed changes, bonus, and out of scope

ORGANIZER What you may change, what earns bonus, and what is explicitly not asked for.

✓ALLOWED ORGANIZER

The organizer explicitly permits:

  • Building more of or changing the supplied agents
  • Using NVIDIA Nemotron for agent development
  • Using Nemotron for adversarial testing of your solution
  • Modifying the deployment to add a second enforcement point — same design, moved

This is your permission to modify the lab agents and demonstrate adversarial testing.

+BONUS ORGANIZER

Rough order of value — chase only after core is stable:

+ second control, then a third + continuous compliance (verdict inline, in the path) + adversarial testing (show what broke) + second or more enforcement points (same design, moved)

Guardrail: never bonus-chase while a core control is incomplete or non-judgeable.

✗OUT OF SCOPE — do not spend time on these ORGANIZER

Signatures, hash chains or Merkle proofs over the ledger. Not asked for, not scored.

A slide deck. A whiteboard and a working tool beat both.

A human actually answering an alert. The alert, the named recipient and the silence/response evidence are sufficient. The response must still be recorded, even if no human responds.

⚠

C9 alert semantics — be precise. ORGANIZER A human actually answering is not required. But the alert, the named recipient, and the silence or response must still be recorded. "Alert issued" alone is insufficient — record who was told and whether they responded or stayed silent.

06

How LogSense and HAIEC map onto the event

OUR STACK Two tools, one contract, one verdict owner. This is our operating system.

OUR HARD LOCK
LOGSENSE MEASUREMENT ≠ HAIEC VERDICT
LSLogSense — independent forensic measurement OUR STACK

What it owns: raw/local evidence preservation · source discovery + mapping · run reconstruction · C7/C9/C16 measurement · coverage + timing measurement · gaps + limitations · evidence references + exports · operator + event guidance.

What it never emits: SATISFIED / NOT_SATISFIED as a judge-facing result · threshold enforcement · governing policy selection · a final Control Test result.

measurementState: MEASURED / PARTIAL / NOT_MEASURED verdict: null verdictOwner: HAIEC
HAHAIEC — governing policy, evaluator, Control Test OUR STACK

What it owns: control register · frozen control versions · thresholds, baselines, budgets · owners / effective dates · deterministic control evaluators · SATISFIED / NOT_SATISFIED / NOT_EVALUATED · judge-facing Control Test · live telemetry binder / intake · optional inline enforcement.

Platform capabilities: AI Action & Access Map · Compliance Twin · evidence-bound assurance evaluation · MCP surface for IDE/CI integration.

ALLOW / REVIEW / BLOCK Evidence-bound evaluation

The handoff path OUR STACK

Step 1

Measure in LogSense — active run → prerequisites → measurement (MEASURED preferred; PARTIAL with honest limitations is still handoff-able evidence).

Step 2

Read the handoff panel — measurement state, run/window, bundle ID, snapshot, evidence set, manifest/baseline/accounting refs, limitations.

Step 3

Export the Competition Evidence Bundle — deterministic JSON the evaluator consumes. Carries verdict: null, verdictOwner: "HAIEC".

Step 4

Run the Control Test in HAIEC — judge asks "Was Control 16 satisfied for run FM-B?"; HAIEC resolves control + run (+ window) → frozen policy → deterministic result.

Step 5

Drilldown stays in LogSense — the evidence behind the answer opens in the pack's per-run HTML views.

Capability truth — current repo state OUR STACK

ControlLogSense measurementHAIEC evaluator coreOrganizer-complete
C7 — AIA-LOG-001 AVAILABLE AVAILABLE closure gaps
C9 — AIA-ARC-006 AVAILABLE AVAILABLE closure gaps
C16 — ACN-COST-001 AVAILABLE AVAILABLE closure gaps

All three deterministic evaluator cores are on Main. "Organizer-complete" means: named PASS + BREACH runs, frozen threshold document, six Judgment-Day artifacts, and the two-minute judge rehearsal all closed. Update from repository truth before judging.

07

HAIEC platform capability

HAIEC CAP Platform features, not organizer requirements. Use them to strengthen your assurance case — do not present them as competition facts.

📌

Read this first. Everything in this section is HAIEC platform capability. None of it is an organizer requirement. The judge scores the six required artifacts and the three controls. HAIEC's capabilities help you produce that evidence — they are not the evidence itself.

The AI Action & Access Map HAIEC CAP

The Map connects seven layers that determine what an AI system or agent can reach, what it can change, and what happens when it acts. Each layer is backed by evidence — and each layer can have gaps. Unknowns are first-class citizens.

  • Identity & Access — who and what can act
  • Agents & Models — the AI systems in the request path
  • Tools & APIs / MCP — what the agent can invoke
  • Data & State — what it can read from or write to
  • Consequences — what outcomes an action can trigger
  • Evidence Basis — what evidence actually establishes
  • Unknowns — what is NOT established
🗺

How to use it: Before the assessed run, map the three agents against their tool/API/MCP surfaces to justify your C7 scope. During the gap list, use the "Unknowns" layer honestly. Do not present the Map itself as a scored deliverable — it informs your one-page architecture sketch and your gap list.

Evidence-bound evaluation and Compliance Twin HAIEC CAP

🔐Compliance Twin

Versioned assurance evaluation records with audit trails for compliance history tracking. Each threshold entry gets a version, a digest, and a frozen timestamp.

📜Evidence packages

Audit-ready evidence packages. Content digests are fine for provenance — but per the organizer, cryptographic sealing and hash chains are not scored. Do not position them as a winning feature.

MCP surface for IDE integration HAIEC CAP

HAIEC exposes a bounded MCP surface over a stateless HTTP transport. MCP is the preferred surface for AI IDEs and coding agents (Claude Code, Cursor, Devin, etc.). It provides the same canonical operations as the REST API, one bounded tool per operation.

haiec_preflight

Returns setup state, per-check truth, and exact required/optional actions with URLs.

haiec_run_assurance

Runs the full assurance pipeline (static + runtime + compliance) with provenance tracking.

haiec_get_agent_audit

Returns the report with resolution groups — deterministic grouping of verification items.

haiec_runtime_preflight

Read-only safety plan showing executable vs HELD categories.

haiec_run_runtime_test

Requires explicit user confirmation + authorization acknowledgment. Held categories never execute.

haiec_ingest_structured / logs

Evidence upload tools requiring explicit user approval and appropriate scope.

⚠

MCP and the second-enforcement-point bonus — be precise. The MCP surface can operate and test HAIEC itself. It does not prove the second-enforcement-point bonus. That bonus requires the same control architecture / policy / evaluator running at an actual second enforcement point in the lab, with nothing material rewritten in the governing rule. Use the MCP surface as a workflow enabler — do not present it as the bonus evidence.

✓

What the MCP surface is genuinely useful for in this event: running adversarial testing workflows from your IDE, capturing the attack/expected-defense/observed-break/change/retest cycle, and operating HAIEC during the judging flow. It reduces friction — it is not the assurance claim itself.

Five evidence planes — an optional deeper explanation HAIEC CAP

HAIEC models agent claims across five independent evidence planes. This is useful context for your own reasoning and gap analysis — but it is not an organizer requirement. The judge's required one-page architecture sketch asks for: components, execution path, and how an enforcement point invokes your control. Keep the sketch simple. Use the five planes as an optional deeper explanation in the gap list or the appendix, not on the required sketch.

  • Requested — what the request asked for
  • Policy Authorized — what policy permits
  • Effectively Granted — what credentials establish
  • Code Capable — what the code can technically do
  • Observed — what runtime evidence shows
  • A gap between any two is a finding.
08

Architecture pattern — separate the decision from the enforcement

ORGANIZER A starting point, not an answer. Draw your own version in the first hour and show a mentor.

The live path — already running before you arrive

Live
AI systemThe agents at work.
→
Gate
Enforcement pointsCan watch, or refuse.
→
Emit
TelemetryWhat the system emits.

The governance side — what turns signals into evidence tied to a control

📋Control register

Controls, thresholds, owners, versions. Hands out the current version. Never in the path.

⚖Evaluator

Compares a value against its threshold range and returns a verdict at the decision point. Stateless.

📚Evidence ledger

Binds a record to control, version and run. Sealing itself is out of scope for scoring.

🔍Control test / query

Answers "was control X satisfied between T1 and T2?" On demand.

One contract, four transports

⇄evaluate()

Synchronous, in the path — for in-transaction controls that must be able to refuse. Best fit: C16 at the shared model gateway.

↗emit()

Asynchronous, on an event bus — for everything that only has to be recorded. Best fit: C7 event records.

↩callback()

A webhook for a verdict that arrives later — deferred actions, human approval.

↻pull()

On a schedule — for scheduled and continuous controls, off the live path. Best fit: C9 drift measurement.

Five habits that keep the design honest

The telemetry store never judges

The evaluator never remembers

The control register is never in the path of execution

The evidence ledger never interprets

Records are sealed where they are made [sealing out of scope]

?ONSITE VERIFY — governed execution path details

The following are not established by the nine organizer screenshots. Confirm at the venue or from a separate official environment brief before treating them as fact:

  • Gateway name modaas-agw (agentgateway)
  • Cedar-policy authorization
  • Direct paths to Bedrock / AgentCore blocked in the workshop environment
  • Fail-closed routing behavior

Label these ONSITE VERIFY in your architecture sketch until confirmed. Do not present them as organizer facts.

09

How you will be judged

ORGANIZER Three axes. Axis 1 orders the field, Axis 2 decides the winner, Axis 3 separates teams that are level.

Axis 1 — How many of the three controls did you complete, end to end?

None

Nothing to score

1 of 3

The minimum · you're in the running

2 of 3

Ahead of a team with one

3 of 3

The full set

Axis 2 — Quality · where the prize is actually decided

The live test

A judge names a control and a window. Two minutes. If we have to trawl your logs, you lose the mark.

The architecture

Would it survive a second enforcement point? Is the register out of the live path? Is the evaluator stateless?

The evidence

Bound to control, threshold version, run. Verifiable by somebody who is not you. Gaps recorded, not left blank.

The honesty

"Built properly: yes. Actually worked: no" is the most useful answer a team can bring.

Axis 3 — Bonus points, and what earns each one

+ Extra controls

A second or third completed to the same standard.

+ Continuous compliance

A verdict formed inline, in the path, while it runs.

+ Adversarial testing

You attack your own design and show us what broke.

+ Second enforcement point

Same architecture, moved, with nothing rewritten.

✗

What loses marks. ORGANIZER A threshold with no date, or dated after the test · records covering only some of the governed actions · evidence only your own team can verify · a finding closed without a passing re-test · a control that stores its own threshold · a design that only works in the request path.

10

Deliverables — what to bring

ORGANIZER If an artefact is missing, we cannot score it — however good the work behind it was. Assemble these Saturday evening, not Sunday morning.

#ArtefactWhat it isRequired?
1Your evidence fileIn whatever form you chose. Exported, openable by us, and covering the runs you want scored.MUST
2Your threshold documentOne entry per control you took on: metric, limit, exception tolerance, measurement basis, owner, version and date.MUST
3Your control test toolThe thing we will drive. A command line, a script, an API, a small app. We type the question; you do not narrate the answer.MUST
4Two runs, namedOne that passes and one that breaches, with their run identifiers written down before we arrive.MUST
5A one-page architecture sketchComponents, what sits in the path of execution, and how an enforcement point invokes you. A whiteboard photograph is fine.MUST
6Your gap listWhat you could not close, and why. This is scored, and it scores well.MUST
✓

Bring one honest failure. ORGANIZER A team that says "built properly, did not work, here is the evidence and here is why" scores above a team with nothing but green ticks.

⏱

The two minutes. ORGANIZER We name one control and one run. You drive your own tool. Rehearse it twice before you present.

11

Schedule

ORGANIZER All times local. Work may continue outside scheduled hours — technical support will be limited.

DAY 1 Sun 4 Oct Lunch 12:30–1:30
AT&T HQ · 208 S Akard, Dallas
  • 9:30–10:30Check-in & badge pick-up
  • 10:30–10:35Day 1 Welcome — Stuart Dunn
  • 10:35–10:45Hackathon Opening — Guy Lupo
  • 11:00–11:30Opening Keynote Panel
  • 11:30–12:30Masterclass · How to Win
  • 12:30–3:00Teams receive access · working time
  • 3:00–3:30AWS Lightning Talk
  • 3:30–6:00Team working time
  • 6:00–6:30NVIDIA Lightning Talk
  • 6:30–6:35Day 1 Closing
DAY 2 Mon 5 Oct Lunch 12:00–1:00
AT&T HQ · 208 S Akard, Dallas
  • 9:00–9:20Check-in
  • 9:20–9:30Day 2 Welcome
  • 9:30–10:00Accenture Lightning Talk
  • 10:00–1:30Team Working Time
  • 1:30–2:00Dell Lightning Talk
  • 2:00–4:30Team working time
  • 4:30–5:00ServiceNow Lightning Talk
  • 5:00–6:30Team working time
  • 6:30–6:35Day 2 Closing
DAY 3 Tue 6 Oct Lunch 1:00–2:00
Marriott Dallas Allen · Starlight Ballroom 3
  • 7:00–7:45Check-in & Innovate Americas Check-in
  • 7:45–8:00Day 3 Welcome
  • 8:00–10:00Final Team Working Time
  • 10:30–10:45Preliminary Judging Kick-Off
  • 10:45–1:30Preliminary Judging
  • 2:00–2:20Finalists Announced & Day 3 Closing
  • 2:30–4:30Finalist Judging
DAY 4 Wed 7 Oct Lunch 1:00–2:00
Marriott Dallas Allen · co-located with Innovate Americas
  • 10:00–11:30Innovate Americas Keynotes
  • 11:10–11:25Hackathon Awards Ceremony
⚠

Location change after Day 2. ORGANIZER Days 1–2 are at AT&T HQ (Dallas). Days 3–4 are at the Marriott Dallas Allen (Allen, TX) — co-located with Innovate Americas 2026.

📌

Clarification needed: Day 4 shows Innovate Americas Keynotes 10:00–11:30am and the Hackathon Awards Ceremony 11:10–11:25am — these overlap by 15 minutes. Confirm with organizers.

12

Onsite logistics

ORGANIZER Check-in, badges, laptops, lunch, team changes.

📍Days 1–2 · AT&T HQ

208 South Akard, Dallas, TX — AT&T Discovery District.

Pre-event: you'll receive numerical codes by email from AT&T before arrival. Use them to check in at Building 1 each day. Bring ID. Badges issued Sunday; hackathon in Building 3, 12th-floor auditorium (visitors escorted after check-in).

📍Days 3–4 · Marriott Dallas Allen

777 Watters Creek Blvd, Allen, TX — co-located with Innovate Americas 2026.

Pre-event: register using the pass link code previously provided. Contact CatalystsOps@tmforum.org if you need it resent. Check in at event registration; badge issued. Hackathon room: Starlight Ballroom 3.

💻Bring your own laptop

Computers will not be provided.

⏰Arrive early

Allow ample time for morning check-in each day.

🍽Lunch provided daily

Notify organizers of any special dietary requirements.

🎟

Innovate Americas access. In-person Hackathon participants receive complimentary access to Innovate Americas 2026. Developers may attend conference sessions once Hackathon work concludes Tuesday afternoon. Access is exclusive to registered participants who take part in the full program Sunday–Tuesday — it is not standalone admission. Non-attendance on all days may change registration and incur a fee.

✉

Team changes: contact CatalystsOps@tmforum.org immediately.

Support mechanism — Teams + onsite help desk

💬Microsoft Teams
  • Invites sent by Friday, 02 Oct 2026.
  • Each team gets its own Meeting Room.
  • Nominate one SPOC per team; only the SPOC posts in the Main Meeting Chat, starting with the team name.
  • Once acknowledged, an organizer joins your room.
  • Flag urgent issues with "Blocker".
🛎Onsite help desk
  • Staffed help desk runs at the venue throughout.
  • Walk up to it, or have your SPOC post in chat and an organizer will come to your table.
  • Urgent issues start with "Blocker".
13

Where the main focus should be

If you read only one section, read this one. Priority order matters.

1
Get one control end-to-end before touching a second
Axis 1 is a gate, not a bonus. A team with one control fully working beats a team with three half-built. Finish it: threshold dated → control emitting → control test returning a verdict → passing + breaching runs recorded.
2
Date and freeze your threshold before the assessed run
A threshold dated after the test loses marks — full stop. Write metric, observation limit, exception tolerance, measurement basis, owner, version and date down early.
3
Do not move the threshold between PASS and BREACH
ORGANIZER Same frozen governing policy for both runs. Changing threshold, baseline, cap or tolerance between them manufactures the result and is a documented deduction.
4
Rehearse the two-minute judge query against both runs
The live test is Axis 2's first dimension. Practise twice on Saturday evening with someone not on the team.
5
Respect C7's multi-point scope and C16's over-cap policy field
ORGANIZER C7 must work at every enforcement point at once. C16 declares whether any over-cap run is permitted at all — as a policy field, not example text.
6
Keep the register out of the live path; keep the evaluator stateless
The architecture dimension. A control that stores its own threshold is a documented deduction.
7
Preserve the exception, not just the pass
Both worked examples keep the exception visible: the 8 ms gap in C7, the 12% drop in C9. Show the measurement, the count of exceptions, and the threshold version behind it.
8
Write the gap list honestly
"Built properly: yes. Actually worked: no" with evidence beats a wall of green ticks. This is a scored artefact.
9
Only chase bonus after the core is complete
Never bonus-chase while a core control is incomplete or non-judgeable.
10
Do not spend time on out-of-scope items
Signatures, hash chains, Merkle proofs, and a slide deck are explicitly not scored.
🧭

Order of operations on Day 1: nominate the SPOC → verify Teams room and workshop access → choose the first control → draft its dated threshold specification → assign ownership → draw the one-page architecture on a whiteboard and show a mentor in the first hour.

14

Validation notes — corrections applied in v4.0

What changed from v3.0 and why. The v2.0 predecessor should be archived.

ItemCorrectionSource
Six required steps ADDED Promoted to a first-class section (§0) — exact organizer sequence now explicit. ORGANIZER
C7 scope wording FIXED "every meaningful action" → "every in-scope event" across all three zones. Organizer wording restored. ORGANIZER
C7 enforcement-point requirement ADDED Callout in §3: C7 must work at every enforcement point at once. ORGANIZER
C16 worked-example shape ADDED Dedicated section (§4) with DECLARE → ENFORCE → RECORD → TEST → SHOW. ORGANIZER
C16 over-cap allowance ADDED Now a policy field, not example text. ORGANIZER
Same threshold rule ADDED Callout in §3: do not move the threshold between PASS and BREACH. ORGANIZER
ALLOWED block ADDED §5: agent changes, NVIDIA Nemotron, adversarial testing explicitly permitted. ORGANIZER
C9 alert semantics FIXED Now: human response not required, but alert + named recipient + silence/response still recorded. ORGANIZER
"Full set" wording FIXED 3/3 is the full set. Continuous/adversarial/second EP are separate Axis 3 bonus. ORGANIZER
C16 dollars statement SOFTENED "Token-based per the supplied control and worked example; slides do not define dollar cost as the governing measure." ORGANIZER
agentgateway / Cedar / Bedrock blocked LABELED Moved to ONSITE VERIFY — not established by the nine screenshots. ONSITE VERIFY
MCP = second enforcement point CORRECTED MCP enables workflow, does not itself prove the bonus. Bonus requires the same control architecture at an actual second enforcement point. HAIEC CAP
Five evidence planes on required sketch DEMOTED Moved to optional deeper explanation. Judge's required sketch is components + execution path + enforcement invocation. HAIEC CAP
Signing / hash-chain emphasis REMOVED From competition strategy. Digests fine for provenance; cryptographic sealing not scored. ORGANIZER
HAIEC capability truth UPDATED All three evaluator cores now AVAILABLE on Main; "Organizer-complete" tracked separately as closure gaps. OUR STACK
LogSense fixture count FIXED 12 → 13 bundled competition fixtures (2 C7 + 6 C9 + 5 C16). OUR STACK
LogSense mutex/build warning REMOVED From permanent guide. Revalidate the distributed EXE separately as an operational action item. OUR STACK
Source separation ADDED Every claim now tagged: ORGANIZER / OUR STACK / HAIEC CAP / ONSITE VERIFY. META
✓

Overall validation: v4.0 separates organizer facts from our stack and HAIEC capabilities. The six required steps are first-class. The C16 worked-example shape is preserved. The C7 multi-point rule, the same-threshold rule, and the ALLOWED block are explicit. The HAIEC capability truth table reflects current repo state. Cryptographic framing is de-emphasized per organizer.

⚠

Operational action items remaining: (1) Clarify Day 4 overlap with organizers. (2) Revalidate the distributed LogSense EXE separately from the repo — the mutex fix status is a packaging question, not a repo question. (3) Confirm modaas-agw, Cedar policy, and blocked direct paths with the venue or an official environment brief before treating them as fact.

15

Appendix — LogSense pre-event sanity check

OUR STACK Tooling and the pre-event verification.

🖥LogSense.exe

Windowed launcher → full production Streamlit workbench. Drives the 12-step stepper (Setup → Case → Discover → Evidence → Source Setup → Analyze → Runs → Summary → Ask → Controls → Gaps → Export).

⌨logsense-cli.exe

Console access. verify runs the bundled C7/C16 fixtures through the real IntegrationService.

📦13 bundled fixtures

C7: healthy, breach · C9: healthy, incompatible, latency, official, partial, zero-baseline · C16: pass, denied, pending, evidence-failure, retry-storm.

🔒Data boundary

User data writes to %LOCALAPPDATA%\LogSense\{workspace,config,logs,cache}. Nothing writes to the install directory.

pre-event sanity check
LogSense.bat verify
# or:
logsense-cli.exe verify

# Expected: "ok": true, six fixture checks,
# all MEASURED / PARTIAL as designed,
# verdict: null, verdictOwner: "HAIEC" throughout.
🧩

Where LogSense fits: OUR STACK it produces the deterministic measurements and the judge-ready evidence bundle (TMF-JUDGE-EVIDENCE-PACK.zip). The verdict is never LogSense's — every measurement carries verdict: null and verdictOwner: HAIEC.

TM Forum Trustworthy AI & Data Hackathon · Event Field Guide v4.0 · Source-separated, corrected
Organizer facts drawn from the nine official screenshots. Our stack from LogSense/HAIEC repo truth. HAIEC platform capabilities labeled as such.
verdictOwner: HAIEC · evidence beats green ticks · v2.0 predecessor archived
AI Advisor →