The evaluator compares measured values against a declared rule and returns a verdict.
Trustworthy AI & Data Hackathon
One operational page for the challenge model, controls, working HTML tools, shared files, screenshots, judging expectations and event logistics.
Open the working tools first
The Field Guide and Intake HTMLs are the two primary team tools. Keep the Drive folder open beside them during the event.
The challenge in one view
Build evidence-backed controls around the running system without confusing telemetry, evaluation and enforcement.
The enforcement point acts on the verdict. Keeping this separate prevents hidden policy state in the path.
Records bind the run, control, threshold version and observation so another person can verify the result.
The judge should be able to ask about one control and one run and get the verdict with the records behind it.
The three event controls
Complete one end to end at minimum. A second and third control increase coverage and bonus potential.
AIA-LOG-001Automatic event recording
Measure: Expected-event coverage, timing gaps, exception rate
Evidence: Expected IDs, events, timestamps and run links
AIA-ARC-006Drift and performance
Measure: Relative drop against a frozen baseline by comparable window
Evidence: Baseline, live values, windows, alerts, owner and response
ACN-COST-001Per-run spend cap
Measure: Input plus output tokens reconciled across agents and retries
Evidence: Call IDs, token counts, budget version and stop records
What to build and what judges will test
The highest leverage is a reproducible path from declared threshold to measured value to verdict to evidence.
Required sequence
- Choose the control or controls.
- Declare and date threshold ranges before the assessed run.
- Decide when in the timeline the control is checked.
- Build the control at a selected enforcement point.
- Build the control test and evaluator.
- Run one passing case and one breaching case, then record the gaps.
Judging focus
- Completion: how many of the three controls work end to end.
- Live test: can the team answer a named control/run question quickly.
- Architecture: can the design move to another enforcement point without rewriting the control.
- Evidence: are verdicts bound to a control, threshold version and run.
- Honesty: gaps and failures are recorded instead of converted into green checks.
Team resources
These links are intentionally large and direct. The two HTML tools and Drive folder should be the team's working tabs.
Team Field Guide
Challenge structure, controls, judging, operating model, first-hour sequence and onsite verification points.
OPEN FIELD GUIDE HTML ↗HTML · LIVE WORKSHEETEnvironment & Integration Intake
Capture environment access, gateways, telemetry, run identity, C7/C9/C16 details, judging constraints and missing artifacts.
OPEN INTAKE HTML ↗TEAM FILESShared Google Drive
Shared working folder for source material, evidence, handoff files, competition artifacts and team collaboration.
OPEN GOOGLE DRIVE ↗HTML · REPORT TEMPLATEAssurance Report Template
Illustrative AL0 to AL2 forensic assurance structure. Sample statuses are placeholders, not measured competition findings.
OPEN REPORT HTML ↗Briefing screenshots
All nine source screenshots are shown individually. Select any image to open the full-size version.
Control register, evaluator, evidence ledger, control test and the four transport patterns.
Open full image in Drive ↗Customer, IT and network zones with zone gateways and the shared model gateway.
Open full image in Drive ↗Event recording, drift and performance, and per-run spend cap.
Open full image in Drive ↗Claim, control objective, risk, control test and evidence structure.
Open full image in Drive ↗Declare, enforce, record, test and show for the shared model gateway.
Open full image in Drive ↗Illustrative measurement patterns for timing gaps and drift against a frozen baseline.
Open full image in Drive ↗The required implementation order plus the bonus paths.
Open full image in Drive ↗Completion, quality and bonus criteria, including what loses marks.
Open full image in Drive ↗Evidence file, threshold document, control test tool, named runs, architecture sketch and gap list.
Open full image in Drive ↗Event logistics that affect execution
Only the items that can affect team availability, support or judging are repeated here.
AT&T Headquarters, 208 South Akard, Dallas. Check in at Building 1. Hackathon space is Building 3, 12th-floor auditorium.
Marriott Dallas Allen Hotel & Convention Center, 777 Watters Creek Blvd, Allen. Hackathon room: Starlight Ballroom 3.
Preliminary judging is Tuesday, with finalists announced afterward and finalist judging later that afternoon. Awards are Wednesday.
Microsoft Teams is the official remote support channel. Each team has a room. The designated SPOC posts in the main chat, with urgent issues prefixed “Blocker”.
Fast path during the event