L1 was the school. (one of several apprenticeship pathways — the one most exposed to automation today). Autonomous SOC may be removing the layer that built the analysts who validate it. Can structured AI disagreement rebuild that learning signal — without the alert queue?
· live aggregate →Autonomous SOC platforms reduce alert volume, improve response speed, and cut operational cost. They may also reduce the analyst exposure that historically produced future L2, L3, and SOC-leader talent. This study investigates whether structured AI disagreement can recreate part of that learning experience — so organisations can adopt automation and still develop the experts they will need.
The strategic question isn’t whether AI can triage alerts. It’s the one your board will ask in two years:
This research exists because the apprenticeship pathway that answered those questions is being automated away. RTAC is one testable mechanism for rebuilding it.
As autonomous SOC platforms absorb more Level-1 work, what becomes the new pathway for developing Level-2 and Level-3 analysts?
If you’re asking how many analysts you still need, what juniors should learn, and how to train them when AI does first-level triage — this study is directly relevant. Take it, share it, critique it.
Traditional analysts learned by working through thousands of alerts. Future analysts may learn differently. This study explores one possible pathway — reasoning through disagreement, not just consuming answers — and measures whether it works. Ten minutes; runs in your browser.
Security operations may have long run a quiet apprenticeship. L1 alert triage was repetitive — but it exposed analysts to thousands of attack patterns across hundreds of hours. The hypothesis this research starts from is that such exposure helped build investigative instinct, and that this is part of how L2 and L3 analysts were trained. On this view, the seat was also a school — a claim this program sets out to investigate, not one it assumes.
AI now automates L1 triage. The learning pathway is disappearing faster than anyone is replacing it. The operational gain shows up now in cost and latency. The structural loss shows up in five years — a hollowed-out L2/L3 bench, weaker pattern recognition under pressure, and reduced ability to validate AI decisions or investigate novel threats.
The question: can structured AI disagreement — the contradiction between independent reasoning agents that an analyst must resolve — reproduce the cognitive challenge L1 used to provide, without the alert queue?
AI can do the L1 work. It cannot do the learning the L1 work used to produce.The thesis, in one line — and the thing this study sets out to measure.
Say it plainly: the goal is to help someone who never got the L1 seat reach an L2 or security-monitoring role. On our hypothesis, L1 used to be the bridge from no-experience to L2, and AI is removing that bridge. Triage School is an attempt to rebuild it — and, first, to measure whether that is even possible. Whether it works is the open question, not the premise.
Practice the technique-discrimination the vanished L1 queue used to build — and get your own measured baseline + a record of the cases you reasoned through. The on-ramp when there's no entry seat.
An independent, open-data research exercise anyone can take in the browser in ten minutes. The anonymised aggregate is public — take it to practise analyst reasoning, or replicate it to challenge the method. Open science, not a course.
As L1 automates, new hires may land in L2 seats without the exposure that—on our hypothesis—once built judgement. This measures one narrow proxy for that reasoning — to help make a possible workforce gap discussable, not to prove one exists. Not a hiring, screening, or ranking filter (see the protocol).
The bigger problem — workforce continuity if autonomous SOC removes the L1 learning seat — is not RTAC's to solve, and is itself a hypothesis rather than a settled fact. RTAC is the specific testable mechanism this study runs, to see whether structured AI disagreement reproduces the cognitive challenge the queue used to provide. The headline is the proposed workforce risk; RTAC is the proposed lever.
Three independent agents reason over the same evidence and reach different conclusions. You see the disagreement, the reasoning behind it, and resolve it. That act of reasoning — not the answer — is the learning signal.
In the vocabulary of human-in-the-loop: the agents do the reasoning, the human resolves the contradiction. This research asks the cognitive precondition for HITL — can the human still reason after L1 is automated — not the accountability or oversight senses of the term.
A few events, no AI. Your branch-navigation accuracy, recorded.
Cases where the agents disagree. You decide before seeing them — then the reasoning is revealed.
No AI. Did the discrimination transfer?
Pre/post in your browser. Nothing transmitted unless you choose.
The mechanism under test (RTAC) is simple to take and precise to measure. The score never comes from the AI-assisted step.
You triage a baseline set with no help. This is your un-assisted instinct, fixed as the starting point.
On the same alert, three independent AI reasoners reach different verdicts. You're shown the disagreement, not an answer.
You decide who's right, and why. Adjudicating conflicting expert reasoning is the cognitive work L1 triage used to build.
A fresh held-out set, no agents. The difference between your cold and your post score is the learning signal — and the reasoning is revealed afterward.
Your cold instinct, on the record — before any assistance.
You'll triage a baseline, work through cases where three agents disagree, then triage a fresh set. At the end you see your own pre/post delta. We don't collect it — the result is yours, computed here, and you may optionally copy an anonymised summary to contribute to the public aggregate.
You don't need experience to take this. In plain terms:
A few terms you'll see: L1 / L2 = junior / more-senior analyst. Triage = the first "is this real?" check. RTAC = our name for the "three AIs disagree, you decide" exercise. That's all you need to start.
A 12-case interactive slice of the study corpus, for accessibility. The full protocol (74 events) is identical in structure. Eligible for the learning aggregate: ≤ 24 months L1 exposure. Senior analysts, SOC leaders, and CISOs are welcome — with > 24 months you become the expert benchmark (your cold pre-test score is the corpus's face-validity ceiling), reported separately rather than in the learning aggregate.
Phase 1 measures one narrow thing — does branch-navigation accuracy improve after structured AI contradiction — and makes no claim about SOC performance, durability, or that the effect is caused by contradiction. Those are Phase 2. Branch accuracy is a proxy for one skill, not "analyst expertise."
Two different finish lines. Research succeeds when it produces honest knowledge — including a clean null. An operation succeeds when its metric moves (faster triage, lower MTTR). Phase 1 is running toward the first. Judging a feasibility pilot by operational rulers will (correctly) find it inadequate — the rulers are mismatched. Phase 1 de-risks the operational study, it does not pre-empt it.
"Forged" is a measurement, not a metaphor. RQ0 needs an operational definition of forging or the comparison isn't real. The protocol defines a Forging-Exposure Index (FEI, 0–18) — coarse-banded attestations across six inputs: months in triage, alerts/month, IR-bridge attendance, escalation reps, hands-on investigations, and mentorship hours. Forged = FEI ≥ 12 with breadth across ≥ 4 inputs; unforged = FEI < 6 or zero months in triage. Cut-points are declared a priori. Study A runs in two tiers: Tier 1 measures the FEI×reasoning gradient on this platform itself from self-reported FEI (executable now, no partner — exploratory/correlational); Tier 2 is the attested-cohort confirmatory upgrade, run only if Tier 1 shows a signal. See methodology.md §3b–3c.
The seed run confirmed the instrument works end-to-end. One measurable playbook correction was observed. This is not a study result — it is evidence the mechanism produces the expected output and is ready for external validation.
A feasibility pilot — it reports at n=5 eligible (a publication trigger, not a powered sample). The headline is the between-arm difference (RTAC vs control), which is the only figure that separates the method from a practice effect.
No causal figure yet.The between-arm Δ — the only figure that separates RTAC from a practice effect — needs ≥10 participants per arm before it reports.
| Metric | Status | Notes |
|---|---|---|
| Participants (eligible) | 0 / 5 | Feasibility trigger, not a powered n. |
| Between-arm Δ (RTAC − control) | — | The causal estimate. Nets out practice. Needs ≥10 per arm; reported with a 95% CI. |
| Single-arm Δ (pre→post, pooled) | — | RTAC + practice combined — not causal; shown for context. |
| Share of participants who improved | — | % with a positive delta. |
| Expert benchmark (>24mo cold pre-test) | — | Seniors/CISOs are the face-validity ceiling — benchmarked, not aggregated into the learning gain. |
| Override accuracy / consensus-zone (RQ2/3) | Predicted | Exploratory; populate with more data. |
The aggregate is reproducible from the public, anonymised submissions; contribution carries deltas + arm only — no answers, no identifiers. Sample composition (sector / exposure / region / arm) is reported so skew is visible: representativeness is measured here, not manufactured — recruited cohorts in a powered Phase-2 study would address it (conditional on Phase-1 results).
A short, plain summary of how the study works. The full protocol (corpus, scoring rules, statistical plan, every threat we've named) lives in methodology.md — written for reviewers and replicators, and intentionally detailed. This summary is for everyone else.
For reviewers and replicators — the full, rigorous version (design, corpus, scoring, statistical plan, threats, ethics, amendments): methodology.md (v2.22). The summary above is plainer on purpose.
You already know the thesis. What this gives you is concrete: taking it tells an analyst — privately — how their cold triage reasoning compares to the senior bar, and which signals they're missing. Across a team, that's a training-needs map on a risk that's otherwise invisible until a novel attack needs a human who understands attacks.
The non-evaluative guarantee, exactly what's collected, and how to self-host for a private, de-identified team diagnostic. Written so you can hand it to HR and your analysts as-is.
Read the employer brief →A one-page guide for faculty and program leads: how to run it in a lab or at home, what your students get, and how it supports graduate employability.
Read the educator brief →Do I have to install or download anything?
No. The whole experiment runs in this page, client-side. No account, no package, no infrastructure.
What happens to my answers?
They stay in your browser. Your pre/post delta is computed locally. Nothing is transmitted unless you explicitly copy and submit the anonymised summary.
Is this a product?
No. RTAC is a training instrument and Triage School is the experiment using it. It is not sold; the aggregate is public.
Are you affiliated with an AI-SOC vendor?
No. The instrument is source-agnostic by design — the disagreement signal is the object of study, not any particular producer of it.
Why fix a feasibility line up front?
The question is interesting only if the answer can be no. Committing the failure condition (mean Δ < 5pp = no signal) before data removes the temptation to find a positive result post-hoc. This is a protocol; OSF archival is intended before external data collection.
Are senior analysts excluded?
No — they're welcomed, just measured differently. With > 24 months L1 exposure you're the expert benchmark: your cold pre-test score is the corpus's face-validity ceiling (if the cases are sound, you should score high cold). That's reported separately from the learning aggregate, where seniors' near-ceiling scores would otherwise dilute the signal.
If RTAC works, how do you know it was the disagreement?
You don't — not from Phase 1, and we don't claim it. That's the central Phase-2 question (RQ1M): a 3-arm design separates contradiction from merely seeing more explanations or being more engaged. A positive RTAC-vs-control result is necessary but not sufficient.
What if the result is null?
It publishes. The protocol commits both directions, and the consensus-zone blind spot already anticipates one form of it.
I'm a CISO / SOC head — what's the business case?
Honestly: there isn't one yet, and the protocol says so. Phase 1 answers "does the mechanism produce learning?" — not ROI, throughput, or risk reduction. The operational metrics a leader weighs (transfer to live triage, time-to-decision, escalation quality, calibrated confidence, and whether it beats your existing onboarding/QA training, at an acceptable analyst-time cost) are what a powered Phase-2/3 study would measure (conditional on the Phase-1 signal warranting it). The strategic risk it addresses is workforce continuity as L1 automates. See methodology.md §12a.
Could this become a hiring or screening filter?
No — that's a hard line. If validated, the path is training/enablement for developing analysts, with guardrails. It is never designed or validated as a hiring, ranking, promotion, or termination instrument. Using a learning instrument to gate employment destroys the honesty that makes people engage with it.