Status
Phase 1 · collecting
Taken so far
runs
Format
in-browser · ~10 min
Your data
never transmitted
Reports at
n ≥ 5 eligible (feasibility)
Feasibility line
mean Δ < 5pp = no signal
Affiliation
independent · no vendor
Triage School Independent research on the queue that made experts.
SOC Workforce Research · Phase 1 Mechanism Study · 2026 · Open

AI can automate L1 alerts. But those alerts were how analysts learned to become L2 and L3.

L1 was the school. (one of several apprenticeship pathways — the one most exposed to automation today). Autonomous SOC may be removing the layer that built the analysts who validate it. Can structured AI disagreement rebuild that learning signal — without the alert queue?

Collecting now — be among the first. · live aggregate →
Take the experiment — 10 min ↓ In your browser, nothing transmitted. Independent research — no vendor, no product.
The thesis, in two pictures PROBLEM The apprenticeship thesis — then the breakdown. THEN L3 L2 L1 alerts NOW L3 L2 AUTO · SOC alerts The seat that taught the next seat — automated away. THIS STUDY · RTAC Structured disagreement, then you. Alert ? three reasoners reach different conclusions You resolve. . Adjudicating contradiction is the work L1 used to teach. IF IT WORKS · PHASE 2+ · CONDITIONAL Automation kept — pipeline restored. IF L3 L2 TRAINING structured disagreement + AUTOMATION VALIDATE AUTO · SOC alerts Automation does L1. Humans keep the validating expertise.
Figure 1 · The thesis in three pictures (a hypothesis, not a finding). Top: the L1 learning seat the thesis describes, automated away. Middle: the mechanism this study tests — structured AI disagreement, resolved by a human. Bottom (conditional, dashed border): what success could look like at scale — only if the mechanism is shown to work in a powered Phase-2 study.
~10 min
In your browser. Nothing transmitted unless you choose.
2 arms
RTAC vs control — the only figure that separates method from practice.
0 identifiers
No name, email, IP, or cookies — ever. Published either way.

Why this matters for autonomous SOC

Autonomous SOC may be removing the layer that built the analysts who validate it.

Autonomous SOC platforms reduce alert volume, improve response speed, and cut operational cost. They may also reduce the analyst exposure that historically produced future L2, L3, and SOC-leader talent. This study investigates whether structured AI disagreement can recreate part of that learning experience — so organisations can adopt automation and still develop the experts they will need.

The logic in one line: Autonomous SOC → less L1 exposure → potential expertise gap → can structured AI disagreement rebuild the missing learning signal? → RTAC study.
For CISOs beginning an autonomous SOC journey

If automation succeeds, who fills the next seat?

The strategic question isn’t whether AI can triage alerts. It’s the one your board will ask in two years:

  • Who becomes my future L2?
  • Who becomes my future L3?
  • Who validates the AI’s decisions when they’re wrong?
  • Who investigates novel attacks the swarm hasn’t seen?
  • Who becomes my next SOC manager?

This research exists because the apprenticeship pathway that answered those questions is being automated away. RTAC is one testable mechanism for rebuilding it.

For SOC leaders planning workforce transformation

The key research question

As autonomous SOC platforms absorb more Level-1 work, what becomes the new pathway for developing Level-2 and Level-3 analysts?

If you’re asking how many analysts you still need, what juniors should learn, and how to train them when AI does first-level triage — this study is directly relevant. Take it, share it, critique it.

For freshers & career-changers

Can this help me become a better analyst?

Traditional analysts learned by working through thousands of alerts. Future analysts may learn differently. This study explores one possible pathway — reasoning through disagreement, not just consuming answers — and measures whether it works. Ten minutes; runs in your browser.

The structural risk

The seat that taught the next seat is being automated away.

Security operations may have long run a quiet apprenticeship. L1 alert triage was repetitive — but it exposed analysts to thousands of attack patterns across hundreds of hours. The hypothesis this research starts from is that such exposure helped build investigative instinct, and that this is part of how L2 and L3 analysts were trained. On this view, the seat was also a school — a claim this program sets out to investigate, not one it assumes.

AI now automates L1 triage. The learning pathway is disappearing faster than anyone is replacing it. The operational gain shows up now in cost and latency. The structural loss shows up in five years — a hollowed-out L2/L3 bench, weaker pattern recognition under pressure, and reduced ability to validate AI decisions or investigate novel threats.

The question: can structured AI disagreement — the contradiction between independent reasoning agents that an analyst must resolve — reproduce the cognitive challenge L1 used to provide, without the alert queue?

If this pattern holds, it may be broader than SOC — AI could be removing apprenticeship layers across professions. SOC is simply where we think it can be measured first. That broader claim is a motivating hypothesis, not a finding.
AI can do the L1 work. It cannot do the learning the L1 work used to produce. The thesis, in one line — and the thing this study sets out to measure.

The intention

Testing whether we can rebuild the entry path into the job.

Say it plainly: the goal is to help someone who never got the L1 seat reach an L2 or security-monitoring role. On our hypothesis, L1 used to be the bridge from no-experience to L2, and AI is removing that bridge. Triage School is an attempt to rebuild it — and, first, to measure whether that is even possible. Whether it works is the open question, not the premise.

New to the field? Here's what's in it for you. The honest version of the question you're really asking is “am I any good at this yet?” — and you'll find out: you'll see how your cold reasoning compares to senior analysts on the same cases, and exactly which signals you missed. “I have no idea” is who this is built for — that's the measurement, not a disqualification. There's nothing to study and no way to fail. It's a mirror, not a certificate. It won't tell you “you're hired as L2” — it shows you, privately, where your instinct already matches a senior's and where the gap is.
Early-career & career-changers

Build the reasoning without the seat

Practice the technique-discrimination the vanished L1 queue used to build — and get your own measured baseline + a record of the cases you reasoned through. The on-ramp when there's no entry seat.

Open to learners — and to scrutiny

Join it, or pull it apart

An independent, open-data research exercise anyone can take in the browser in ten minutes. The anonymised aggregate is public — take it to practise analyst reasoning, or replicate it to challenge the method. Open science, not a course.

Security leaders & workforce planners

See how your new bench reasons

As L1 automates, new hires may land in L2 seats without the exposure that—on our hypothesis—once built judgement. This measures one narrow proxy for that reasoning — to help make a possible workforce gap discussable, not to prove one exists. Not a hiring, screening, or ranking filter (see the protocol).

Honest about what it is today. Triage School is a research instrument and a practice + measurement tool. It is not (yet) a certification, and not a job guarantee — its value as an employability signal is contingent on the Phase-1 result (does structured AI disagreement actually improve reasoning?). We publish that result either way. We won't sell a bridge we haven't shown holds weight.

The mechanism we test

One workforce question, one mechanism: RTAC — Reasoning Through Agent Contradiction.

The bigger problem — workforce continuity if autonomous SOC removes the L1 learning seat — is not RTAC's to solve, and is itself a hypothesis rather than a settled fact. RTAC is the specific testable mechanism this study runs, to see whether structured AI disagreement reproduces the cognitive challenge the queue used to provide. The headline is the proposed workforce risk; RTAC is the proposed lever.

Three independent agents reason over the same evidence and reach different conclusions. You see the disagreement, the reasoning behind it, and resolve it. That act of reasoning — not the answer — is the learning signal.

In the vocabulary of human-in-the-loop: the agents do the reasoning, the human resolves the contradiction. This research asks the cognitive precondition for HITL — can the human still reason after L1 is automated — not the accountability or oversight senses of the term.

Step 1 · Pre-test

Baseline

A few events, no AI. Your branch-navigation accuracy, recorded.

Step 2 · RTAC session

Reason through

Cases where the agents disagree. You decide before seeing them — then the reasoning is revealed.

Step 3 · Post-test

Same skill, new events

No AI. Did the discrimination transfer?

Step 4 · Result

Your delta, live

Pre/post in your browser. Nothing transmitted unless you choose.

What RTAC is — and is not

How it works

Structured disagreement, in four moves.

The mechanism under test (RTAC) is simple to take and precise to measure. The score never comes from the AI-assisted step.

Move 01

You judge cold.

You triage a baseline set with no help. This is your un-assisted instinct, fixed as the starting point.

Move 02

The agents disagree.

On the same alert, three independent AI reasoners reach different verdicts. You're shown the disagreement, not an answer.

Move 03

You resolve the contradiction.

You decide who's right, and why. Adjudicating conflicting expert reasoning is the cognitive work L1 triage used to build.

Move 04

You're measured un-assisted.

A fresh held-out set, no agents. The difference between your cold and your post score is the learning signal — and the reasoning is revealed afterward.

The RTAC loop Alert ? they disagree You resolve Held-out check

Your cold instinct, on the record — before any assistance.

Take the experiment

Ten minutes. In your browser. Nothing leaves this page.

You'll triage a baseline, work through cases where three agents disagree, then triage a fresh set. At the end you see your own pre/post delta. We don't collect it — the result is yours, computed here, and you may optionally copy an anonymised summary to contribute to the public aggregate.

New to security? Read this first — in plain words

You don't need experience to take this. In plain terms:

  • A security alert pops up. Is it a real attack, or harmless? That judgement call is what analysts do all day.
  • You'll make a series of these calls. First on your own, then on a few cases where three AI assistants disagree with each other — you decide who's right, then see the reasoning.
  • Then a fresh set, to see if the reasoning stuck. Guessing is fine — the point is learning how experienced analysts think, which you'll see explained for every answer.

A few terms you'll see: L1 / L2 = junior / more-senior analyst. Triage = the first "is this real?" check. RTAC = our name for the "three AIs disagree, you decide" exercise. That's all you need to start.

Ready

A 12-case interactive slice of the study corpus, for accessibility. The full protocol (74 events) is identical in structure. Eligible for the learning aggregate: ≤ 24 months L1 exposure. Senior analysts, SOC leaders, and CISOs are welcome — with > 24 months you become the expert benchmark (your cold pre-test score is the corpus's face-validity ceiling), reported separately rather than in the learning aggregate.

The study

Two studies: does the gap exist — and can it be rebuilt?

RQ0 · the gap
The thesis (a hypothesis to test). Do analysts who came up through the L1 learning seat reason better than those who entered after it was automated, at comparable tenure? "L1 was required for a reason," made testable. May come back null — published either way. Everything below only matters if this gap is real.
RQ1F · now
Feasibility (not causal). Can a single RTAC session move playbook branch-navigation accuracy on a held-out post-test — enough to justify a powered Phase 2? Phase 1 estimates; it does not confirm causation.
RQ1 · Phase 2
Causal. Does RTAC improve accuracy net of practice? Answered by the between-arm difference (RTAC vs an active no-disagreement control), at a power-calculated n.
RQ1M · Phase 2
The central question. If RTAC works, is the learning caused by the contradiction specifically, or just by more explanation / engagement? Isolated with a 3-arm design (RTAC vs multiple-explanation vs single-explanation). A positive result is necessary but not sufficient for a "contradiction" claim.
RQ2 · exploratory
For which playbooks can the agent swarm act reliably alone? Per-playbook swarm accuracy × post-RTAC analyst accuracy.
Reports at
n ≥ 5 eligible — a feasibility-publication trigger, not a powered sample. Mean Δ < 5pp = no feasibility signal, published either way. (An illustrative working guess only — not a target, not a prediction the pilot is powered to test, and not a claim: e.g. ~55% → ~90% on this narrow proxy. On a 4-item slice a single item is 25pp, so treat any such figure as noise, not effect size.)
Eligibility
Learning aggregate = ≤ 24 months L1 exposure (target: < 6 months). > 24 months = the expert benchmark (face-validity ceiling), reported separately — welcomed, not excluded.
Phases
1 feasibility (active). Any causal or efficacy claim is reserved for a powered, controlled Phase-2 study, conditional on Phase-1 results: verified cohorts, mechanism isolation, retention at 30/90/180 days, powered n. Phase 3 would test operational transfer.

Phase 1 measures one narrow thing — does branch-navigation accuracy improve after structured AI contradiction — and makes no claim about SOC performance, durability, or that the effect is caused by contradiction. Those are Phase 2. Branch accuracy is a proxy for one skill, not "analyst expertise."

Two different finish lines. Research succeeds when it produces honest knowledge — including a clean null. An operation succeeds when its metric moves (faster triage, lower MTTR). Phase 1 is running toward the first. Judging a feasibility pilot by operational rulers will (correctly) find it inadequate — the rulers are mismatched. Phase 1 de-risks the operational study, it does not pre-empt it.

"Forged" is a measurement, not a metaphor. RQ0 needs an operational definition of forging or the comparison isn't real. The protocol defines a Forging-Exposure Index (FEI, 0–18) — coarse-banded attestations across six inputs: months in triage, alerts/month, IR-bridge attendance, escalation reps, hands-on investigations, and mentorship hours. Forged = FEI ≥ 12 with breadth across ≥ 4 inputs; unforged = FEI < 6 or zero months in triage. Cut-points are declared a priori. Study A runs in two tiers: Tier 1 measures the FEI×reasoning gradient on this platform itself from self-reported FEI (executable now, no partner — exploratory/correlational); Tier 2 is the attested-cohort confirmatory upgrade, run only if Tier 1 shows a signal. See methodology.md §3b–3c.

Pilot · phase 1 seed run

One real correction. Three honest blind spots.

The seed run confirmed the instrument works end-to-end. One measurable playbook correction was observed. This is not a study result — it is evidence the mechanism produces the expected output and is ready for external validation.

What RTAC taught. A legacy service account logged in interactively. The triage agent called it valid-account abuse; the identity agent called it token theft. The contradiction forced the distinction between "valid account being abused" and "token replay" — different vector, different containment. After: the distinction was permanent. Reasoned once, owned permanently.
What RTAC missed. Three cases where the agents agreed (a documented load-balancer account; encoded PowerShell from a patch tool; a confirmed DR database dump) and an analyst still went wrong. No disagreement → no learning signal. RTAC cannot teach what the swarm already agrees on — the 30% problem, quantified.

Live research dashboard

The aggregate — populating as people take it.

A feasibility pilot — it reports at n=5 eligible (a publication trigger, not a powered sample). The headline is the between-arm difference (RTAC vs control), which is the only figure that separates the method from a practice effect.

Triage School · Phase 1 · live aggregate Collecting · awaiting eligible n

No causal figure yet.The between-arm Δ — the only figure that separates RTAC from a practice effect — needs ≥10 participants per arm before it reports.

MetricStatusNotes
Participants (eligible)0 / 5Feasibility trigger, not a powered n.
Between-arm Δ (RTAC − control)The causal estimate. Nets out practice. Needs ≥10 per arm; reported with a 95% CI.
Single-arm Δ (pre→post, pooled)RTAC + practice combined — not causal; shown for context.
Share of participants who improved% with a positive delta.
Expert benchmark (>24mo cold pre-test)Seniors/CISOs are the face-validity ceiling — benchmarked, not aggregated into the learning gain.
Override accuracy / consensus-zone (RQ2/3)PredictedExploratory; populate with more data.

The aggregate is reproducible from the public, anonymised submissions; contribution carries deltas + arm only — no answers, no identifiers. Sample composition (sector / exposure / region / arm) is reported so skew is visible: representativeness is measured here, not manufactured — recruited cohorts in a powered Phase-2 study would address it (conditional on Phase-1 results).

Methodology — in plain terms

What we're doing, how we're checking it, and why a null result publishes too.

A short, plain summary of how the study works. The full protocol (corpus, scoring rules, statistical plan, every threat we've named) lives in methodology.md — written for reviewers and replicators, and intentionally detailed. This summary is for everyone else.

How the study works
Everyone takes the same three steps — baseline, middle set, fresh set — but is randomly assigned to one of two versions of the middle set. Half see AI agents disagree and resolve it themselves; half see the same cases without disagreement. Comparing the two groups is what tells us whether the disagreement is what teaches, rather than just doing the test twice.
What counts as the result
Your improvement from pre-test to post-test (in percentage points). Across many people, the difference between the two groups is the headline number — not your individual score.
What we measure (and don't)
We measure one specific reasoning skill: picking the correct response on look-alike alerts. We're honest that this is narrow — it isn't "are you a good analyst." A larger, powered Phase-2 study would widen that — if Phase 1 warrants it.
When we publish
The aggregate goes public at n ≥ 5 eligible participants. If the average improvement is under 5pp, we publish that as a null — "the signal isn't there." No quietly waiting for better numbers.
Privacy
Anonymous, in-browser, you only submit if you choose. No name, email, IP, cookies, or accounts.
What this is not
Not a certification. Not a hiring or screening tool. Not a vendor product. Independent research, full stop.

For reviewers and replicators — the full, rigorous version (design, corpus, scoring, statistical plan, threats, ethics, amendments): methodology.md (v2.22). The summary above is plainer on purpose.

For CISOs & SOC leads

Running this with your team? Lead with the diagnostic, not the thesis.

You already know the thesis. What this gives you is concrete: taking it tells an analyst — privately — how their cold triage reasoning compares to the senior bar, and which signals they're missing. Across a team, that's a training-needs map on a risk that's otherwise invisible until a novel attack needs a human who understands attacks.

CISOs, SOC leads & HR

The non-evaluative guarantee, exactly what's collected, and how to self-host for a private, de-identified team diagnostic. Written so you can hand it to HR and your analysts as-is.

Read the employer brief →

Universities, colleges & bootcamps

A one-page guide for faculty and program leads: how to run it in a lab or at home, what your students get, and how it supports graduate employability.

Read the educator brief →

Frequently asked

Honest answers.

Do I have to install or download anything?
No. The whole experiment runs in this page, client-side. No account, no package, no infrastructure.

What happens to my answers?
They stay in your browser. Your pre/post delta is computed locally. Nothing is transmitted unless you explicitly copy and submit the anonymised summary.

Is this a product?
No. RTAC is a training instrument and Triage School is the experiment using it. It is not sold; the aggregate is public.

Are you affiliated with an AI-SOC vendor?
No. The instrument is source-agnostic by design — the disagreement signal is the object of study, not any particular producer of it.

Why fix a feasibility line up front?
The question is interesting only if the answer can be no. Committing the failure condition (mean Δ < 5pp = no signal) before data removes the temptation to find a positive result post-hoc. This is a protocol; OSF archival is intended before external data collection.

Are senior analysts excluded?
No — they're welcomed, just measured differently. With > 24 months L1 exposure you're the expert benchmark: your cold pre-test score is the corpus's face-validity ceiling (if the cases are sound, you should score high cold). That's reported separately from the learning aggregate, where seniors' near-ceiling scores would otherwise dilute the signal.

If RTAC works, how do you know it was the disagreement?
You don't — not from Phase 1, and we don't claim it. That's the central Phase-2 question (RQ1M): a 3-arm design separates contradiction from merely seeing more explanations or being more engaged. A positive RTAC-vs-control result is necessary but not sufficient.

What if the result is null?
It publishes. The protocol commits both directions, and the consensus-zone blind spot already anticipates one form of it.

I'm a CISO / SOC head — what's the business case?
Honestly: there isn't one yet, and the protocol says so. Phase 1 answers "does the mechanism produce learning?" — not ROI, throughput, or risk reduction. The operational metrics a leader weighs (transfer to live triage, time-to-decision, escalation quality, calibrated confidence, and whether it beats your existing onboarding/QA training, at an acceptable analyst-time cost) are what a powered Phase-2/3 study would measure (conditional on the Phase-1 signal warranting it). The strategic risk it addresses is workforce continuity as L1 automates. See methodology.md §12a.

Could this become a hiring or screening filter?
No — that's a hard line. If validated, the path is training/enablement for developing analysts, with guardrails. It is never designed or validated as a hiring, ranking, promotion, or termination instrument. Using a learning instrument to gate employment destroys the honesty that makes people engage with it.