Case study · insurance talent · live since January 2025

Results from production. Evidence you can inspect.

One Fortune 500 insurance carrier reported $1.58M in Q1 2025 net savings, validated by its CFO. Requisition-to-hire time moved from 127 to 38 days. The deployment has scored 900,000+ candidates since January 2025.

One insurance talent deployment · Observed customer records · These findings do not establish causality or transfer without separate validation.

Requisition to hire · 2025127 → 38 daysObserved
Candidates scored since Jan 2025900,000+Observed
Deployment footprint215+ locationsObserved
Research cohort · 2022 to 202510,765 agentsObserved
The case

The filter had never been tested against production.

The carrier operated a high-volume hiring funnel across 215+ locations. The study tested familiar screening signals by joining historical hiring records to later production outcomes.

It examined whether the evidence used to assess candidates was associated with the defined production milestone, while preserving the limits of a retrospective analysis.

01 · Filter audit

The familiar signals did not hold up.

Of 8,181 parsed skills, 3,597 had enough data to test. After Bonferroni correction, zero predicted the first production milestone and 30 were anti-predictive.

02 · Ranking evidence

Connected evidence supported further testing.

The keyword analysis reported AUC 0.558. In a separate small sample of n=229, behavioral assessment alone produced cross-validated AUC 0.647 and fused ATS, assessment, and behavioral evidence produced AUC 0.735. These analyses are bounded by their samples and methods.

03 · Time

The production milestone arrived 47 days earlier.

The reference cohorts moved from 109 days to 62 days to the first production milestone. The live requisition-to-hire loop moved from 127 days to 38 days.

04 · Modeled value at risk

One rule would have rejected 2,863 producing agents.

The historical replay identified producing agents the industry-experience filter would have removed. Their annual production represented $17.7M at risk in the retrospective counterfactual. This does not represent incremental revenue caused by Nodes or realized savings. Inspect the modeled claim.

AUC describes ranking performance in the stated sample. The time and value findings come from one carrier. They do not establish universal accuracy, causality, or future lift.

Observed production record

The deployment moved from contract to production in 34 days.

Legal approval took 17 days after six AI hiring vendors had been rejected over 18 months on architecture. These are results from one deployment, not a timeline promise for another environment.

17 days · legal approval · one customer record
34 days · contract to production · one deployment
Live since January 2025
Expanded two quarters ahead of schedule
Application and research boundary

What this deployment establishes.

The production application evaluates resume and assessment evidence against configured hiring criteria. Named people retain hiring authority. Quarterly production reports and retrospective research provide outcome evidence; their existence does not establish that the application automatically learns from them.

01 · Ingest

Read approved records.

Production evidence stayed inside the approved customer VPC boundary.

02 · Evaluate

Review candidate evidence.

The application supports candidate evaluation against configured criteria. A named person makes the hiring decision.

03 · Analyze

Study the historical outcome.

The retrospective research links records to the customer's defined production milestone, with its methods and limits published separately.

04 · Distinguish

Keep the runtime claim separate.

These results do not prove autonomous workflow creation, cross-system recovery, or automated outcome learning.

Broader platform · evaluation standard

Ask to see the unfamiliar job, the exception, and the next case.

A separate capability demonstration should show relevant context and gaps, reused or newly tested capabilities, a human-refined plan, permitted execution, and recovery from an exception. Then test whether a later result changes the next applicable recommendation.

Measure engineering effort, customer intervention, verified effects, and learning against a prior or no-memory baseline. The homepage example illustrates this direction. It is not production footage or additional customer proof.

Evidence record

Observed, validated, and modeled claims stay distinct.

Each number retains its population, period, method, and limitations. The full evidence register remains available as a separate inspection surface.

IDClaimValuePopulationMethodStatus
C01Candidates scored900,000+One Fortune 500 carrierLive production record since January 2025Observed
C05Research cohort10,765 agentsAgents at one carrierRetrospective observational study, 2022 to 2025Observed
C20Q1 2025 net savings$1.58MOne live deploymentCFO-validated customer recordObserved
C21Annual production at risk$17.7M2,863 producing agentsRetrospective counterfactualModeled
C25Time to first production milestone109 to 62 daysReference carrier cohortsHistorical median comparisonObserved
C26Requisition to hire127 to 38 daysOne live deploymentCustomer production record, 2025Observed
C41Predictive keywords after correction0 of 3,597Testable keyword setBonferroni-corrected testingValidated

Methodology: Decision Traces, arXiv:2604.19819. Complete claim records and supporting artifacts are available through the evidence register and deployment data room.

Find your starting point

What should Nodes take on in your business?

Explore one workflow in a product walkthrough, with no dataset required. If you want to test an existing decision against historical outcomes, Decision Replay is a separate next step.