✦ A New Bias · Independent AI fairness audit
Same facts. Different person.
Does the answer change?
Independent counterfactual auditing of the AI systems that make consequential decisions about people — across criminal justice, healthcare, employment, lending, housing, and education. We hold every fact about a person constant and change a single attribute — race, gender, age, class, disability, religion, or origin — so any shift in the outcome is attributable to that attribute alone. One method, every domain where an AI's judgment carries weight.
6+
domains in scope
12
demographic axes
2
evidence layers
100%
reproducible
Scope of the method, not a tally of completed audits — coverage is expanding domain by domain. See the roadmap below for what’s live now and what’s next.
Start with the reports
Four ways into the evidence — from the headline disparities to a single model’s own words.
Technical analysis
Findings
Every substantive disparity with its effect size, confidence, and cross-model replication. The reproducible core of the audit.
One domain at a time
Scenario report
The full experiment for a single scenario — design, the verbatim prompt, the groups compared, and the headline result.
The raw evidence
Responses
Filter the response corpus and open any individual answer in full — prompt, response, and judge reasoning.
Editorial deep-dive
Case study
A narrative end-to-end report on one domain — what changed when only the person did, in the models' own words.
Open any study from the Studies page for the published reports.
How the audit works
Hold everything constant
We write one realistic scenario — a loan application, a résumé, a patient chart — and fix every fact about the subject.
Change one detail
We generate demographic variants that differ by a single attribute: race, gender, age, class, disability, or origin.
Measure the divergence
Each model answers every variant. Any shift in the decision or its framing is attributable to that one detail alone.
Domains we audit
The same counterfactual method applies anywhere an AI weighs a person’s fate. We are building coverage across every domain where that judgment has real stakes.
Criminal justice
Sentencing · bail · parole · police review
Does the model hand down a harsher outcome when only the defendant's identity changes?
Healthcare
Pain management · diagnosis · mental health · child welfare
Are symptoms taken as seriously, and care offered as readily, across every patient identity?
Employment
Hiring · salary · promotion
Do identical résumés and identical performance draw different offers and advancement?
Finance
Lending & credit · insurance · financial advice
Does an identical risk profile get priced — or counselled — differently by group?
Education
School discipline · academic placement · grading
Is the same work, or the same incident, judged differently depending on the student?
Housing & civic
Housing · media coverage · customer service
Do everyday gatekeeping decisions shift with a name, an accent, or a background?
Roadmap
The work is sequenced from breadth of coverage toward statistical permanence — and toward making every figure independently reproducible.
Now · live
Youth career guidance
The first domain is live: a 16-year-old asks for guidance, and only their demographic label changes. Next we stand up the full catalog of decision domains.
Next
Repeat-sample to significance
Re-run every scenario until the leads cross the evidence bar — turning single-run signals into findings with confidence intervals.
Next
Widen the model panel
Audit every frontier model side by side, so cross-model replication — not any one system — carries the conclusion.
Later
Longitudinal & public methodology
Track each model version over time and publish the full method and raw responses, so anyone can re-run the audit.