PolicyBench has a new #1: GPT-6 Sol gets 94.2% of tax and benefit answers within $1 with no tools, ending GPT-5.6 Sol's run at the top since July. Among other models release yesterday, Claude Opus 5.5 enters second at 92.9%, GPT-6 Luna fourth at 91.4%.
policybench.org/notes/2026-0…
This morning, the Census announced that the SPM child poverty rate was 13.4% in 2025, down 0.1 pp from 2024.
Today we're launching the Child Poverty Impact Dashboard: a tool to analyze how reforms can affect this rate across all 50 states and DC.
policyengine.org/us/child-po…
Configure a reform: a state Child Tax Credit, EITC, dependent
exemption, child allowance, and more to see its impact on child poverty, individual households, income distribution, and congressional districts.
BLS re-estimates SPM poverty thresholds every year from consumer spending, not inflation. Renter threshold for two adults and two children: +6.3% from 2024 to 2025, more than twice CPI. PolicyEngine now builds and projects those thresholds the way BLS does, from the Consumer Expenditure microdata with BLS's procedure. A calculator and a paper:
Census releases the 2025 poverty numbers today at 10am ET. At 2pm ET we go through them live with guests @BlahaOsborne and @JoshuaTMcCabe and demo the tools: "The new child poverty numbers, and how state reforms would affect them." Register: us06web.zoom.us/meeting/regi…
PolicyBench tests how accurately AI models compute US taxes and benefits. GPT-6 Astra debuts second of 39 at 88.0% of answers within $1, 1.2 points behind GPT-5.6 Sol.
We've also added new Gemini, DeepSeek, and GLM models.
Most of the gap between GPT-6 Astra and GPT-5.6 Sol is two rules Astra invented regarding payroll taxes and Medicare eligibility for people with disabilities. policybench.org/notes/2026-0…
New poverty numbers land the morning of Tuesday, September 15. That afternoon at 2pm ET, join us for guest remarks on the release and a live demo of PolicyEngine's new Child Poverty Impact Dashboard.
Register: us06web.zoom.us/meeting/regi…
What would a state Child Tax Credit, an expanded EITC, or a child allowance do to child poverty in your state? The dashboard answers in minutes, for all 50 states and DC, down to the congressional district.
PolicyBench tests how accurately AI models compute US taxes and benefits. Claude Fable 5.1, released yesterday, debuts second of 33 at 86.3% of answers within $1, a tenth of a point past Kimi K3. GPT-5.6 Sol stays first at 88.7%.
PolicyBench tests how accurately AI models compute US taxes and benefits. Two new rows: Ox Alpha, a cloaked preview model with no maker named, debuts fourth of 32 at 84.2% of answers within $1. @grok 4.6 debuts eighth at 82.8%, 1.9 points past Grok 4.5.
policybench.org
PolicyBench tests how accurately AI models compute US taxes and benefits. New: Gemini 3.7 Flash scores 79.2% of answers within $1 — tenth of 30, just past its predecessor, at $0.014 per household.
Full board: policybench.org
PolicyBench, our benchmark of AI accuracy on US taxes and benefits, found forcing a tool call silently turns off Claude's thinking — others reason regardless. Under tool_choice auto, Claude Fable 5 jumps 79.9→86.9% within $1: would-be #2.
More: github.com/PolicyEngine/poli…