Technical consulting for coding, predictive analytics, and optimization. Focused on public sector applications, mostly with police departments.

AI project for the day. My work (both in criminology and with healthcare records) I have come to the rough guideline that you need around 20k observations before tree based models (random forests or boosted model variants) out-perform more typical regression for prediction problems. (That is at least my experience across a variety of mostly binary prediction problems with admin data.) There are different tabular foundation models though that are meant to be potential alternatives in the low-N scenario. Similar in spirit to K-shot prompts for LLMs, these are large models that can be trained on a different outcome with few observations. I put one of the newer tabular foundation models just released from NVIDIA, Kumo, through its paces on the NIJ recidivism dataset, and it does quite well. It would have been on the large team leaderboard in the NIJ for basically every category without feature engineering (so just piping in the data as is).
1
28
Scott wrote a very nice response letter to Benn Jordan's call critiquing the recent Flock pre-print. It is not crimrxiv's place to moderate scholarly disagreements, but to provide a platform for those disagreements to be aired in public.
2
7
733
VerusCite has its first retraction (that I am aware of)
1
1
79
Replying to @VittorioNastasi
So FYI apwheele.github.io/CrimeDeco… that shows a model of same sample and just underlying national and city trend. Agree MVT shows some of the most volatility, but that doesn't mean the synth result in the linked post is wrong.
1
34
New VerusCite blog post, some editors/authors have given pushback on what I flag as an AI hallucination. Often these are articles that exist, but have whole-cloth fabricated elements, like made up authors or a totally wrong periodical. I don't have a smoking gun, but lets be real why would a human type out totally wrong information?
3
2
347
I do not know about infant mortality as much, but one of the topics I familiarized myself with a few years ago was US's higher rate of maternal pregnancy morbidity. Which is almost entirely due to US having a broader definition of what counts (like homicide and drug uses that can occur months later).
Hey @grok, does the US have higher infant mortality because we hate babies? Or are there better explanations?
1
144
Replying to @LessCrime
This article has a likely AI hallucinated reference veruscite.com/share/fFEnQG92…
1
79
Hallucination of the day in Probation Journal I have made the tools (when sharing these) as transparent as possible. So the NIJ article my tool originally classifies as minor error, and the Moon article classifies as not found. Before sharing, I review them though. And IMO these are hallucinations. Specifically built the tool so it is easy to go and review the identified errors and change classification/edit components as you see fit.
1
106
Maybe it is a meta message in the paper itself -- "generating metacognitive feelings of confidence that are easily mistaken for epistemic justification".
14
Hallucination of the day in a Philosophy journal
1
79
New blog post, I have been using extra credits to just churn out ideas I have had for years. Peer reviewed papers were always symbolic. The ease you can write them now I am hoping spurs academics to consider "doing things in the real world" as a measure of success.
2
1
4
308
Hallucination of the day in BMC Research. Journals could definately afford my service -- they are charging over $3k for APC fees in this journal. Costs only $2 to review in my tool.
1
69
Most recent AI project -- ignorance bands on synthetic control results. So you have many potential solutions. What this does is generate Manski style bands on the feasible solution space (so these are not standard error estimates, but actual treatment effect estimates that are reasonable). It is not as bad as I expected actually.
1
102
Replying to @CharlesFLehman
13 fewer per 100k per month, so adds up.
1
1
55
San Fran motor vehicle thefts is a good example I like to look at the cumulative differences over time. Any individual month is within the error bar, but it is persistently under predicted. Post June 2024 when they rolled out Flock.
1
2
1,593
Four new papers added to the Hall of Hallucination today
2
140
Here for illustration if you shift the intervention a year early, counterfactual just hovers at 140 per 100k.
31
Replying to @joshmccrain
That is a reasonable interpretation. It is just 1 estimate, two different ways to estimate the standard error around that estimate. Here is the longer term trend relative to the US, apwheele.github.io/CrimeDeco… (in that same sample of cities). So ignoring the downward spike could argue the DiD by eye shows pretty similar trends to US post 2023.
1
42
One of the reasons I built the synth tool is I don't like the defaults currently in use. One of those is the use of placebo's, which have way too wide of intervals relative to other conformal approaches. Placebo for city level here needs to reduce/increase crime by nearly 50% to be able to detect an effect!
2
1
327
The most recent Retraction Watch blog post is an author cleared of wrong doing by pretty much saying these are real citations, I just messed up. These are so brazenly AI hallucinated references this defense is absurd.
2
1
190
This runs client side in pyscript WASM. For all the estimators except for placebo intervals only takes a few seconds. And has a nice print report. (Works just fine on my iphone.) Part of the reason to do this is I do have preferred estimators (in particular the cumulative estimate using Lasso with an intercept). For now just a single city, will have to think about multiple treated, multiple intervention, and including multiple outcomes in the future. (And maybe even multi-verse analysis given the different potential options.) I think it can probably all be done in a nice WASM app though.
70
I have created an app to conduct on-demand synthetic control analysis for crime trends at the city level (using Jeff Asher's RTCI data). Here for example, while it is true that violent crime was declining before the Memphis task force started. But Memphis does appear to continue on lower than expected compared to the counter-factual cities.
2
2
10
728
This is a weird Google Scholar error, not from scanning PDFs but from its recommended reading email to me this morning. This is a legit article from Center for American Progress, Scholar indexed it to a maybe porn site.
128
One of the reasons that having regulatory policy is hard is that it tends to turn into a checklist on a spreadsheet. Marina Nitze on the Eating Policy blog writes up her struggles with foster care bureaucracy. It is worth going through her paragraphs on how 6 states ended up requiring foster care parents to have recycling. I am seeing this now at the day gig with "AI regulations" by various states. It is mostly box checking for generic questions.
1
111
Hallucination of the day in the American Journal of Criminal Justice Was speaking to a professor yesterday, and she was a bit exacerbated that she needs to be more vigilant with her grad students over this. A paper having this many authors does not insulate it from these problems.
1
4
617
Durham, NC crimes (orange) compared to the typical crime decline (blue). apwheele.github.io/CrimeDeco… Violent has largely matched, property crime though going up.
2
2
149
Hallucination of the day I have not been the one submitting a paper for years now, can someone confirm that the DOI's are inserted by Springer copyeditors? (Not authors?) So when they cannot find the right DOI they just insert the next best guess (which indeed is not the correct paper).
1
124
I have created a stats page on VerusCite. My guesses of minor errors have been too low, but hallucinated citations (even acknowledging false positives) may only be 2% of all cites. At least among the selected papers scanned.
1
91
ojp.gov/pdffiles1/Photocopy/… From 1992 -- while the Glock post was about civil, No soul to damn no body to kick is a long standing problem for criminal liability (with still no real great solution).
1
1
88
Part of the difficulty in writing a book on LLM APIs is that they change so often. LLMs for Mortals is 6 months out, and I wrote a recap of updates needed so the book recompiles (mostly just newer model strings). But a few things probably by next year will need to be changed in the book. AWS uses a token now, should update with some of the newer open source models (and show calling from Databricks or Baseten US endpoints). And may need to take out sections on temperature and logprobs. See blog post for promo code for EPUB and print versions of the book.
1
1
107
Hallucination of the day is an article that has all wrong journals. So another example where current title only estimates will be too low.
1
1
255
Hotlanta assaults are rising
1
78
They are not doing a great job Thomas! veruscite.com/blog/springer-… Again happy to gift arxiv some credits to evaluate my tool.
1
8
Quick blog post based on some the recent @cremieuxrecueil posts on the prevalence of hallucinated citations (which are often around 1% of citations in recent preprint servers). I think these are likely underestimates, as those papers only check titles. I did a quick coding for some of the recent papers in my hall of hallucinations, and close to half would not have been caught just looking at the titles. Totally wrong authors are the most common example. So it might be twice as bad as people are estimating!
1
5
13
4,691
Uggh another google scholar entry for a well cited book that is not even close.
1
135
My tool identifies a few additional hallucinations based on pretty wrong author lists Can see the report at veruscite.com/share/FU2qLek4… In Crémieux's other thread, it is actually a reason why the estimates may be too low. Most of the research groups doing this analysis only look at titles. I do not have hard numbers, but I am seeing author hallucinations probably just as frequently. (Pure journal hallucinations are rarer but do happen as well, as you can see here!)
This paper should be retracted. The authors did not write it, they outsourced thinking to AI, and they published hallucinations. 1. All the arguments are nonsense; 2. It's substantially AI-written; 3. Several citations are made-up or in error. The authors are frauds.
2
3
27
23,205
Hallucination of the day #2, an article in BMC Geriatrics with 6 hallucinated references. My tool is $2, these journals charge article fees of over $3000.
1
118
Hallucination of the day in Psychology & Marketing. One of the complaints I have been getting are these are not actually hallucinations, but other tools make up bad data. (I had another paper also have a P. Regulation author today in a few of the papers I scanned). Even if you don't think the source of the error is genAI (but some other tool), it is an obvious error none-the-less and should be corrected.
1
103
Pro-tip for folks who want to get a private sector job in data science
1
117
Again, you can't just look at the tables and pre-fit and be able to tell the differences between the real results and the fake ones. It is because the system is under determined. You can likely have many different weight tables that produce nearly the same pre-fit.
1
1
86
Here is the post, andrewpwheeler.com/2026/09/1… And here are the real results, the synth estimate mirrors thefts for the first year+, then it does appear to be higher than you would expect.
1
1
96
Newest blog post (I am lazy and just had AI write it from my notes full disclosure). If criminologists do not disclose data and code for synthetic control designs, they can easily fake the model fit to provide whatever results you want. Here are three fake results for Gascon's tenure in LA as a prosecutor on thefts.
1
2
3
399
Some minor updates to the Crime Decomposition explorer. Included series for violent and property crimes Also made a sub-panel so it is easier to see trends across each crime type. Also a toggle to index those graphs so global and local start at the same line (I am not a big fan of this, may revisit a smarter way to shift up the line later so it is easier to compare). This is for Scranton, can see there was likely some funny business going on for assault reporting mid 2018 (from <10 to ~100 per month, then back down a few years later). Other crime types (with the exception of rape) largely follow the national trends though.
1
1
3
160
Views over the lifetime of my personal blog. I have a total of 580 posts going back to 2012. So almost a post every other week on average. (Although it is very bursty, I just write when I want.) I did take a hit with lower traffic from Google with the onset of ChatGPT. But it is also somewhat cannibalized with starting my Crime De-Coder site. (It would still be decreasing since 2023, but not by quite as much.) Since even before ChatGPT, I suspected much of the traffic was bots. (Maybe half, not sure how to really tell with the analytics wordpress has available to me.) But writing a blog has clearly been more influential for me than any peer reviewed paper I wrote.
106
Newest AI test in writing papers using Claude Sonnet. I had some code I wrote back in late 2023 about how I would find near duplicate surveys (prompted by posts by @RealJonBrauer). Idea is some pollsters fake surveys by duplicating them, but change a few of the responses. A researcher from Survey Monkey and Princeton (Kuriakose and Robbins, 2016) proposed if the items overlap 85% they should be flagged as duplicates. This is bad advice, a survey with rare outcomes for example is simple to get that many overlap responses, so I propose a way to estimate outliers in overlap. So I had already written the code to do this, I just had Claude work on it overnight to turn it into a paper. Can see my prompt and notes in the images for this post. It is admittedly worse than my tests with OpenAIs Sol/Luna. More so for taste than for anything concretely wrong. (It put things in the lit review that should probably be in the discussion, it is succinct but does not explain things in a way a human could easily understand, removed some tables and graphs I think make the results easier to understand.) Probably will not post this one to crimrxiv because I don't like it as is. The crime decomposition paper to be fair I spent maybe two to three days of putzing back and forth with, this I just straight gave the goal and let it go brr while eating dinner. Could likely get this in a place I like better with the same level of effort (that I do not care to put in right now). Given the tools as they are now, I could see doing a paper (that is sufficient to publish in a mid-tier social science journal) maybe once a week. The issue is more having good ideas that are worth spending your time on.
1
2
198
Love how legacy private sector contractors to public sector operate under a cloak and dagger pricing model. Just to be clear, there is not a single dollar value anywhere on this page. Those two images are the entire page (minus a contact form and footer for the website). Also appear to only have software developers in India.
1
159
One of the sources of false positives in VerusCite are actually inserted via the publishers themselves (at least I think they are, have not had a paper accepted at Criminology) This Pubmed link goes to a totally different article
82
LA assaults and their regular seasonal periodicity. apwheele.github.io/CrimeDeco…
56
Hallucination of the day in the Bulletin of Mathematical Biology. One of the things I have not fully wrapped my head around for this, is that the AI tools are basically just parroting what they think a peer reviewed paper should look like. They unlikely have the papers in the weights (or in context) and reasonably cite. Although to be fair, that is not all that different than the superficial cites most humans write now.
1
94
The auto-regressive covariances matter more. This is sometimes called "overdifferencing" in the time series lit, because if the process is actually a unit root this equation equals 0.
1
1
52