A long post!
The Election Commission never had the bandwidth or manpower to carry out such a gigantic clean up of the electoral rolls. It instead relied on an AI-powered application called ERONET.
It uses AI to identify Photo Similar Entries (PSE) and Demographic Similar Entries (DSE) to flag potential mismatches between current voter rolls and legacy databases, like the 2002 base rolls.
The software is essentially a black box, with little to no publicly available information about the AI models actually used in deployment. We do not know the model specs, similarity threshold used, its performance metrics, evaluation benchmarks, false positive rates, how it was trained or evaluated for Indian languages, whether it was developed indigenously or based on a known model/provider. We simply do not know enough.
Anyone with even a basic understanding of AI will tell you that these systems are prone to hallucination errors, and out of distribution (OOD) problems. Even the best large language models, vision models, state-of-the-art OCR technology in the market still struggle to peform with perfect accuracy, particularly on non-Latin scripts.
The 2002 electoral rolls were in native languages and scripts across different states, such as Kannada, Telugu and Bengali. Many AI systems continue to have substantially weaker performance on low-resource languages and scripts compared with English. For ex, a model trained on English will generalize poorly to a low resource language like Telegu or Odia. The old electoral rolls were not originally structured digital databases. They had to be digitised and processed, including through automated conversion into English in some workflows, before they could be compared against contemporary electoral data.
Effectively, it creates multiple points where errors can enter the pipeline through scanning, OCR, transliteration, and finally the AI based matching itself.
That means some level of rigorous human oversight is essential to keep such a system robust and control false positives. We do not know whether adequate human oversight was actually in place, or whether the entire workflow, from OCR and data extraction to AI based matching and automated notice generation, was effectively handed over to the system. The ECI needs to be completely transparent about how this entire workflow was carried out and make the relevant technical documentation, evaluation results, and human-verification procedures publicly available.
There is also a major question around the system's access control and privilege architecture. We do not publicly know the role-based access control (RBAC) structure of ERONET, who has read/write access to the underlying electoral databases, whether EROs or other field level officials can directly modify records, what privileges are available at each administrative level. The Indian Express report seems to indicate that EROs, who have the statutory authority to add/ delete voters, were denied this privilege through the software.
We also need to know whether there is a complete, immutable audit trail recording who accessed, modified, approved or deleted a record, when the action was performed, and what changes were made. Without this information, it is difficult to independently establish the chain of custody and accountability for changes made to electoral roll data.