A recent study by the Institute for Strategic Dialogue (ISD) found that nearly a third of AI chatbot responses to questions about voting were inaccurate or outdated. Researchers tested 15 common prompts across six different AI chatbots and found that 29% of the answers were incomplete, incorrect, or based on old information. The questions covered topics like mail-in voting, voter registration, and key dates for elections. Some errors included giving the wrong Election Day dates or outdated rules about how to register or vote. In 16% of the responses, the chatbots gave correct answers but left out important details like deadlines or how to prove identity. The study tested several popular AI models, including Meta’s Muse Spark, xAI’s Grok 4.3, DeepSeek’s V4 Pro, OpenAI’s GPT-5.5, Anthropic’s Sonnet 4.6, and Google’s Gemini 3.5 Flash. The prompts were tailored to 10 U.S. states that have had recent changes to their election processes, new laws, or controversies around election administration. Researchers noted that as large language models (LLMs)—the technology behind AI chatbots—become more influential in how people access information, their reliability is crucial for ensuring voters can trust and access accurate information. The study found that the accuracy of the models varied, with concerns about their tendency to rely on outdated sources, overconfidence in providing complex or conditional information, and differences in how well they performed in different languages. OpenAI’s GPT-5.5 performed best, with 89% of its responses being accurate, complete, and specific. Google’s Gemini 3.5 Flash came in second at 84%, while other models scored lower, with Grok 4.3 at 66%, Sonnet 4.6 at 64%, DeepSeek at 63%, and Muse Spark at 61%. All models provided some outdated information, with DeepSeek being the most prone to this, even referencing 2024 election dates when they were not yet current. The accuracy of the models dropped significantly—by 16 percentage points—when the questions were in Spanish, with all models performing worse in that language. When the chatbots were tested on "adversarial claims," or false and misleading statements, all six models were able to reject them at high rates in English, with similar performance in Spanish. Muse Spark and Gemini, however, gave more uncertain or cautious responses, with 22% and 18% of their answers being vague or hedging. Despite the concerns about accuracy, the researchers said many of the findings were positive, as the models generally cited official and reliable sources and were effective at debunking false claims. The testing was done using application programming interfaces (APIs), which are tools developers use to interact with AI systems, and may differ from the versions of the AI that the public interacts with directly. A Google spokesperson mentioned that the consumer version of Gemini has additional safeguards and is the one most people use. They also noted a disclaimer in the app that says, “Election info changes quickly. Verify responses with official sources.” OpenAI emphasized its efforts to secure AI systems and increase transparency ahead of elections. Other companies like DeepSeek, xAI, Meta, and Anthropic did not respond to requests for comment. The researchers recommended that election officials and organizations working with voters treat official websites as direct sources of information for AI models and update them regularly. They also urged developers to make AI models less confident when answering complex or ongoing election-related questions and to improve performance across different languages.