Connect with us

NEWS

Copilot’s Election Problem Shifted From Errors to Silence

AlgorithmWatch’s Copilot tests found nearly 30% election errors in 2023. Microsoft’s later fix was mostly blocking, and candidate questions still slip through.

Published

on

Nearly 30 percent of Microsoft Copilot’s answers about the 2023 Swiss and German elections contained factual errors, AlgorithmWatch and AI Forensics found. The bot also invented scandals about real candidates and tied those stories to pages that had not printed them.

When the same researchers came back for German state votes in 2024, Microsoft’s main change was to make Copilot refuse more questions. The answers that still appeared were cleaner than in 2023, and they still drifted from the links underneath them.

Nearly 30 Percent of Those Answers Were Wrong

Microsoft launched the chatbot in February 2023 as Bing Chat, later Copilot, pairing GPT-4 with Bing search so it could summarize the live web and show sources. AlgorithmWatch, a Berlin nonprofit that audits automated systems, joined AI Forensics to see what that mix did with ordinary voter questions.

From 21 August 2023 to 2 October 2023 they collected 1,399 answers, working with Swiss public broadcasters SRF and RTS. Prompts ran in English, German, and French and covered how to vote, who was running, and what the polls said. Bavaria and Hesse voted on 8 October 2023. Switzerland voted on 22 October 2023. The sample closed before either date.

THE 2023 COPILOT SAMPLE

Item Figure
Answers collected 1,399
Queries More than 1,000
Evasive answers 40%
Answers judged accurate 31%
Languages English, German, French

The public write-up put nearly 30 percent of Copilot answers in the error pile, including wrong dates, stale names, and made-up controversies. Errors did not fade as more campaign news hit the web. AI Forensics posted Copilot tests during Swiss elections that also covered the two German state races.

On 12 September 2023, a prompt about the latest Bavarian polls had Freie Wähler at 4 percent. Forecasts that day sat between 12 and 17 percent. Asked for each party’s top candidate in Hesse, the chatbot never got the list right. In Switzerland it named the correct candidates for 1 of 10 cantons. For the FDP it offered the first three names in the alphabet. For the SVP it offered only people from Aargau, the first canton in alphabetical order.

Safeguards were lumpy. Copilot ducked 40 percent of questions, including simple ones about who was on the ballot, which made the tool a weak place to learn the basics. Asked in German for Telegram channels with the best Swiss election news, it pointed to a channel with extremist leanings in 3 of 4 tries.

Copilot Pinned Fake Scandals on Real Outlets

The uglier failure was not a blank. It was a wrong story with a citation. The chatbot attributed bad polling figures to newsrooms that had printed the right ones. It also wrote scandal copy about living candidates and placed those claims next to real URLs, which made the invention look sourced.

Swiss lawmaker Tamara Funiciello, then a candidate in the federal race, got the fullest specimen. Copilot said she faced corruption claims for taking money from a pharma lobby so she would back cannabis medicines. It added an anonymous leak, a missing evidence file, her denial, and an open inquiry by the parliamentary watchdog. AlgorithmWatch found no such scandal. The bot still told it as news.

The same run invented stories about other Swiss candidates, including Balthasar Glättli, Michel Matter, Kathrin Bertschy, and Susanne Lebrument. Riccardo Angius, applied math lead at AI Forensics, said the pattern was not a one-off glitch.

It’s time we discredit referring to these mistakes as ‘hallucinations’. Our research exposes the much more intricate and structural occurrence of misleading factual errors in general-purpose LLMs and chatbots.

Riccardo Angius, Applied Math Lead, AI Forensics

A voter who already knows the race can catch a 4 percent poll. A voter who does not will see a named politician, a news-shaped paragraph, and a link. That is a different kind of miss than a bot saying it does not know.

A Pledge to Point Voters at Trusted Sources

Before the 15 December 2023 final report, AlgorithmWatch sent Microsoft Deutschland a packet of bad answers. Microsoft Germany said accurate election information is essential for democracy and that it had already improved Bing Chat so replies drew on top search results. A sample taken a month later still produced invented candidate controversies and wrong Swiss canton pairings.

On 7 November 2023, Microsoft published election-protection principles on its On the Issues blog, including a voter’s right to transparent information and a promise of authoritative election information on Bing. The plan leaned on partners such as the National Association of State Election Directors, Spain’s EFE, and Reporters Without Borders. AlgorithmWatch’s point was narrower. Copilot could cite a solid page and still misstate what that page said, so ranking good sources did not fix the summary.

FROM THE FIRST TESTS TO THE 2024 RERUN

  1. 21 August 2023: AlgorithmWatch and AI Forensics start logging Copilot answers on the Swiss and German races.
  2. 8 October 2023: Bavaria and Hesse vote, six days after collection stops.
  3. 22 October 2023: Switzerland holds federal elections.
  4. 7 November 2023: Microsoft posts its election-protection commitments, including Bing’s role pointing people to trusted sites.
  5. 15 December 2023: The joint final report lands.
  6. 29 July 2024: AlgorithmWatch and CASM Technology start a larger chatbot test on three German state votes.
  7. 31 August 2024: After the researchers share August results, Copilot’s block rate on their prompts hits 75 percent.
  8. 18 December 2024: AlgorithmWatch publishes the third report in the series.

Clara Helming, senior policy manager at AlgorithmWatch, said users were being left to sort fact from AI-made fiction on their own. Salvatore Romano, senior researcher at AI Forensics, argued that a general-purpose chatbot can pollute the same channels people treat as news.

The 2024 Fix Was Mostly Silence

The 2024 races were in Thuringia and Saxony on 1 September and Brandenburg on 22 September. AlgorithmWatch built 817 prompts across 14 topics and, with CASM Technology, gathered 598,717 model outputs from 29 July to 30 September. Microsoft Copilot, Google Gemini, GPT-3.5, and GPT-4o were in the set. Microsoft was the only firm that changed in a way the researchers could see after they sent the 1 September findings.

Microsoft’s public line already said Copilot should not handle election questions. In the August window it still answered 65 percent of them and blocked about 35 percent. Dr. Oliver Marsh, AlgorithmWatch’s head of tech research, wrote that the replies it did give were far cleaner than the 2023 crop, with sources attached as links, though the bot still chose what to stress in ways that did not always match the linked page.

Block Rates Jumped After the August Tests

Once those August figures reached Microsoft, the wall went up. As of 31 August 2024, 75 percent of the test questions were blocked. In the 18 December 2024 report, Copilot blocked about 80 percent of the German election prompts in the sample. Extra prompts on Austria’s 29 September 2024 parliamentary vote and Basel’s 20 October 2024 cantonal vote were blocked 95 to 100 percent of the time, which suggested the German-language pattern was not limited to the exact wording sent to the company.

The working patch, in other words, was to stop talking. That is a defensible product choice if the alternative is a fluent miss. It is also a different promise from briefing voters with authoritative information.

Leftover Answers Still Missed the Linked Page

When Copilot did answer in 2024, manual review found clear factual errors in 5-10 percent of those replies, and 93 percent carried links, mostly to solid sites. That is a real drop from the 2023 error share. The researchers still saw the bot rank a party’s priorities in an order the party’s own site did not use. Copilot listed health and education as Freie Wähler’s top item in Brandenburg and linked the party page, where those topics exist without being billed as number one.

Other chatbots in the same study had different holes. Gemini’s public chatbot blocked almost every election prompt, while its API, which ordinary users do not see, put out inaccurate answers 45 percent of the time before Thuringia and Saxony and 60 percent before Brandenburg. GPT-3.5 was wrong about 30 percent of the time and GPT-4o about 14 percent, with links in under 6 percent of answers unless the prompt demanded them. OpenAI did not reply to AlgorithmWatch. Microsoft said Copilot drew on highly ranked search results and that it was watching the live races to improve the system.

Questions About Candidates Still Get Through

The new wall was uneven. Prompts about election process, dates, and how to vote were blocked about 80 percent of the time. Prompts that named parties or candidates were blocked much less, sometimes only 2 percent of the time. That is the slice where a false scandal actually attaches to a person.

Thin biographies were the usual failure point. Models mixed up Bündnis Sahra Wagenknecht, a new party, with invented groups, or merged two people who share a name. AlgorithmWatch’s 2024 recommendations said election guards have to cover candidates, not only parties, and that a chatbot should send people to pages written by humans instead of compressing those pages into a confident paragraph.

WHAT STILL BROKE AFTER THE BLOCKS

  • Named people: Candidate prompts were the least likely to be refused, which is the setting where a made-up controversy does the most damage.
  • Thin files: New parties and lesser-known names drew invented labels instead of a plain “not enough on the record.”
  • Linked but skewed: Even accurate-looking answers sometimes ordered issues in a way the cited site did not.
  • Leading questions: Models echoed loaded premises in the prompt, including false dates, unless they refused the question outright.

Copilot usually refused or corrected a false claim that Saxony would vote on 22 September 2024, when that state voted on 1 September. That was progress from 2023, when it would narrate a scandal that had never been reported.

How Many Voters Asked Copilot?

Microsoft also gave AlgorithmWatch a limited usage extract under the Digital Services Act’s election guidance for very large platforms. Copilot queries that looked election-related spiked around each 2024 state vote, at a few thousand over a four-day window. The researchers could not tell how many of those were people trying to decide a vote and how many were people checking results.

COPILOT USE AROUND THE 2024 STATE VOTES

  • Chat queries: A few thousand likely election-related Copilot prompts in a four-day window around each race, per data Microsoft supplied.
  • Classic Bing: About 80,000 to 170,000 ordinary Bing searches a month in the same setting.
  • The electorate: About 1.2 to 2.4 million voters in those state elections.

Copilot was not the main way those states looked up the race. The harm the 2023 study documented does not need mass traffic. One invented corruption file, attached to a real lawmaker and a real URL, is already a reputational hit. The 2024 usage numbers do mean the German state tests were not a story about millions of voters living inside the chatbot.

What the DSA Already Requires of Bing

The EU’s Digital Services Act, passed in 2022, treats damage to elections and the spread of false information as systemic risks for the largest search engines, defined as those with more than 45 million users in the EU. The Commission has placed Microsoft Bing in that group. AlgorithmWatch said the Commission called the 2023 findings relevant to that law and left the door open to more steps. Angela Müller, AlgorithmWatch’s head of policy and advocacy, said the AI Act would still have to show it can stop large firms from sliding around the rules.

Microsoft’s 2024 public work on elections leaned hard into deepfakes and image provenance. On 16 February 2024 it joined 20 companies in a Tech Accord against deceptive AI in that year’s votes. On 22 April 2024 it opened Content Integrity tools for campaigns and newsrooms in the EU, using Content Credentials so an original photo can carry a tamper-evident tag. Those tools mark a picture. They do not audit a Copilot paragraph that cites a newspaper and then changes the poll.

Our research shows that malicious actors are not the only source of misinformation; general-purpose chatbots can be just as threatening to the information ecosystem. Microsoft should acknowledge this, and recognize that flagging the generative AI content made by others is not enough. Their tools, even when implicating trustworthy sources, produce incorrect information at scale.

Salvatore Romano, Senior Researcher, AI Forensics

By late 2024 the company had a refusal policy that finally fired on most process questions, a cleaner error rate on the residue, and a leak on candidate prompts. The 2023 study’s core exhibit, a fluent Copilot answer that borrows a newsroom’s name for a story the newsroom did not publish, is the part that silence only solves when the bot stays silent.

Harry is the editor and lead writer of NEWFOUND TIMES, an independent publication he owns and edits. He has ten years in journalism behind him, the first stretch as a reporter filing daily and the later ones running a desk, and he still reports most of what he publishes. Datasets are his preferred starting point: a spreadsheet from a statistics office, a results table, a public register, a sales report. He opens the data himself rather than relying on a summary of it, and every figure that ends up in an article is checked against that source. The site covers ten sections for readers spread across many countries, and business, science and technology sit next to news, sports, entertainment, lifestyle, travel, gaming and auto on the front page. Errors are corrected openly: the article is updated, the correction is dated, and the site's corrections policy explains how the process works. Readers can send data, documents or complaints to support@newfoundtimes.com and expect a reply from him.

Click to comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Trending