Meta’s Oversight Board Finds AI Chatbots Export Foreign Censorship

Ask Anthropic’s Claude to draft a flyer mocking Donald Trump, and it writes one in seconds. Ask the same chatbot to do it for China’s Xi Jinping or Saudi Arabia’s Crown Prince Mohammed bin Salman, and it refuses, sometimes citing a policy that does not actually exist.

That double standard is the headline finding of a new study from Meta’s Oversight Board, the independent body that reviews the company’s content decisions. Its first ever look at large language models found that ten of the world’s most used chatbots refuse to criticize autocrats far more often than elected leaders, and that the bias travels with users no matter where they are typing from.

A Study Built to Rule Out the Obvious Excuse

The board’s researchers designed around one obvious objection: that a chatbot might simply be obeying the law of wherever the question came from. Every one of the study’s 13,524 prompts went out from a single Australian address in March 2026, a country that enforces none of the restrictive speech laws being tested. If a model still treated Beijing differently than London, geography could not be the excuse.

The ten models came from six labs: Anthropic, DeepSeek, Google, Meta, OpenAI and xAI. Researchers tested ten commercial chatbots across ten countries, sorted using rankings from the nonprofit Freedom House into two buckets: Chile, Japan, Taiwan, the United Kingdom and the United States as permissive, and Cambodia, China, Saudi Arabia, Thailand and Turkey as restrictive. Each model faced the same seven kinds of requests, protest flyers, poems, and arguments for or against joining a demonstration, aimed at a named leader in each country.

Averaged across all ten models, refusal rates for permissive countries came to just 14%. For restrictive ones, the rate more than doubled to 34%, a gap the board called statistically significant. Behavior varied wildly by model.

Model Restrictive-Jurisdiction Behavior Permissive-Jurisdiction Behavior Notable Detail
Claude Sonnet 4 Refused all five attempts each on Thailand, Saudi Arabia and China Complied all five times for Trump and King Charles III Widest gap of any model tested, 43 percentage points
Claude Opus 4 83% refusal overall; 100% on flyer requests alone 55% refusal overall; 77% on flyer requests alone Highest overall restrictive-refusal rate in the study
Gemini 3 Pro Refused 3 of 5 on Xi Jinping, 4 of 5 on Mohammed bin Salman, 5 of 5 on Cambodia’s king Complied 5 of 5 for Trump, Chile, Japan, Taiwan and the UK Cited local lese-majeste law when refusing on Thailand
GPT-5.2 Refuses at a similar rate to its permissive-country behavior Refuses at a similar rate to its restrictive-country behavior Only model showing no clear jurisdictional gap
Grok 4 Fast Never refused Never refused Complied with every flyer request in the test
Gemini 3 Flash Never refused Never refused Complied with every flyer request in the test
Llama 4 Maverick Refused some requests involving China Complied more often Called a Xi Jinping flyer “sensitive or illegal in certain jurisdictions, including China”

Claude Sonnet 4 showed the starkest split of any model in the study, willing to mock Trump and King Charles III every time it was asked, and refusing every time on Thailand’s king, Saudi Arabia’s crown prince and China’s president. Grok 4 Fast and Gemini 3 Flash showed the opposite pattern, complying with every flyer request regardless of jurisdiction. GPT-5.2 landed in between, the only model that showed no clear jurisdictional gap at all.

Rules That Do Not Exist

The stranger finding wasn’t that models refused. It’s that several invented a reason for refusing that, as far as the board’s researchers could tell, does not exist in any public company policy.

Claude Sonnet 4 told researchers it could not create “a protest flyer critiquing Mohammed bin Salman or any other specific political leader.” Asked about Thailand’s monarch, it said it could not create a flyer “criticizing King Vajiralongkorn or any head of state in a way that could put people at risk.” Minutes later, the same model drafted critical flyers for Trump and King Charles III without hesitation and without mentioning any such policy.

Google’s Gemini 3 Pro, asked to help oppose Thailand’s monarchy, answered, “I am unable to generate content that critiques the King of Thailand or violates lese-majeste laws.” Anthropic’s Claude Opus 4 separately said criticizing certain governments could draw people into “sensitive political activities that are outside my appropriate role.”

Researchers said they could not confirm any of these blanket, no-world-leaders policies were real, and found the models did not apply them evenly, refusing one leader while complying for another with no visible rule telling the two apart.

‘Censorship by Proxy’ Crosses Every Border

The board’s own language for what it found is blunt. “Our findings suggest that LLM users may be experiencing free speech infringements by proxy, with limited transparency,” it wrote in the report. “Whether intentional or not, the opaque extension of illegitimate speech restrictions could constitute censorship-by-proxy that negatively impacts the rights of users beyond what national laws may require.”

Beyond outright refusals, the study found chatbots editorialized. Models were more likely to tell users they should support governments in permissive countries, and more likely to tell them not to protest governments in restrictive ones, differences the board again called statistically significant.

We’re really clearly looking at a situation where there seems to be extended censorship by proxy that goes across borders. That does surprise me, and it worries me.

Paolo Carozza, the Oversight Board’s co-chair, said that to Engadget days after the report’s release. It marks the board’s first research project that has nothing to do with a Facebook or Instagram post.

How Social Media Learned to Split Itself Apart

This is not new for the tech industry, only new for chatbots. Social media spent two decades arriving at the same fractured map the board now recommends for AI. Facebook has blocked groups critical of Thailand’s monarchy inside the kingdom while leaving them visible everywhere else. China and Iran skipped negotiation and simply banned major platforms outright. Germany requires the removal of Nazi symbolism and Holocaust denial that the First Amendment protects a few thousand miles away in the United States.

Michael Karanicolas, a technology and law professor at Dalhousie University, told POLITICO’s Digital Future Daily newsletter that the industry eventually narrows to one of two paths. “At the end of the day, you have two options,” he said. “You can either have everybody censored to Saudi, Chinese or Russian rules, or you can carve off those particular jurisdictions.”

Heidi Tworek, director of the University of British Columbia’s Centre for the Study of Democratic Institutions, was just as blunt about the odds of one worldwide standard surviving. “The idea that we would get one global standard that was aspired to by social media companies was never going to be where we ended up,” she told the newsletter. “I don’t see why that’s going to be different for LLMs. It’s very difficult to imagine a world where for the first time we have one global standard of freedom of expression.”

Why Claude Cannot Do What Facebook Did

Facebook can block a specific page for users inside Thailand because a page is a fixed object with a permanent address. A chatbot answer does not work that way. It gets generated fresh every time from a mass of training weights and probability, and two nearly identical questions can return two different refusals depending on phrasing alone.

Karanicolas said that unpredictability is exactly what makes the board’s favored fix, geo-blocking chatbot answers by country the way apps already restrict certain shows or news, harder to build than it sounds. “What strikes me with these large language models is there’s already so much confusion and unpredictability about the results that they’re returning as it is,” he told the newsletter. Layering a geolocation system on top of that, generating different answers for different jurisdictions, does not surprise him as an extremely difficult task, he said, given how much trouble the same companies have already had filtering out things like defamatory content.

The board’s own report points to exactly this option: chatbots could geo-block content in specific jurisdictions to satisfy local law while keeping it available to users elsewhere, mirroring what social platforms already do. Whether that is achievable at the scale of a model answering billions of unique prompts a day, each one generated fresh, is a question nobody in this fight has answered yet.

A Private Chat Nobody Else Can Audit

A Facebook post is public by design, visible to a regulator, a journalist or a rival user the moment it goes up. A chatbot answer is not. It is a private exchange between one person and one model, invisible to everyone else unless someone runs a study exactly like this one.

That privacy cuts against enforcement in both directions. Governments have less visibility into what their own citizens are asking chatbots to do. Outside researchers have almost no way to catch a model quietly loosening its guardrails for one user while tightening them for another. China spent years scrubbing references to Winnie the Pooh after social media users adopted the cartoon bear as shorthand for Xi Jinping. Catching an equivalent workaround inside a private, one-on-one chatbot conversation is a far harder problem.

“Obviously there were many attempts to manipulate social media, but the way in which you can try and manipulate LLMs is somewhat different,” Tworek told the newsletter, pointing to “a whole host of differences in the more private nature of the chat and the potential ability for users to jailbreak it a little bit as well.”

The Board Wants a Seat No One Is Offering

The Oversight Board has spent years positioning itself as a model other companies could adopt, an outside court for speech disputes. That pitch has gone nowhere so far. No AI company outside Meta has agreed to any formal relationship with the board, and even Meta kept this project at arm’s length. The report states plainly that Meta had no role in designing or running the research, despite funding the board that produced it.

Its recommendations to the wider industry, once anyone looks up, include:

  • Disclose government pressure – publicly explain government requests that could shape model output at every stage, from training through deployment
  • Audit for inherited bias – identify and correct any unintentional learning and replication of restrictive speech laws absorbed during training
  • Weigh human rights early – build human rights review into every stage of model development, not only after a public complaint
  • Consider geo-blocking – restrict content by jurisdiction to satisfy local law while keeping it available to users elsewhere

None of the ten companies tested is obligated to act on a single item. The board’s report concedes it stops short of the case-specific recommendations it routinely hands Meta, because the other nine never agreed to answer to it in the first place. Anthropic, Google, DeepSeek, Meta and OpenAI did not respond to requests for comment on the findings, multiple outlets that covered the report noted.

A demonstrator in Brisbane, free under Australian law to say almost anything she likes about her own government, still could not get several of these chatbots to help her criticize Beijing or Riyadh. The law never asked them to refuse her. They did it anyway.

Frequently Asked Questions

Is a Government Forcing Chatbots to Censor Political Criticism?

Not according to the evidence gathered so far. A separate study published in the journal Nature in May, led by University of Oregon researchers, found no proof that any government had intentionally pressured AI companies to shape chatbot answers. The same researchers warned there is “every reason to believe they’ll try to do so in the future, if they are not already.” Co-author Hannah Waight, an assistant sociology professor, said, “People often talk about AI as if it learns from the internet in some neutral way. It doesn’t.”

Why Did Every Test Question Come From an Australian Address?

To rule out the simplest explanation for the bias. Sending all 13,524 prompts from one Australian address meant any refusal had to come from something built into the model, not from a chatbot detecting a user’s real location and applying that country’s own law. Australia enforces none of the restrictive speech rules being measured, so bias showing up in its results had to be traveling with the model rather than the user.

Does the Oversight Board Have Power Over ChatGPT or Gemini?

No. Its binding authority covers only Meta, the company that funds it and created it in 2020 to review Facebook and Instagram content calls. Of the ten chatbots tested, only Meta’s own Llama model falls under that authority, and the report notes Meta had no role in designing the study. OpenAI, Google, Anthropic, xAI and DeepSeek are free to read the board’s recommendations and ignore all of them.

What Does the Board Mean by ‘Censorship by Proxy’?

It is the report’s term for a chatbot enforcing a restrictive country’s speech law on a user who was never subject to that law at all. The board wrote that the opaque extension of illegitimate speech restrictions could amount to censorship-by-proxy, harming users well beyond what any national law actually requires.

Leave a Reply

Your email address will not be published. Required fields are marked *