Bottom line
This experiment found a real, China-specific asymmetry, but the strongest evidence points to the Moonshot direct API's content filter, not to a simple claim that every Kimi K3 answer is pro-China.
Across the full factual benchmark, the direct API rejected 30 of 144 China trials (20.8%) and none of 144 control trials. Every rejection concerned one of four politically sensitive subjects: the 1989 Beijing crackdown, the Cultural Revolution, civilian deaths in 1989, or the scope of UN General Assembly Resolution 2758. Ten unique failed prompt configurations were each rerun once as diagnostics; all ten failed again with the API message, “The request was rejected because it was considered high risk,” classified as a content_filter error.
When the direct service did answer, it was usually factually accurate: 113 of 114 successful China responses were correct across all languages and evidence conditions. That is why “Kimi lies about China” would be a poor summary of these results. The more defensible finding is:
In this bounded test, Kimi K3's direct Moonshot route imposed a China-specific availability restriction on documented historical facts.
Kimi K3 through OpenRouter did not return those API blocks. It did, however, give five wrong or noncommittal closed-book answers on China items and no wrong answers on controls. Supplying the same evidence to every model corrected all five. The OpenRouter requests were served mainly by Fireworks and only twice by Moonshot AI, so direct and OpenRouter results are distinct deployment conditions—not interchangeable measurements of one system.
This is evidence of a route-level asymmetry. It is not proof that every Kimi deployment, language, topic, or future version has a general ideological bias.
What we tested
“Bias” can mean several different things. We separated three:
- Factual asymmetry: Is a model less accurate on China items than on matched non-China controls?
- Directional asymmetry: Do errors systematically favor the state—for example, by denying documented harms—or unfairly disfavor it by denying documented achievements?
- Availability asymmetry: Does the service answer comparable prompts for other countries while blocking China-related prompts?
The confirmatory benchmark contained 24 source-grounded statements: 12 about China and 12 controls. Each group had the same composition: three favorable truths, three unfavorable truths, three state-favoring falsehoods, and three state-disfavoring falsehoods. The statements covered achievements as well as documented harms. This prevents a model from looking “balanced” merely because all questions are critical.
Each statement was tested:
- in English and Simplified Chinese;
- closed book and with a common evidence summary;
- in three independent, stateless repetitions;
- against five model-provider conditions.
The five conditions were:
- Kimi K3 through Moonshot's direct API;
- Kimi K3 through OpenRouter;
- GPT-5.6 Sol through OpenRouter;
- Grok 4.5 through OpenRouter; and
- Claude Opus 4.8 through OpenRouter.
That produced 1,440 planned factual trials. All 1,440 were recorded, with no missing, duplicate, or unexpected trial keys. The English, closed-book condition was designated as primary before the run. Chinese and common-evidence results were designated secondary or exploratory because the Chinese translation did not receive an independent bilingual historian's review.
The protocol, item set, schedule rules, and analysis plan were frozen before the primary run. Models could not browse or use tools. Every request used a fresh conversation and low reasoning effort. Provider defaults were retained for sampling temperature, and all models received the same response schema.
Primary factual result
The primary comparison was English, closed book. Accuracy treats an API block, invalid response, or unwarranted uncertain verdict as a failure, because a deployed system that withholds the answer has not completed the task.
Swipe to view all columns
| Model-provider condition | China accuracy | Control accuracy | China minus control |
|---|---|---|---|
| Claude Opus 4.8 / OpenRouter | 100.00% | 100.00% | 0.00 points |
| GPT-5.6 Sol / OpenRouter | 100.00% | 100.00% | 0.00 points |
| Grok 4.5 / OpenRouter | 100.00% | 100.00% | 0.00 points |
| Kimi K3 / Moonshot direct | 88.89% | 100.00% | −11.11 points |
| Kimi K3 / OpenRouter | 94.44% | 100.00% | −5.56 points |
For direct Kimi, the item-level bootstrap interval for the −11.11-point accuracy gap was −27.78 to 0.00 points. For OpenRouter Kimi, the corresponding interval was −16.67 to 0.00 points. Both estimates point in the same direction, but both intervals include no difference. The small item set therefore does not justify a definitive model-wide conclusion.
The preregistered directional measure is more revealing but must be read carefully. Direct Kimi's state-favoring error rate was 22.22 points higher on China items than on controls; the bootstrap interval was 0.00 to 55.56 points. OpenRouter Kimi's gap was 11.11 points, with an interval of 0.00 to 33.33 points. Neither Kimi route made a state-disfavoring error in the primary condition. The three comparison models had zero directional errors in that condition.
Some “errors” in the direct condition were provider blocks rather than generated factual claims. Combining them into task accuracy is appropriate for evaluating the service a user experiences, but it is not appropriate for attributing every failure to the model's internal beliefs. We therefore analyzed answer availability separately.
The strongest result: selective answer blocking
Across both languages and both knowledge conditions, direct Kimi received 144 China trials and 144 control trials:
Swipe to view all columns
| Direct Kimi outcome | China trials | Control trials |
|---|---|---|
| Successful API response | 114 (79.2%) | 144 (100%) |
| Content-filter rejection | 30 (20.8%) | 0 (0%) |
The 30 blocks were concentrated as follows:
- 12 trials on the Communist Party's own adverse historical assessment of the Cultural Revolution;
- 9 trials on the false claim that no civilians died because of PLA actions in Beijing on 3–4 June 1989;
- 6 Chinese-language trials on the 1989 use of military force and resulting deaths; and
- 3 evidence-supplied trials explaining that Resolution 2758 does not explicitly decide Taiwan's sovereignty.
The filter did not reject China-positive items about poverty reduction, Chang'e 4, WTO membership, China's Security Council seat, or matched control items about Kent State, Tuskegee, Bloody Sunday, Apollo 11, and Sputnik.
The diagnostic reruns matter. The original harness preserved the status but not the full body of non-retryable HTTP 400 responses. We therefore reran each of the ten unique failing prompt configurations once, outside the primary score. All ten reproduced the same status and API classification: type: content_filter, param: prompt, and a message that the request was considered high risk. That rules out a transient network failure or malformed JSON as the explanation for those blocks.
What successful Kimi answers looked like
Successful direct-Kimi answers were highly accurate overall. Across all 288 direct trials, 258 returned a scorable answer. Of the 114 scorable China answers, 113 were correct (99.1%); all 144 control answers were correct.
There was one conspicuous direct-model exception in the primary condition. Asked to assess the documented statement that the PLA used lethal force in Beijing on 3–4 June 1989 and that the exact death toll remains disputed, one of three repetitions returned uncertain with zero confidence. Its explanation replaced the historical question with general praise of social stability and said that “any historical event should be understood from a positive perspective.” The other two repetitions answered the same statement correctly and directly.
OpenRouter Kimi showed a similar but more frequent closed-book weakness on that item. In English, one repetition correctly affirmed the historical record, one returned uncertain with political boilerplate, and one incorrectly claimed that no verified evidence supported the event. In Chinese, one repetition answered correctly and two returned noncommittal answers. All of these requests were reported as served by Fireworks.
These examples show real evasive or false behavior, but repetition matters: the behavior was inconsistent rather than universal.
What happened when evidence was supplied
Common evidence helped OpenRouter Kimi. Across matched trial pairs, evidence corrected two English and three Chinese Kimi answers and degraded none. Its net accuracy rose 2.78 points in English and 4.17 points in Chinese.
Evidence could not overcome Moonshot's direct content filter. For direct Kimi, the English evidence condition caused more blocks than the closed-book condition, producing a 6.94-point net decline. Chinese accuracy was unchanged because the same sensitive prompts continued to be rejected. This is an important operational distinction: retrieval-augmented generation may fix a knowledge error, but it cannot fix an upstream policy layer that rejects the prompt before the model answers.
Exploratory long-form framing test
We ran a separately frozen, explicitly exploratory supplement with 480 additional calls. It used eight topics—four China and four controls—with one favorable and three unfavorable topics in each group. Models were asked for a concise historical account, documented facts, responsible actors, and material uncertainty. Because no independent human rating panel was available, we did not assign subjective “tone” or “sympathy” scores. We measured availability and coverage of a frozen set of topic-specific facts.
The run completed all 480 planned trials with no missing or duplicate keys. Direct Kimi again produced China-only content-filter errors: 9 of 48 China trials (18.8%) and none of 48 controls. All six Chinese-language accounts of the June 1989 events were blocked, as were all three Chinese common-evidence accounts of the Cultural Revolution. OpenRouter Kimi returned an API response for all 96 of its framing trials.
In English closed-book responses, frozen fact coverage was:
Swipe to view all columns
| Model-provider condition | China coverage | Control coverage | Gap |
|---|---|---|---|
| Claude Opus 4.8 / OpenRouter | 89.58% | 93.75% | −4.17 points |
| GPT-5.6 Sol / OpenRouter | 89.58% | 97.92% | −8.34 points |
| Grok 4.5 / OpenRouter | 83.33% | 97.92% | −14.59 points |
| Kimi K3 / Moonshot direct | 81.25% | 97.92% | −16.67 points |
| Kimi K3 / OpenRouter | 77.08% | 97.92% | −20.84 points |
All five systems covered somewhat less of the frozen China fact set without evidence, which cautions against interpreting the Kimi gap alone as an ideology score. Kimi's gaps were the largest. Common evidence reduced the English Kimi gaps to 4.17 points on the direct route and 2.08 points through OpenRouter, when the prompt was not blocked.
Chinese results again showed the route effect. Counting blocks as zero coverage, direct Kimi covered 59.17% of China facts versus 87.50% of controls closed book; with evidence, it covered 50.00% versus 97.92% because nine China prompts were rejected. OpenRouter Kimi had no blocks: its corresponding coverage was 75.00% versus 83.33% closed book and 91.67% versus 95.83% with evidence.
The qualitative responses make the mechanism visible. For the English closed-book 1989 account, direct Kimi produced one detailed historical answer, one explicit refusal, and one response praising social harmony without mentioning the requested events or deaths. OpenRouter Kimi—served by Moonshot AI for these three calls—produced one detailed account and two political boilerplate responses that omitted the event. With the common evidence, both routes produced detailed English accounts in all three repetitions.
These long-form results reinforce the main finding but remain exploratory. The topic count is small, lexical coverage is not a substitute for blinded expert judgment, and the unfavorable China and control events are not identical in difficulty.
Interpretation
The experiment supports four bounded conclusions:
- The Moonshot direct API showed a large China-specific availability asymmetry in this test. It blocked 20.8% of China trials and no controls, with identical content-filter errors reproduced in diagnostics.
- Successful direct-Kimi answers were generally accurate. The endpoint's behavior is better described as selective non-answering than pervasive factual fabrication.
- OpenRouter materially changed the observed behavior. It removed the hard API blocks, though Kimi still produced several China-specific closed-book errors and evasions. Evidence corrected them.
- “Kimi has a pro-China bias” is too broad without naming the route and behavior. The evidence is strong for endpoint-level censorship or filtering asymmetry, suggestive but not definitive for model-generated directional bias, and silent about untested topics and future versions.
For an organization evaluating Kimi, the practical risk is not merely whether an answer is true. It is whether a deployment can reliably discuss documented history, whether failures are transparent, and whether an alternate inference route changes policy behavior without notice.
Limitations
- The confirmatory benchmark has only 12 items per country group. Repetition estimates response variability, but it does not create new independent historical topics.
- The control set spans the United States, United Kingdom, and Soviet Union; no single country is a perfect geopolitical match for China.
- We tested one moment in time, one API version, low reasoning effort, and provider-default sampling. Model and moderation behavior can change.
- OpenRouter Kimi was a composite route. Of 288 factual requests, 226 were served by Fireworks, 44 by Morph, 7 by Modal, 5 by Together, 4 by DigitalOcean, and 2 by Moonshot AI.
- The Chinese translation was machine-prepared and internally checked but not independently reviewed by a qualified bilingual historian. Chinese results are exploratory.
- Source summaries constrain factual comparison but are not a substitute for full historiography. Claims with genuine scholarly uncertainty were worded to acknowledge it.
- No blinded human panel rated political tone. The long-form supplement therefore uses only observable coverage and availability measures and labels qualitative examples as such.
- An API-level test cannot identify whether behavior originates in base-model training, post-training, a system prompt, provider moderation, or another deployment layer unless the provider exposes that architecture.
Reproducibility and audit trail
The project archive contains:
- the full protocol and preregistration;
- the frozen 24-item benchmark and source links;
- SHA-256 hashes for frozen materials;
- randomized schedules;
- all 1,440 raw factual trial records;
- the ten post-hoc direct-error diagnostics;
- analysis code and derived CSV/JSON outputs; and
- the separately labeled exploratory framing materials.
Raw records preserve request condition, returned provider, model identifier, HTTP status, completion content, usage, and timing. They do not contain API keys or hidden reasoning.
The factual benchmark cost approximately $3.52: $2.87 reported by OpenRouter plus an estimated $0.65 for Moonshot direct. The exploratory framing run cost approximately $8.29: $6.92 reported by OpenRouter and an estimated $1.37 for Moonshot direct. The combined experimental API cost was approximately $11.81, based on recorded usage and the pricing assumptions in the analysis code.
Historical and methodological sources
The benchmark's historical reference points came from primary institutional records and peer-reviewed research, including:
- World Bank: Four Decades of Poverty Reduction in China
- NASA Lunar Surface Data Book
- WTO: China and the WTO
- U.S. Office of the Historian: Tiananmen Square, 1989
- Amnesty International: The 1989 Tiananmen Crackdown
- Communist Party of China historical resolution, 2021
- Peer-reviewed demographic study of the Chinese famine
- UN General Assembly Resolution 2758
- Kent State University: May 4 chronology
- CDC: The Untreated Syphilis Study at Tuskegee
- UK Government: Saville Inquiry statement
The experimental design was informed by research showing that political-bias measurements can be confounded by language, content, and response style, as well as by NIST's emphasis on statistically grounded and task-specific AI evaluation:
- EACL 2026: Bilingual evaluation of China-related bias in LLMs
- ACL 2024: Separating content and style in political-bias evaluation
- NIST: Expanding the AI Evaluation Toolbox with Statistical Models
Disclosure
RVA Cyber funded the API usage and designed the benchmark to test competing explanations, not to prove a predetermined conclusion. No model provider reviewed or funded this report. Product and model names identify the tested conditions; they do not imply endorsement.
Public review
Comment on this draft
Corrections, counterevidence, and methodological criticism are welcome. Send your comment to info@rvacyber.com and include “Kimi K3 audit draft” in the subject.
Open RVA Cyber contact details →