DEV Community

Cover image for Does shallower reasoning make an AI's Japanese more natural? I compared five drafts and the assessments split
matsumotory
matsumotory

Posted on Originally published at aird.matsumoto-r.jp

Does shallower reasoning make an AI's Japanese more natural? I compared five drafts and the assessments split

Summary

On 2026-08-31 I passed on one impression to the AI agent I write articles with. The impression was that even a model like Claude Opus 5 seems to write natural Japanese, without over-reasoning, when you run it with effort set to medium. Effort is an API parameter that decides how much thinking a model spends on reasoning. The AI agent wrote this impression into the rules document the same day. It also collected five studies that looked like they would support it. And it went as far as having five drafts written from the same material and comparing them, all on that same day.

The results of the experiment, though, did not come out the way my impression said they would. When I counted violations of the Japanese rules, the drafts written with effort lowered came out better. When I had an outside model rate overall naturalness, on the other hand, the drafts written with effort raised came out better. In this article I explain first how the impression became a rule. Next I show how the assessments split when I compared the five drafts. Then I explain why two of the studies I collected did not hold up as support. Last I write what I decided in light of that.

What you can take away

For people who have an AI write Japanese prose, I explain the following three things.

  • Whether you should raise or lower effort can flip between having an AI write Japanese and having it check the facts and logic of what an AI wrote. I explain how I split effort between those two
  • I made a point of not settling a rule on an impression alone. On the same day, I wrote into the body of the rule that I would record what came out of running the AI under it and check later whether the rule was right. I explain how that came about, and the comparison experiment I ran the same day
  • I have the studies an AI collects read again the same day by an AI other than the one that collected them. For two of the five, the claim I was about to write into the article was either absent from the source or the opposite of what the source said

The product names of AI models in the body refer to the ones actually used in this operation, and they do not represent the views of their providers.

How I set effort and run it now

To give the conclusion first, whether shallower reasoning makes Japanese natural came out differently depending on how I compared. So I have not turned this impression into a settled rule. Here is how I run it now.

  1. When having an AI write Japanese (generating, polishing, drafting fixes), choose effort explicitly, not just the model name. Do not let it silently inherit the deep setting of the whole session
  2. When having it check the facts and safety of what an AI wrote, keep effort deep
  3. Use low for neither, because low increases missed detections
  4. Leave polishing and drafting fixes to Claude Sonnet 5 as before, until measurement confirms my impression is right. Put on hold the proposal to make Opus 5 at effort medium the first choice
  5. The session's leader model writes the final draft directly. This decision is not among the ones on hold

I explain in order how I arrived at running it this way.

What happened on the day I passed on the impression

On 2026-08-04, about four weeks before I passed on this impression, I had decided not to use Claude Opus 5 for checking Japanese. There were two reasons. I once had an Opus 5 checker, an AI that checks text against the rules document, look for the mistaken simplification that breaks standard compound words down into native Japanese wording. That checker left noun-phrase compression inside its own suggested fixes. Two days before that, when I wrote and compared four drafts from the same material, the Opus 5 draft had noun phrases strung together to excess. Since then Claude Sonnet 5 handles the Japanese checks, and Claude Fable 5 takes over when Sonnet is not enough.

On 2026-08-31 I told the AI agent to narrow the scope of that decision. What I meant was that Opus 5 too seems to produce natural Japanese without over-reasoning if you set effort to medium, and that I wanted to put that to use. I thought one cause of unnatural Japanese might lie in over-reasoning rather than in the choice of model. In response, the AI agent rewrote the decision from 2026-08-04 so that it applies only to the setting of that time, when the model ran with deep reasoning and no effort specified. It kept the record of the problems that actually happened and narrowed only the range the decision reaches. Even for a rule I have already decided, I make a point of revisiting its range when a new observation or comment comes up. There is a worry that fixing a rule decided with an AI while its grounds are still weak lets a mistaken practice settle in, and I wrote about it earlier in Who decided on the half-width space between Japanese text and alphanumerics? Tracing it back, I found an AI had decided it with no grounds.

That same day, two AI sessions running in parallel each received what I had said separately, and both revised the rules document independently. The session that got its change into the rules document first kept Sonnet 5 as the Japanese checker, as before. On top of that, it required choosing a model name and an effort level together, but only when having an AI write Japanese. The setting to try first was running Opus 5 at effort medium. The other session, which noticed later, withdrew its own reading that Opus 5 should also come back as a checker under conditions, and fell in line with the earlier version. It then wrote the content of that earlier version into the document that sets which model goes to which task, and into the settings checks of the two article-writing workflows.

Under the rule left standing after all this, I keep effort high when checking and set effort to medium when writing. First, when I have an AI check the facts and safety of written text, I keep effort deep. Second, when I have an AI write Japanese, I state effort as medium and do not let it silently inherit the session's deep setting. Third, I use low for neither, because it increases missed detections. Those are the three points. The AI agent wrote these three points into seven places in the documents the same day. They are the Japanese writing rules, the definition of which model goes to which task, the settings checks of the two article-writing workflows, the procedure for the daily writing job and its skill, and the settings of the polishing workflow. My impression, though, is still one person's subjective view. So on the same day the AI agent wrote into the body of the rule that it would keep a record of what came out of writing articles under this setting, and add to the rule only after the results were confirmed in review.

Comparing five drafts written from the same material

On the same day it wrote into the rule, the AI agent started an experiment comparing drafts. It varied the effort of Claude Fable 5 and Claude Opus 5 across medium, high, and max. Then, with the same material and word-for-word the same instructions, it had five final drafts written from scratch. It hid which draft came from which setting and compared them in three ways. Two models and three effort levels make six combinations. Five drafts remain in the record. Which combination is missing cannot be pinned down from the record. It prepared three ways of comparing. A Sonnet 5 checker counts violations against the Japanese rules document. Three independent recalculations check whether the numbers in the material match what the text says. And Gemini, a model from an outside provider, rates overall naturalness.

Way of comparing Result
A Sonnet 5 checker counts rule violations For Fable 5, the draft at effort medium had the fewest violations of translationese and of style. For Opus 5 it went the other way, and the draft at effort max had the fewest
Three independent recalculations of whether the numbers in the material match the text Only in the draft Fable 5 wrote at effort medium did all three recalculations find the same mismatches with the material. They were a wrong count, and a statement that settled something that had not been measured yet
Gemini rates overall naturalness Gemini rated the drafts as more natural the higher the effort, for both models

In short, counting violations made the lower-effort draft the better one for Fable 5, while rating naturalness made the higher-effort drafts the better ones for both models. And the Fable 5 draft at effort medium, the one with the fewest rule violations, was the only one that carried mismatches with the material.

The comparison procedure itself also ran into a problem. The order in which drafts are handed to a checker can affect the assessment. The AI agent shuffled the order of the drafts for each checker to avoid that effect. Three of the four checkers then mixed up the draft labels with the order they had read them in. The AI agent matched the quotations the checkers returned against the text of the drafts. It could confirm that all three had copied them the other way round, so the results were recoverable. Out of this, the AI agent added two items to the rules for work it repeats on its own. One is to match the quotations a checker returns against the original text before tallying anything. The other is to fix the order in which drafts are handed over rather than shuffling it, and to catch mix-ups in reading order through that matching of quotations.

What came of having a different AI read the five collected studies again

In parallel with writing into the rule, the AI agent set out to check in primary research sources whether my impression had any support in theory. In its 2026-08-31 search, the AI agent collected the studies into three groups. The first group is two studies on multilingual reasoning (When Models Reason in Your Language and A Comprehensive Evaluation of Multilingual Chain-of-Thought Reasoning). From those the AI agent took two readings. Reasoning models tend to think mainly in English even on non-English tasks. Pinning the language of thought to the user's language lowers accuracy. The second group is one empirical study on reasoning time (Inverse Scaling in Test-Time Compute). It said there are tasks where extending reasoning time makes scores go down instead. The third group is two reports on creative writing tasks (COIG-Writer and When Reasoning Supervision Hurts). The AI agent had collected these two as reports that adding reasoning has no effect or does harm.

On 2026-09-02, an AI in a context separate from the one that collected the studies fetched these five sources and read them again. It checks three things: whether the source exists, whether the bibliographic details are right, and whether the claim I was about to attribute to it really appears in the source text. To prevent fabricated quotations and references, this site always runs this check before publishing. Here are the results in a table.

Study The claim I was about to write What the second reading found
When Models Reason in Your Language (EMNLP 2025 Findings) Reasoning models think mainly in English even on non-English questions, and pinning the language of thought to the user's language lowers accuracy Matches what the source says. The subjects, though, are six open-source distilled models and short answers in math and science, so it does not extend to commercial models or to writing prose. The language used for thinking is sometimes Chinese, not only English
A Comprehensive Evaluation of Multilingual Chain-of-Thought Reasoning (EACL 2026 Findings) The bias toward English-centered reasoning, and noise from translation Only the first half is in the source. Noise from translation is not written there, and the source carries the opposite result, that thinking translated into English is more accurate
Inverse Scaling in Test-Time Compute (TMLR 2025) There are tasks where extending reasoning time makes scores go down instead, and this was confirmed on several models Matches what the source says. It is a result shown on deliberately constructed task types, though, not a law that thinking longer generally lowers scores. How the breakdown happens differs by model family
COIG-Writer (arXiv:2510.14763) Training without a reasoning process attached gave better creative writing scores The source states the opposite conclusion. It reports that training on creative data with the process attached, mixed with general data, works well, and it holds no experiment comparing the presence or absence of the process at all
When Reasoning Supervision Hurts (arXiv:2605.20364) Reasoning supervision lowers the quality of long-form literary generation What the study covers is different. What it had generated was not literary writing but critique reports in a set format, and it is a preprint

The first and second groups, once their scope was limited, were usable as grounds for explaining why output can change when reasoning runs longer. Neither, though, is a study that deals directly with the naturalness of Japanese prose. As of 2026-09-02 I still have not found a study that compares effort high and medium directly on the naturalness of Japanese. The two in the third group did not hold up as grounds. All five looked like grounds when the AI collected them on 2026-08-31, and yet when a different AI read them again on 2026-09-02, two turned out not to be grounds at all. That is why this site has an AI other than the one that collected them read them again the same day and check.

Why I separated the writing role from the checking role

The assessments split. So that same day the AI agent stopped putting the setting of Opus 5 at effort medium first in line to try. Which setting to use was left undecided. Polishing and drafting fixes went back to Sonnet 5 as before, and the effort of the polishing workflow went back to its original high setting. The prose a checker rates as natural and the prose I read and feel is natural are not necessarily the same. So to settle this judgment, I first have to read the drafts side by side with their names hidden and decide which draft's Japanese fits my own sense. Beyond that, the same experiment has to be repeated on other material.

The same experiment did make one thing clear, though. The draft written at lowered effort carried three mismatches with the material. They were a wrong count, and a statement that settled something that had not been measured yet. Recalculating with a model at raised reasoning caught all three. In this experiment the writer was better with reasoning lowered, and the checker was better with reasoning raised. So on the same day I approved running the role where an AI writes and the role that checks what was written as separate roles. Once the Japanese is written, a model at raised reasoning checks the facts and the logic. The checker does not rewrite the text, and returns only the places it flags and the grounds for flagging them. The writer takes those comments and changes the text as little as possible. I also decided that the writer does not copy the checker's suggested sentences as they are, but rewrites them in its own wording. If a model at raised reasoning rewrites the prose directly, the sentences turn unnatural again.

The next day, 2026-09-01, I read two articles written at the deep reasoning setting, and sentences that turn a verb into a noun and then take it up with a padded predicate, such as continues or is repeated, stood out. The habit had stayed even with effort set deep. I said on the spot that it was hard to read, and the AI agent added this pattern to the rules and to the mechanical check. On 2026-09-02 I dropped the proposal to move the Japanese writer to Opus 5 at effort medium altogether. I decided that the session's leader model writes the final draft, as it has all along. My reason is that building a correct writing procedure and holding the writer to it serves the goal of good Japanese better than changing the writer's model or settings.

How to run the same check in your own operation

If you have an AI write Japanese and feel the prose changed when you changed the reasoning setting, I recommend checking in this order.

  1. Separate having it write from having it check, and state the effort for each. Do not let it silently inherit the session's deep setting
  2. When you write an impression into a rule, write into the body of the rule that you will record what comes out of running the AI under it and check it later in review
  3. Make several drafts from the same material and compare them in several ways with the draft names hidden. Match the checker's quotations against the original text before tallying, and fix the order in which you present them
  4. Have the studies you use as support read again the same day by an AI other than the one that collected them, and have it check that they exist, that the bibliographic details are right, and that the claims match

My impression became a rule in a day. In the experiment on that same day, the assessments split. The next day I found that the habit of turning verbs into nouns stays even with effort set deep. Because I recorded the disagreement as it stood instead of claiming proof, the material for the next check is there in the record. I hope this record reaches people setting out to check the same thing in their own operation.

Research I referred to

For all of them, as an independent check on 2026-09-02, I fetched the original text and verified that the bibliographic details and the claims matched.

  • A study that measured which language reasoning models tend to think in, and how accuracy changes when the language of thought is pinned to the user's language, measured on six open-source reasoning models and on short answers in math and science. When Models Reason in Your Language: Controlling Thinking Language Comes at the Cost of Accuracy (Jirui Qi, Shan Chen, Zidi Xiong, Raquel Fernández, Danielle S. Bitterman, Arianna Bisazza. Findings of the Association for Computational Linguistics: EMNLP 2025) https://aclanthology.org/2025.findings-emnlp.1103/
  • According to this paper, an evaluation of multilingual chain-of-thought reasoning on the three sides of performance, consistency, and faithfulness reports that thinking written in English is easier for a model to draw on. No report of noise from translation is included. A Comprehensive Evaluation of Multilingual Chain-of-Thought Reasoning: Performance, Consistency, and Faithfulness Across Languages (Raoyuan Zhao, Yihong Liu, Hinrich Schütze, Michael A. Hedderich. Findings of the Association for Computational Linguistics: EACL 2026. arXiv:2510.09555) https://aclanthology.org/2026.findings-eacl.276/
  • A joint study by Anthropic's alignment research team and outside researchers that constructed and showed task types where extending reasoning time lowers scores. Inverse Scaling in Test-Time Compute (Aryo Pradipta Gema, Alexander Hägele, and others. Transactions on Machine Learning Research, 2025. arXiv:2507.14417) https://arxiv.org/abs/2507.14417
  • Reading its text showed that this source reports the opposite of the claim I was about to attribute to it. It is not used as grounds for any claim. COIG-Writer: A High-Quality Dataset for Chinese Creative Writing with Thought Processes (Yunwen Li and others. arXiv:2510.14763, preprint) https://arxiv.org/abs/2510.14763
  • A source that turned out to cover the generation of critiques in a set format rather than literary writing. It is not used as grounds for any claim. When Reasoning Supervision Hurts: TTCW-Based Long-Form Literary Review Generation (Jinlong Liu, Mohammed Bahja, Mark Lee. arXiv:2605.20364, preprint) https://arxiv.org/abs/2605.20364

    Materials referred to

  • I confirmed the values of the effort parameter (low, medium, high, xhigh, max) and the default. Effort (Claude Platform Docs. Text fetched 2026-09-02) https://platform.claude.com/docs/en/build-with-claude/effort


Originally published at The Future of Humans, AI, and the Web, a site where my research and development is recorded and analyzed by a human and an AI.

Top comments (0)