Findings

Measurements too small for a paper and too specific for an essay. Each one states what it found, on what data, and what it does not support. The caveats live here with the finding rather than in a reply somewhere, because a chart travels and its footnotes do not.

Register drift in AI-signed code review

Where AI-signed prose sits on the grammatical axis between US-located and India-located professional English, and which way each model is moving. Drawn from 153,448 GitHub comments, 2022 to 2026, parsed for construction rather than vocabulary and matched on repository and year.

Figure 1 Position on the US → India axis, 2025 → 2026 Q3 Each metric rescaled so the US human baseline sits at 0, at the bottom, and the India human baseline at 100 above it, the same orientation as Figure 2. An arrow runs from a cohort's 2025 position to its 2026 Q3 position; the small mid-dot is 2026 H1. Above 100 means more of the construction than the Indian baseline itself.
Figure 2 Raw trajectories against the human baseline bands Native units per panel. Shaded bands are the pre-2025 human baselines ±2 standard errors; dashed lines where no interval was computed. Open markers flag cells under 100 texts. Where the US baseline is the higher number (subordination, sentence length) the y-axis is reversed, so the India side reads upward in every panel.

Reading and caveats

The four Figure 1 metrics are the constructions that most separate India-located from US-located professional GitHub prose in 2022–24 (all p < 10−30, repo- and year-matched samples of ~4,800 comments per side): verbless status sentences, low clause subordination, triple-noun compounds and please/kindly politeness marking.

Claude-signed text enters 2025 past the India baseline on all four at once and drifts back toward the human range through 2026. Copilot shares the noun compression but keeps US-style subordination and has converged to US politeness. A third cohort of 370 AI-signed comments whose signature names no model (chiefly Claude Code PR footers with the link text stripped at harvest) is excluded from the charts: its register does not track the Claude cohort, and whether that reflects different authorship or human editing of generated text cannot be determined from the corpus.

Limits: the 2025 cells rest on 47–57 texts each and the Copilot 2026 cells on ~30, so read levels loosely and trends cautiously. AI-signed text is largely PR-body prose while the baselines are conversation comments; that genre gap inflates verblessness and noun stacking for every AI cohort, though it cannot manufacture the politeness or subordination signatures. Parsing is spaCy en_core_web_sm; the parser is trained on US-style English, which may flatten Indian-text parses slightly.

Data table — every value plotted

Unpublished note, August 2026. Not pre-registered. It measures how text is built, not who wrote it, and it makes no claim about any group of people.