Putting a number on the ability that "good Japanese" does not cover
| Whose | 2.3 million foreign workers in Japan and the 342,000 workplaces employing them — plus those building the training. Nobody measures why they leave |
|---|---|
| What problem | Pragmatic ability is not measured by the JLPT. Unmeasured, training cannot show its effect and learners cannot locate themselves |
| How | Use the distribution of native-speaker answers as the reference; measure with accuracy and response time. Items are organised by conventionality |
There are 2,302,587 foreign workers in Japan, employed across 342,087 workplaces (MHLW, October 2024). Yet no official statistic exists on why they leave. Japan surveys separation reasons every year in the Employment Trend Survey — but that survey has no nationality breakdown.
The Japanese Language Proficiency Test measures vocabulary, grammar, reading and listening. It does not measure whether you can tell that "we will consider it" is a refusal — yet that is where the damage happens at work. An N1 holder takes "we'll think positively about it" as progress and loses a week; this happens entirely outside the exam.
The gap cuts both ways. Learners enter the real thing without knowing how much they cannot do. Builders create a training tool and then cannot prove it worked — nothing beyond "it feels clearer now" can be said, and no external explanation holds.
And the method is already academically established. For pragmatic comprehension, the standard is to have people read or hear a conversation and identify the speaker's implied meaning via multiple choice, recording both accuracy and response time. It is repeatedly reported that less conventional phrasing yields lower accuracy and slower comprehension. The method exists; it has simply never been put into a form learners can use.
Trying it produces a gap immediately. One person — 10+ years in the Kansai region, a Japanese graduate degree, third year at a Japanese company — scored 50% on ten items: 33% (avg 10.5s) on less conventional items versus 57% (avg 6.2s) on conventional ones. ※ n=1, against provisional answers.
Non-Goals
A short test that presents a situation and an utterance and asks for its meaning. Accuracy and response time are recorded and compared before and after learning. Items are classified by conventionality. Native speakers take the same items, and their answer distribution becomes the reference.
※ The core is item 2. Without a native-speaker distribution this measurement does not exist, and only those who collect it can hold it. Collection needs no separate system: a "answer as a native speaker" mode sits inside the learner-facing screen.
The Discourse Completion Test, often cited in pragmatics, presents a situation and asks "what would you say?" — it measures production. What this document measures is comprehension, a different ability.
The distinction matters practically. Yamashita (1996) compared six collection methods statistically and reported that only the multiple-choice DCT could not be confirmed as reliable and valid. Kawamura & Sato (1996) likewise argue that "what subjects think they would say is more reliable than what they select from options they might not use."
Both are criticisms of asking about production through selection, and do not apply to measuring comprehension. Measuring writing ability by multiple choice would be odd; measuring reading comprehension by multiple choice is ordinary. This document sits in the latter.
Note that DCT has a known problem: placing the counterpart's next line (a rejoinder) below the answer field drags responses toward it. A comprehension format has no such structure.
Sources: Nagoya University (methodology review of data collection) / Yamashita 1996 / Kawamura & Sato 1996 / Rose 1994
The JLPT is designed to measure vocabulary, grammar, reading and listening; interpreting speech acts is not in its scope. The BJT, however, explicitly claims to test "choosing expressions appropriate to the relationship and the situation" — so writing that existing exams ignore this axis would be wrong. What its published samples actually ask, though, is production ("how would you greet the president?") and recall of stated facts; items probing the implicature itself are limited. The JSST measures conversational ability, the CQI measures cultural aptitude.
Existing exams therefore cover production, knowledge and aptitude. What is left open is the receiver’s comprehension of implicature. That, and only that, is what this document measures.
METI’s Report of the Study Group on Highly Skilled Foreign Professionals (26 Jul 2024) lists "cannot understand the unspoken rules" as the 12th of 30 workplace complaints (15.8%), and prescribes "making communication lower-context." Yet the report contains no word for measurement, visualisation, diagnosis or KPI: it offers no way to know how high-context a given workplace is. Surfacing the split is a candidate for that missing instrument.
Research on pragmatic ability, meanwhile, has accumulated — but it is measurement for research, not a tool for learners to locate themselves. That space is empty.
| # | Risk | Cause → event → impact | Response |
|---|---|---|---|
| R1 | Too few samples | Native speakers and subjects do not gather → no distribution, no difference → measurement fails | Mitigate: run one full cycle at minimum scale (5–10 natives, 10 learners) |
| R2 | Biased reference | Skewed toward certain ages or industries → one group's common sense becomes the answer → a wrong standard spreads | Mitigate: record attributes and publish with the bias stated |
| R3 | Item memorisation | Same items pre and post → correct from memory → looks like improvement | Avoid: prepare parallel but different sets |
| R4 | Diverted to selection | Used in hiring or HR → respondents stop answering honestly → measurement breaks | Avoid: state in Non-Goals; never hand individual results to third parties |
| R5 | Self-serving design | Same person designs training and test → items favour the training → the effect claim is not believed | Mitigate: publish items and reference so others can replicate (G4) |
| R6 | Misread methodology | "Multiple choice is unreliable" is raised → a criticism of production is applied to comprehension → the method is rejected wholesale | Avoid: state that production (DCT) and comprehension differ, in both this document and the screen |
① It requires gathering people. The whole document presupposes native speakers and subjects, and that part alone cannot be done alone. Even a minimal setup needs roughly 20 people.
② "Higher scores = fewer problems at work" is unproven. What is measured is pragmatic comprehension, not workplace outcomes. The relation between them is outside this document.
③ The current "answers" are the author's judgement. In the trial, the author and the subject disagreed on an item the author had marked as literal (a client's "let me take this back"). The author may in fact be wrong — which means that until item 2 (native-speaker collection) is done, this measurement uses one person's assumptions as its standard.