A claim circulating online that the AI service Astra scored 99.9 points on a "human vs AI discrimination test" โ an evaluation measuring whether judges can tell machine output from human output โ has been held up as proof that the system is all but indistinguishable from a person. This desk reviewed 12 primary sources connected to the claim and found the headline figure is real, but the sweeping interpretation attached to it is not.
The claim, as it has spread through Korean AI communities, presents the 99.9-point result as a near-flawless pass: that in a head-to-head test of human discernment, judges mistook Astra for a human virtually every time. Posts sharing the figure treat it as a landmark for the service's conversational realism.
The 12 primary sources do support that a 99.9-point figure tied to Astra was reported in connection with the discernment test. What the materials do not establish is the test protocol behind that number โ the judge pool, the scoring rubric, or the conditions under which responses were compared. A point score in this style of evaluation is not the same as a measured rate of human misidentification, and the reviewed materials do not document independent replication of the result.
In short, the number exists; the meaning projected onto it does not follow from the evidence at hand.
As Korean AI developers lean on benchmark-style results to market the human-likeness of their LLM products, raw scores travel far faster than the methodology behind them. Figures like this one can shape consumer expectations and investor narratives well before any third party has examined how the test was run. The gap between a score and its interpretation is precisely where this claim comes apart.
On the materials reviewed, the reported score checks out while the implied near-human indistinguishability remains unproven โ a finding we rate Partly True.