Conference
Whose Words Are These? Forensic Linguistics, Authorship Evidence, and the Courts in the Age of Large Language Models
المستخلص
For half a century, forensic linguistics has rested on a quietly powerful premise: that people
leave themselves in their language. Habits of lexis, syntax, spelling, and discourse
accumulate into an idiolect distinctive enough, in favourable conditions, to link a questioned
text to its author, and courts in many jurisdictions have admitted analysis built on that
premise in cases ranging from threatening letters to disputed confessions. Large language
models unsettle the premise at its root. When any literate person can produce fluent text
through a model, when a coerced author can be assisted, imitated, or replaced by a machine,
and when the tools marketed to detect machine text prove fragile under paraphrase and biased
against non-native writers of English, the evidential status of the written word requires re-
examination. This paper conducts that re-examination across the three disciplines the problem
now spans. From linguistics, we assess which components of idiolect survive machine
mediation and which dissolve. From computer science, we review the technical state of AI-
text detection, watermarking, and stylometric attribution under adversarial conditions, and
formalise the underlying inference problem. From literary and interpretive studies, we ask
what authorship itself now means when composition is distributed between human intention
and machine articulation, and what that shift does to legal categories, confession, threat,
defamation, plagiarism, built on unitary authorship. We argue that forensic authorship
analysis is not obsolete but must be rebuilt on likelihood-ratio reporting, validated error rates,
and explicit uncertainty about machine mediation, and that courts should treat current AI-text
detectors as investigative leads rather than evidence. Recommendations for practitioners,
courts, and researchers follow.
leave themselves in their language. Habits of lexis, syntax, spelling, and discourse
accumulate into an idiolect distinctive enough, in favourable conditions, to link a questioned
text to its author, and courts in many jurisdictions have admitted analysis built on that
premise in cases ranging from threatening letters to disputed confessions. Large language
models unsettle the premise at its root. When any literate person can produce fluent text
through a model, when a coerced author can be assisted, imitated, or replaced by a machine,
and when the tools marketed to detect machine text prove fragile under paraphrase and biased
against non-native writers of English, the evidential status of the written word requires re-
examination. This paper conducts that re-examination across the three disciplines the problem
now spans. From linguistics, we assess which components of idiolect survive machine
mediation and which dissolve. From computer science, we review the technical state of AI-
text detection, watermarking, and stylometric attribution under adversarial conditions, and
formalise the underlying inference problem. From literary and interpretive studies, we ask
what authorship itself now means when composition is distributed between human intention
and machine articulation, and what that shift does to legal categories, confession, threat,
defamation, plagiarism, built on unitary authorship. We argue that forensic authorship
analysis is not obsolete but must be rebuilt on likelihood-ratio reporting, validated error rates,
and explicit uncertainty about machine mediation, and that courts should treat current AI-text
detectors as investigative leads rather than evidence. Recommendations for practitioners,
courts, and researchers follow.
الكلمات المفتاحية
forensic linguistics
authorship attribution
large language models
AI-generated text
stylometry
expert evidence
admissibility


