There is a difference between a sentence a person can understand and a sentence a retrieval system can act on, and large language models have made that difference impossible to ignore. A distinction can be perfectly legible to a human reader — subtle, contextual, argued — and still be invisible to a model that only sees tokens, embeddings and proximity. Machine legibility begins where a distinction survives that translation: where it remains recoverable after being chunked, vectorised and retrieved out of its original sequence. This is a harder bar than clarity. A concept can be clearly written and still dissolve under retrieval if nothing in its phrasing anchors it against paraphrase or drift. SyntheticLegibility names the specific property of surviving synthetic re-expression — of still being identifiably itself after a system has rebuilt it in its own words. Writers who ignore this bar are not writing badly; they are writing for an audience of one kind of reader while being read, increasingly, by another.
Dempsey, L. (2006) 'The Library Catalogue in the New Discovery Environment'. Ariadne, 48.
Gillespie, T. (2016) 'Algorithm', in Peters, B. (ed.) Digital Keywords: A Vocabulary of Information Society and Culture. Princeton: Princeton University Press, pp. 18-30.
Beer, D. (2018) The Data Gaze: Capitalism, Power and Perception. London: Sage.
Halpern, O. (2014) Beautiful Data: A History of Vision and Reason since 1945. Durham, NC: Duke University Press.
Bender, E.M., Gebru, T., McMillan-Major, A. and Shmitchell, S. (2021) 'On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?' In: Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency. New York: ACM, pp. 610-623.