THE SMALLEST UNIT OF AN IDEA - Minimal semantic units between prose, citation and computation - Anto Lloveras · LAPIEZA-LAB, Madrid - semiotics; natural-language processing; knowledge graphs; ontology design; digital humanities





Digital systems increasingly divide scholarship into fragments, but computational fragments are not necessarily conceptual units. A search engine retrieves a passage, a language model processes tokens, an annotation system points to a sentence and a knowledge graph may extract a statement. None of these boundaries guarantees that the resulting object can stand intellectually on its own. The more useful question is semantic rather than textual: what is the smallest configuration that still contains enough structure to be recognized, interpreted, cited and revised as the same idea? Minimality does not mean brevity. A two-word expression may be too ambiguous to function independently, while a longer formulation may contain exactly the structure needed to preserve a distinction. A viable semantic unit requires a recognizable label or formulation, a concise explanation, provenance, contextual occurrences, relations and some way of registering change. Removing too much makes the object unintelligible; adding too much turns it back into a document. The interesting threshold lies between those scales. This intermediate object also has to negotiate a tension between human interpretation and machine processing. Formal ontologies achieve precision by restricting meaning, while theoretical and humanistic concepts often remain productive because their boundaries can move. Complete closure improves inference but may destroy the ambiguity through which concepts travel between disciplines. Total openness has the opposite problem: a term becomes impossible to distinguish from neighboring ideas. A minimal semantic unit should therefore preserve bounded ambiguity—enough structure for identification, enough openness for reinterpretation. Once such a unit is addressable, it acquires several simultaneous bodies. It has a rhetorical body through which it persuades; a documentary body through which it is published; an institutional body through which it is classified; a computational body through which it is indexed and retrieved; and a circulatory body through which it returns in new contexts. This hybrid condition matters more than the exact length of the unit. The smallest viable unit of knowledge is not simply the shortest phrase a system can isolate. It is the smallest formation able to retain a recognizable intellectual identity while moving across media, institutions and technical environments. Designing that threshold carefully may become essential as scholarship is increasingly read by machines at scales smaller than the document.



Briet, S. (1951) Qu’est-ce que la documentation?
Bucur, C.-I., Kuhn, T. and Ceolin, D. (2020) ‘A Unified Nanopublication Model’, arXiv:2006.06348.
Peirce, C. S. (1931–1958) Collected Papers of Charles Sanders Peirce.
Sennrich, R., Haddow, B. and Birch, A. (2016) ‘Neural Machine Translation of Rare Words with Subword Units’, ACL.
Schmidt, C. W. et al. (2024) ‘Tokenization Is More Than Compression’, arXiv:2402.18376.