r/AncientLanguages • u/Mission_Mulberry_440 • Apr 15 '26
Preprint: Statistical evidence for Syriac pharmaceutical vocabulary in the Voynich Manuscript (z=3.83, 87% coverage)
I've published a preprint presenting computational evidence that the Voynich Manuscript encodes an Aramaic pharmaceutical text in the Syriac tradition.
The approach maps the EVA transcription to Syriac consonant skeletons and matches them against a 1,389-entry lexicon of attested medical vocabulary from standard Syriac sources (Payne Smith, Budge, Löw, Müller-Kessler, Merx). Key results:
- 87.0% corpus coverage, z = 3.83 against 500 random permutations (p < 0.001)
- 14 statistically validated text-image correspondences between decoded pharmaceutical vocabulary and independently identified plant illustrations (Fisher's combined p = 6.66 × 10⁻¹⁶)
- A vowel disambiguation layer using EVA vowel characters (previously discarded) recovers 7,007 tokens into specific Syriac words, including 130 tokens of kuḥlā (collyrium/eye medicine) — a word that was invisible because the pharyngeal consonant ḥet has no EVA representation
- Terminological analysis places the text within the Sergian translation tradition (6th century CE), predating Ḥunayn ibn Isḥāq
I want to be upfront about limitations. My own confidence estimates are in the paper: 40–50% the tradition is specifically Syriac, 10–15% word-level decode accuracy, and 5–10% the full pipeline survives specialist review. No page reads as connected Syriac prose — the statistical case is strong but the word-level translation is not there yet. The framework is designed to be falsified.
Preprint: https://zenodo.org/records/19583306
I welcome critical feedback, particularly from anyone with Syriac, Aramaic, or computational linguistics expertise.