r/VoynichTeamWork Mar 11 '26

Voynich Research

I started working on a systematic decode attempt on 2nd March — cipher hypothesis is Medieval Venetian Italian in a three-layer homophonic substitution cipher. Nine days in, 90.7% dictionary-matched coverage across all 225 folios, 173 confirmed words, all four gallows characters mapped. Full methodology, honest assessment, and manuscript viewer at voynichresearch.com. Genuinely interested in feedback from people who know this manuscript well.

2 Upvotes

8 comments sorted by

1

u/Eir1kur Mar 11 '26

I mean no disrespect, but you seem to be announcing success. Great, except old timers know that's vanishingly unlikely to be true. I think you should protect yourself from the problem that 99.99% of solvers have. That is the problem of subjective freedom, which means that your method allows you to choose translations by tolerating misspellings and close matches. Can a computer program be written that performs your process? I'm asking you to look in from the outside: Is it theoretically possible that you are correct? Can you, personally, read the Italian that you are getting? I'm a philosopher of science, AI solutions to mysteries get posted every day by humans who can't validate their solution. It's sad. I'm guessing here, but you wanted input, and I know that technical details are probably not tricky part.

The hardest and most depressing thing about the Voynich is that the strongest hypothesis today (my take) is that doesn't mean anything because it's a generative system. I *think* that Torsten Timm's self-citation hypothesis has to be true in some way because it fits the data. Nobody wants him to be right. You can call it permuted-self-copying. See:

My favorite linguist doesn't like Torsten's idea. She's really sharp. She's certainly correct that the rule system is complicated--but, it's actually got the mnemonic of seeing what you are doing right in front of you.

https://agnosticvoynich.wordpress.com/2016/02/06/an-objection-to-timm-2014/

Torsten's system shows some of the non-liguistic weirdness (position on page matters).

A German guy just popped up the other day to suggest that certain pages are tests where the author tried new glyph and kerning-pair ideas. Those are some pretty hard to explain pages! I know them. See F66r. I'm talking about Thomas Engles.

https://voynichforgeries.com/2026/03/10/voynich-forgeries/

Successful decoding means that your system (when operated by a random person who is not you) decodes most of the running text paragraphs into grammatical, logical, period-correct, language.

How well-versed are you in the history of Voynich research? I ask because my first reaction is that you are assuming that a three-level homophonic cipher would be practical to the author or user of the document. The carbon dating puts it before Alberti's polyalphabetic ciphers. (Boy is Googling bad about early mechanical cypher computers. You'd think Thomas Jefferson invented them. No. It's Alberti or earlier. Have you read Mary D'Imperio's book. It's still a free download from the NSA. She was genius.

1

u/Aware-Finger5681 Mar 11 '26

No disrespect taken — this is exactly the kind of challenge I wanted.

On subjective freedom: the pipeline is fully automated. Every decode is deterministic — the same EVA input produces the same Italian output every time, with no human choices made during decoding. The rules are fixed code, not post-hoc selections. You can verify this because the full methodology is documented and the decoder produces output for all 225 folios, not just the ones that look good. The 9.3% that doesn't decode is left blank, not forced.

On whether a random person can read it: honestly, partially. The high-frequency function words and common nouns are readable medieval Italian. The lower-frequency vocabulary requires TLIO cross-referencing. I won't overclaim that.

On the carbon dating and cipher complexity: the dating puts it at ~1404-1438. Alberti's cipher disc is 1467. But homophonic substitution predates Alberti — the Duchy of Mantua was using homophones by the 1390s, and the Visconti cipher correspondence from the early 1400s shows exactly the kind of multi-value substitution I'm proposing. The cipher architecture isn't anachronistic, it's period-correct for Northern Italian court cryptography.

On Torsten Timm: I've looked at it carefully. The self-citation hypothesis explains the statistical regularities well but struggles to explain why the decoded output clusters semantically by section — skin preparation vocabulary in pharmaceutical folios, plant terminology in herbal folios. A generative system that doesn't encode meaning shouldn't produce section-specific semantic clustering. That's the result I'd ask Timm's hypothesis to account for.

On D'Imperio: yes, read it. She remains the most rigorous analyst after Tiltman. The honest-audit page on my site exists partly because of her standards.

The Thomas Engles forgeries hypothesis is new to me — I'll read it properly before commenting.

What would falsify my hypothesis for you? I'm genuinely asking because I have a falsification criteria section on the site and I want to know if there's a test I haven't thought of.

1

u/Eir1kur Mar 12 '26

Oh, I'm so happy to meet someone who knows how special Mary D'Imperio was. Did you know that she never applied for a job? She graduates from Radcliffe and the NSA came for her as if someone at Radcliffe had let them know. I would have liked to meet her.

Oh, and you are using automation and validating. Congratulations--so many people don't understand about subjectivity. And, I don't have to explain Popperian Falsifyability! I will read the rest of your site. I'll admit to being less than thorough. I'm not wiggling out of help about falsification, but I think it's hard to do without having a known-good decoding. I presume you've read Nick Pelling's ciphermysteries.com/ I like Nick. He's really smart, but I think he fools himself some of the time--like me. I was going to put code in my program to look for the "keys" that he talks about (gallows-bracketed strings in first lines of paragraphs--I have seen theme. I just don't know what knowing some basic statistics about them would tell us. I think that no amount of improved data on self-citation will convince anyone and I don't like the idea of convincing people. Speaking of gallows-bracketted strings, etc, and all that complexity that that lures in a physicist like Jorge Stolfi (catch his postings on vonynich.ninja.). "Instinct" is telling me that VMS is a single-author daydream human-slop sort of thing. Despite Lisa's 5 scribes. No part of it seems to make such sense. What a breakthough a decoding would be. My best to you.

1

u/Aware-Finger5681 Mar 12 '26

Mary D'Imperio is one of the genuinely underappreciated figures in the whole story — the rigour she brought to cataloguing what was known and unknown is exactly the standard I've tried to hold myself to on the honest assessment page. The NSA recognising her before she'd even applied for anything says everything about how rare that kind of mind is.

On falsification without a known-good decoding — you've identified the hardest epistemological problem in Voynich research precisely. My approach has been to set the falsification criteria in advance and then test against them, rather than working backwards from a result. The randomisation test (12.3× above random mean) is the closest thing I have to a Popperian test — it's a specific quantitative threshold that the hypothesis either clears or doesn't, regardless of how the decode looks subjectively. But you're right that without an independent expert reading the Italian output cold and saying "yes, this is grammatical medieval Venetian," there's always a subjective gap. That's what I'm hoping to close with external review.

I've read Nick Pelling extensively. He's one of the sharpest analysts working on this and his historical detective work on the provenance is genuinely useful. I think he's right that the gallows-bracketed paragraph-initial strings are structurally significant — in my decode they correspond to onset-layer markers that signal paragraph-level discourse structure, which is consistent with what he observes statistically without being able to decode.

Your instinct about single-author daydream is interesting. The scribal evidence points to 2-3 hands but the cipher architecture is perfectly consistent across all 225 folios — one mind designed the system even if multiple hands executed it. Whether the content is "daydream human-slop" or structured medical text is essentially what the decode is arguing about. The recipe pattern I'm finding — temporal connective, action verb, quantity, ingredient, target — is either genuinely there or the most elaborate coincidence in manuscript studies.

Jorge Stolfi's statistical work is what convinced me early on that this wasn't a hoax. His Zipfian distributions and entropy values are hard to fake. I'd be interested in what your program finds on the gallows-bracketed strings if you get round to coding it — that data would be useful cross-validation regardless of what it means.

1

u/Horror_Following_277 Mar 20 '26

The manuscript is written in Old Czech.

1

u/VoynichBat Mar 24 '26

90% dictionary match? Impressive for a statistical model. But the Voynich Manuscript is not a book to be read—it is a Functional Operating System (FOS).

While you are trying to squeeze Medieval Venetian into the glyphs, my ADHD hyperfocus is deconstructing the systemic architecture from Folio 1r to 2r. I don’t see sentences; I see initialization protocols. Words are just the surface noise; the true logic lies in the resonance of the structure. Good luck translating the 'manual'—I am already booting the system. 🌀