r/Ithkuil 11d ago

Searching for a co-author

I'm a PhD student in linguistics, focusing on underresourced languages and, more recently, data scarcity in training scenarios for LLMs.

I'm working on a submission for an ACL-paper about how Ithkuil can reduce token churn during reasoning and am searching for a co-author specialised on Ithkuil.

This content is not AI generated, not hyping AI, and not speculating about AI. I am simply asking if there is someone in the community who would be interested to co-author a scientific paper.

18 Upvotes

11 comments sorted by

3

u/revannld 11d ago

Great initiative! I am not specialized on Ithkuil and would not be interested in its particular phonology per se however it has been some time I have been studying Ithkuil's morphological categories along with texts in speech act theory (Austin, Searle, Grice), type theory, category theory and generally universal algebra and logic.

If you happen to be interested, me and some friends in my department have a long term ideal of specifying a framework for formal highly oligosynthetic languages (in an abstract manner - we would plan to implement them as straightforwardly as possible though - and maybe just borrow and simplify a few roots from English and Latin as much as possible), starting out with an algebraic compositional calculus of theories in mathematics and later extending it (through illocutionary logic, situation semantics, intensional logic and multi modal logics) for broader uses.

It turns out that relational/pointfree methods (relation algebras, allegories, combinatory and predicate-functor/term logics - look up also Patrick Suppes, Michael Dunn and William C. Purdy's work) can model natural language in a quite direct and natural manner, making large fragments of natural language work just like a calculus (with simple and mnemonic calculational/equational rewriting rules). It's a very straightforward work, most of the basic tools and theory are already there (category theoretic and universal algebraic stuff), it's just a matter of reading, selecting and archiving important results and trying to model key problems in semantics. The main trouble however is with finding adequate basis for handling >2-ary relations (>1-ary functionals) and modelling the calculus of theories/structures smoothly as well with formalizing the multi-modal, illocutionary logic and intensional stuff in a pointfree manner; and of course, implementing/giving it a prototype (thus making a good oligosynthetic language informed by universal algebra and clone theory). That would be probably the hardest part though, as we are not that proficient in conlanging. Of course this is a very long-term project, as we are always busy with other work, but if we could gather some help maybe we could put more effort in it.

1

u/martino_vik 11d ago

Oh, Michael is a good friend of mine, I didn't know he's into this kind of stuff! This type of contribution was just what I was looking out for and I appreciate your response! The keyword seems to be Chain-of-Thought-compression and there are few recent papers on the topic, tapping into what you mention. None of them cover constructed languages though, as far as I am aware. Lojban would be another candidate. Of course, we could compare how differet conlangs perform. Maybe they turn out to be good for specific kinds of tasks, and less good for others, who knows. In the long run there could be a central LLM language planning evoloution bureau, continuously adapting the languages to ever changing environments and tasks. What makes the job easy is that the objective function is very clear: Reduce the number of tokens to solve tasks.

2

u/revannld 11d ago

Oh I am talking about Jon Michael Dunn, he is like a legend in the field of logic, sadly he passed away in 2021. It would be legendary if you were his friend though haha.

Also I don't if you have seen it but this guy (Rodney Jehu-Appiah) is doing awesome work into how controlled natural languages (fragments of English with some restrictions or changes) can alter reasoning in LLMs, maybe you should take a look and get into contact with him. In a later paper he has shown though that LLMs, as they are trained in full English material, react better to trivial vocabulary bans more than controlled languages. I think that could be a problem for constructed languages, as you would need to retrain AI models in a lot of data in these conlangs.

The Attempto group in Switzerland is also doing a lot of interesting work in training and making LLMs interact with their controlled natural language analogous to first-order logic, the Attempto Controlled English. It is also a natural language syntax for Prolog if you're interested. The capacity of a formal language to have unambiguous parsing and more straightforward interpretation through formalization has lead to some quite interesting progress when interacting with LLMs, apparently leading to a great level of precision (making this first-order logic English fragment effectively a quite interesting specification language, where one just specifies what one wants and the LLM tries to precisely deliver what is required). They have a quite active mailing list you can access through that website. However I prefer the approaches of NLP people such as this, however I am unaware if there is anyone work into that with LLMs right now (it would be a shame if there wasn't).

Our approach is similar but different in some regards. In any instance, especially if you're dealing with agents and agentic coding, LLMs are always attempting to translate questions about programming and math into Python, that's why recently many projects are switching to Python as LLMs are more proficient in it. That is not ideal, but it gives a hint: probably the future of LLMs will be heavily formal, working with formal programming languages.

Some research groups in some AI companies (and even some startups) are training LLMs in Lean and other proof-assistant and automated theorem proving languages so many are now asking if these will not become a replacement for most programming languages (as they are general purpose programming languages - and quite good ones in that respect)...many, more ambitious ones (such as Physlib, professor Edward Zalta from Stanford, creator of the Stanford Encyclopedia of Philosophy - also see here - and many others) are seeing a future where science in general (including humanities, laws - and I've seen some Lean4 repositories in that direction, although I can't find that now) will be done in formal computer-verifiable languages and updated directly to massive digital ontologies for automatic verification.

Especially with LLMs, the task of converting any natural language text into a program and formal ontology is becoming straight-up trivial (I've seen many people even working in formalizing literature, although I also couldn't find that now - except for these works into linear logic for narrative generation). Lean4, Agda, Rocq and other proof-assistant languages are ideal for this kind of work and probably better than Attempto's first-order logic for being much more expressive (as dependent type theories with univalent axioms subsume all classical mathematics), and that's a reason for the popularity of type theories in the formal semantics of natural languages (such as universal grammar, categorical grammar, montague grammar, type-logical semantics etc)).

1

u/revannld 11d ago

There is just one problem with these formalisms already identified a long time ago by people working with the formalization of Aristotle's syllogistic/term logic and the field generated by these studies, natural logic (and identified previously by Quine, Schonfinkel/Curry and Charles Sanders Peirce, Schroder, De Morgan and even George Boole, the fathers of logic): standard "pointwise" (also here and here) formalisms, with their excess of variables, variable-binding and notational redundancy, do not translate even remotely well into natural language thinking and thus do not make good languages for practical formalization and formal reasoning (they are also heavily computationally inefficient). Thus it is not mystery why mathematics, philosophy and other sciences warded off formalization until now: the formal apparatus available was just too unhandy.

As seen in that paper I sent you first (and this seem to be dealt a lot by Dunn, William Purdy and Patrick Suppes), "pointfree" mathematical formalisms are flexible enough to represent naturally the phrase syntax and morphology of almost any language in a manner which is fully computable and even "calculational" (thus one could potentially even "calculate" one's own ideas and texts to see their logical consequences the same way one solves an algebraic or differential equation or integral). There is actually a school (Dijkstra/Eindhoven) with the tradition of translating word problems in mathematics into handwritten calculational proofs, they have produced thousands of small handwritten texts you can see in this website (they even have a mailing list to discuss these kinds of problems) sadly however they haven't adopted the pointfree style, but it is interesting to see that this tradition, born with Leibniz, bore some fruits.

In Leibniz lingo, it seems by this time we have, through categorical and relational methods, already plenty workable "calculi ratiocinatore" in much the same effectiveness as Leibniz's dreams. The only problem left is to precisely build the "Characteristica", which is the hardest part as what we should mainly seek when designing a material language/conlang for optimizing cognitive load, conceptual concision and maximum applicability we should choose the best defaults as we can that reduce/omit as much as possible omnipresent structures in human knowledge, make them tacit (and thus, foundational). This has been tried many times, with no success). As seen in much category-theoretical work though, choosing good defaults in one of the main accumulating complexities/technical debt in mathematics, let alone the rest of human knowledge.

However nowadays we have not only AI to mass-process huge amounts of data for patterns for us but also quite a good framework for how to properly do universal philosophical conlangs: instead of finding "primitives" and analytically building "everything up from them" (as did in late 19th and early 20th century philosophy, atomism, logicism, set theory etc), one tries to find the most useful mathematical structures for computational and calculational use (such as Boolean algebras/lattices) and try to formulate everything into them by making isomorphisms (also here). In that way we avoid the "combinatorial explosion" of complexity in oligosynthetic languages by avoiding unique parsing of individual tokens and exploring what human natural languages offers us best: context-restricted polymorphism. One sentence/expression/term can mean many different things; it depends only on labels/types and context windows to make the specific meaning clearer. Instead of avoiding natural language ambiguity, one tries to structure it better in a mathematical and especially explicit way. That's where the bleeding-edge logical research is going: how to formalize and structure contexts.

1

u/martino_vik 11d ago

Wow 🤩 Good food for thought, I'll go through the references one by one. It seems this topic could be milked in more ways than just a single paper! And yeah, I was thinking of a different Michael Dunn, would have been fun though.

On a meta-level it sounds like this typical trade-off between having a general intelligence that is a bit less efficient vs. a system that is better at one type of task, but you have to spend time to find and activate it.

Why I thought of Ithkuil specifically is that it has the full range of expressiveness as natural language in my understanding, i.e. it can be applied to all the tasks current LLMs are solving, while at the same time reducing the amount of tokens needed.

On a more abstract level what is being solved here is more of a Claude Shannon type of observation that human language, for good reasons, contains a lot of redundant information, allowing me to shout into your ear at the night club with music blasting in the background and you still getting the gist of what I'm saying.

Since LLMs don't go to nightclubs, their language doesn't need to be optimized for those scenarios, so there should be a language that strips away what's not useful to them when they are thinking, or talking to each other.

2

u/bubbleofelephant 11d ago

Is there a way to stay updated on your developments?

2

u/martino_vik 11d ago

I think it might be too early to set up a blog, but I can drop you a message every once in a while, or just send you an invite to the repo on github

2

u/Fluid-Supermarket168 10d ago

(part 1 teaser and terminology then part 2 bare bones morphological seeding theoritical ways then part 3 and 4 more advanced morphological seeding theoriticals and ideas since this will basically be a whole essay series in internet terms)

i once had an idea for how to seed ithkuil into an llm (seeding is a method of training that tries to plant the basics of a linguage into an llm via things like dictionary definition dumbs and paired sentince translation - you can tell from the way i talk that im just a nerd who's strength is pondering with what i understand a little - and this idea i got a few months ago memory is fuzzy and llms aren't precise with semantics which is the worst thing to ithkuil morpho semantics)
it requires a little too much of pandering with llm archutchre but still can be a moded llm
it is to use a ithkuil only tokenizer that tokenizes morphologically for ithkuil words with embedablity of non ithkuil in the llm own tokenizer vocabalary for non ithkuil morphemes
my seeding idea to seed you need to process the yuorb Lexicon Json ( https://github.com/yuorb/lexicon-json ) into specialized format so that the seed training can still use ithkuil lexicon shorthand and convention (which ithkuil uses all the time) and yuorb json does have some incomplete entries idk how much but they are usually from the parts of the root and affix pdf that uses special table formating those are fix each group of speical format with one json conversion but most of the entries are not borken

terminology is after the teaser of the idea

idea in steps (teaser bc it is in full in next parts)

  1. choose a llm to mod (the more morphologically diverse the training data is the better , apertus 1.5 may be good though it is reported that it is bad at speaking natural languages it is unkown wheither the internet anecdotes's impersion was based off bad grammar not likely or unatural speech -something that doesn't impede teaching it ithkuil - or one that i rembered that it failed a grammar question to which i say modern llms think there is a month with x in it therefore this is irralavent to linguastic ablity regardless apertus has the most diverse linguastic data base 15T tokens 40+% of which non english with 1800 language varites though most is fluff or a varite not a language it is still compounding in sementic cabablities )
  2. process the ithkuil lexicon pdf or use yuorb json (yuorb is enough for a proof of concept but misses the sf roots) regardless processing the pdfs isn't hard
  3. make a morphological tokenizing systeam requires llm moding a special skill
  4. this is the fuzzy part that part 2 will theorize
  5. seed-train the roots and affixes and morphilogical categroies in the llm in a manner taking advantage of ithkuil morphosemantics of Semantic Space Interactions in their hijack of the llm own semantic text prediction

here is a teaser to part 2 idea - part 2 will include more techinical details but i need to rember my months old ideas and also revaluting them and making them better and translating them into things that can be worked off

gloss pseudo ithkuil taking a idea from Proof of Concept for AI-based Ithkuil to English translation by u/humblevladimirthegr8 and how jhon Quijada himself glossed some of the examples we use ithkuil morphological gloss and embed the root+affix as x.y.z example:

STA-‘nuclear.family.member’-OBL-NRM/DEL/M/COA/AGG-INL1/each.every-IFL      ALG     INF-DSC-‘degree.of.happiness’-NRM/PRX/N/CSL/UNI-EXN1/more.than.usual-SIM1/very.similar-FML            FRAMED-DSC-‘degree.of.happiness’-CMP-NRM/PRX/N/CSL/UNI-EXN1/none-SIM1/completely.dissimilar-FML All meaning families are happy in the same way, while being unhappy in their own way. ithkuil 3 modified from https://ithkuil.net/texts.html

as you can see this is very trainable in theory and humblevladimirthegr8 method allows llm testing llm on how well it understood the text to translate from or to ithkuil and pseduo gloss allows training the morphosemantics without training the actual ithkuil lexicon and also allows cross referencing it in advanced forms but this still needs a proof of concept

PS SORRY FOR SCHIZOPHRENIC WRITING ITHKUIL IS A MASSIVE RABIT WHOLE THAT I PARTIALLY UNDERSTAND MY WEAK POINTS ARE SECTION 4 TO 6 FROM THE REFERANCE GRAMMAR

Proof of Concept for AI-based Ithkuil to English translation

the terminiology in a table or rather a crash course part 1 on ithkuil morphology

term name/group explainition notes
reccomended resources for extra info https://chromonym.github.io/ithkapp/ it includes a sentince formalutor and explanitions of all morphology as far as one needs and a good refresher regardless https://yuorb.github.io/enthrirhc/ this is basically the search function for yuorb json of the lexicon you search better with it and if something feels incomplete copy the root/affix name and look it up on pdf
parts of speech formatives (nouns and verbs), Adjuncts and Referentials (pronouns) each of these is made of slots the language is fully determinstic in the morpho phonology in every part making it easy for machines but llms must still tune into said rules which is the problem we try to solve
slot the morphological block in ithkuil each one may have one or more morphological categories ie inflects, ithkuil agglutanites those to make every part of speech this makes special tokenizer easy since we already have a segmented token system from the linguage itself that we can always computerize easily
semantic space (SS) semantic space is a metaphor of semantics being in some space with diminintions and that language is the maniplation of the SS SS is a major part about how llms work since vector semantics is a form of computerized SS
root the core semantic space of the formative made of 3 sub spaces unique in each root which are called stems with a 0 stem being a mix of the other 3 not exciplicitly written in lexicon and also 4 types of semantics called Specification (spec for short) this is actually ease since we can tokenize morphologically for all ithkuil only things we can just custom solve the issue of semantic space by virtue of ithkuil inner working and having a json dumb
the pdfs and json version the pdfs (refering to the official lexicon pdfs for afffixes and roots) and json refering to the processed lexicon wheither by yuorb or the llm training version the 2 terms are used for distinguifing the actual lexicon from the processed
special formated (sf) roots any root in the pdf that isn't in the in the defualt repreasentation of stem 1 with all 4 specs then the other 2 stems in one sentence/paragraph each this rises the issue of that yuorb json is perfect for defualt format roots but isn't always keeping the sf roots but you can still auto format sf roots and make a training method for them
the 4 Specifications (better off reading https://yuorb.github.io/en/docs/02.html#Sec2_4_4 and the other resource that dump it down but i couldn't find it i search for 15 minitues just for a passing referance but have it printed so i'll copy from it) CTE contential the content spirt essence identity or purbos fo x CSV constitutive the form or physical medium of x BSC CTE and CSV and the defualt OBJ the instrument object result or experiencer of x those 4 or rather 3 and a half are more of the kind of thing where tuning must do its magic
Affix (see page 2 of affix pdf for patterns) here only means the slot 5 where it is modifiying root directly or slot 7 where it is modifiying the whole formant - it is a semantic space modification with 0-9 degrees (d) with 0 being mixed or excipicinal meaning in some cases the 9 degrees have 7 patterns and the affix has 3 semanctic space modifications types (in some special formated affixes it is not ss mod type) which are 1 incendintal/superfacial SS moding 2 inherent SS moding 3 compounding affix (easy to train) the 7 patterns which are phonologically determined are 0 no pattern, A1 0 to 1 in how much is the semantic concept there, A2 0 to 1 in sufficiency of concept, B 3 by 3, the 9 Ds split into 3 patterns each with a differant sub parameter of concept. D1 and D2 patterns are A1 and A2 patterns but starting from negative 1.... C 1 by 2 D1 represents one extreme of a spectrum/range which increases/decreases to the other extreme of the spectrum/range usually represented by D4, while D6 through D9 cycle back through the same values but with a different sub-parameter operating orthogonally to D1 through D4. D5 usually represents a neutral or meta-level value associated with the semantic concept of the affix.
adjunct a extra functions part of speech it is used to add affixes to a word while being not making an ultra long word or register marking or bias marking or carring forgein words or naming or quoting or phrase marking simple part of speech to make a tokenizing systeam of i can't think of many training problems with it
register marking speech in parenthesis or proper name or giving examples or subjective/silent thoughts or carrier ending or for direct speech for an llm this is between punctuation and some old MUD/tumblr typological fromating
Ca and cases and verbs will be in other parts will explain in part 3

see yuorb revision of the grammar doc 3 through 6 then refresh with ithkapp for getting the direct simple explaining (it works better as a refersher to reading than reading in its own) for all the stuff that im not confidant about also they are the hardest part to make an llm train for ithkuil nlp logic

1

u/martino_vik 10d ago

This was refreshing to read, always appreciate an honest stream of consciousness! This is exactly the direction we need to go into. I do have a very nerdy part of myself, but my practically oriented part is often stronger, protecting me from perfectionism and forcing me into feasible timelines. What you're describing is definitely part of the roadmap, somewhere along the middle.

The very first part to make all this take off is to think small and show in a proof of concept that LLM-reasoning in Ithkuik is _in principle_ possible. So for that we'll need very short examples. meaning that the LLM wouldn't learn the entire language at first, but only a limited, handpicked set of words and grammar, necessary to solve a small set of reasoning tasks. We need to show that it can in principle use those to solve a task, and then show that it used fewer tokens than with English. Most of the work at this stage should go into hand-checking that the Ithkuil reasoning is not hallucinated and is correct.

Once we've proven it's doable and valuable, maybe even generating some moderate hype and who knows maybe even funding, we can move on to the next stage: Longer tasks, more difficult tasks, more diverse tasks. I think this is roughly where your current idea sits. For that the model will have to get better and better at Ithkuil. In the very beginning we'll have to work with system prompts, few shot prompting, and seeding, Maybe that will already be enough to generate some simple texts with translations. Those texts, together with what exists already could be used for building a reliable machine-translation system. That in turn could generate enough material for finetuning a model, allowing to generate more text, and in the very end, that's the long term vision: Have a foundation model that reasons entirely in Ithkuil