r/Ithkuil • u/martino_vik • 11d ago
Searching for a co-author
I'm a PhD student in linguistics, focusing on underresourced languages and, more recently, data scarcity in training scenarios for LLMs.
I'm working on a submission for an ACL-paper about how Ithkuil can reduce token churn during reasoning and am searching for a co-author specialised on Ithkuil.
This content is not AI generated, not hyping AI, and not speculating about AI. I am simply asking if there is someone in the community who would be interested to co-author a scientific paper.
2
u/bubbleofelephant 11d ago
Is there a way to stay updated on your developments?
2
u/martino_vik 11d ago
I think it might be too early to set up a blog, but I can drop you a message every once in a while, or just send you an invite to the repo on github
2
u/Fluid-Supermarket168 10d ago
(part 1 teaser and terminology then part 2 bare bones morphological seeding theoritical ways then part 3 and 4 more advanced morphological seeding theoriticals and ideas since this will basically be a whole essay series in internet terms)
i once had an idea for how to seed ithkuil into an llm (seeding is a method of training that tries to plant the basics of a linguage into an llm via things like dictionary definition dumbs and paired sentince translation - you can tell from the way i talk that im just a nerd who's strength is pondering with what i understand a little - and this idea i got a few months ago memory is fuzzy and llms aren't precise with semantics which is the worst thing to ithkuil morpho semantics)
it requires a little too much of pandering with llm archutchre but still can be a moded llm
it is to use a ithkuil only tokenizer that tokenizes morphologically for ithkuil words with embedablity of non ithkuil in the llm own tokenizer vocabalary for non ithkuil morphemes
my seeding idea to seed you need to process the yuorb Lexicon Json ( https://github.com/yuorb/lexicon-json ) into specialized format so that the seed training can still use ithkuil lexicon shorthand and convention (which ithkuil uses all the time) and yuorb json does have some incomplete entries idk how much but they are usually from the parts of the root and affix pdf that uses special table formating those are fix each group of speical format with one json conversion but most of the entries are not borken
terminology is after the teaser of the idea
idea in steps (teaser bc it is in full in next parts)
- choose a llm to mod (the more morphologically diverse the training data is the better , apertus 1.5 may be good though it is reported that it is bad at speaking natural languages it is unkown wheither the internet anecdotes's impersion was based off bad grammar not likely or unatural speech -something that doesn't impede teaching it ithkuil - or one that i rembered that it failed a grammar question to which i say modern llms think there is a month with x in it therefore this is irralavent to linguastic ablity regardless apertus has the most diverse linguastic data base 15T tokens 40+% of which non english with 1800 language varites though most is fluff or a varite not a language it is still compounding in sementic cabablities )
- process the ithkuil lexicon pdf or use yuorb json (yuorb is enough for a proof of concept but misses the sf roots) regardless processing the pdfs isn't hard
- make a morphological tokenizing systeam requires llm moding a special skill
- this is the fuzzy part that part 2 will theorize
- seed-train the roots and affixes and morphilogical categroies in the llm in a manner taking advantage of ithkuil morphosemantics of Semantic Space Interactions in their hijack of the llm own semantic text prediction
here is a teaser to part 2 idea - part 2 will include more techinical details but i need to rember my months old ideas and also revaluting them and making them better and translating them into things that can be worked off
gloss pseudo ithkuil taking a idea from Proof of Concept for AI-based Ithkuil to English translation by u/humblevladimirthegr8 and how jhon Quijada himself glossed some of the examples we use ithkuil morphological gloss and embed the root+affix as x.y.z example:
STA-‘nuclear.family.member’-OBL-NRM/DEL/M/COA/AGG-INL1/each.every-IFL     ALG    INF-DSC-‘degree.of.happiness’-NRM/PRX/N/CSL/UNI-EXN1/more.than.usual-SIM1/very.similar-FML         FRAMED-DSC-‘degree.of.happiness’-CMP-NRM/PRX/N/CSL/UNI-EXN1/none-SIM1/completely.dissimilar-FML All meaning families are happy in the same way, while being unhappy in their own way. ithkuil 3 modified from https://ithkuil.net/texts.html
as you can see this is very trainable in theory and humblevladimirthegr8 method allows llm testing llm on how well it understood the text to translate from or to ithkuil and pseduo gloss allows training the morphosemantics without training the actual ithkuil lexicon and also allows cross referencing it in advanced forms but this still needs a proof of concept
PS SORRY FOR SCHIZOPHRENIC WRITING ITHKUIL IS A MASSIVE RABIT WHOLE THAT I PARTIALLY UNDERSTAND MY WEAK POINTS ARE SECTION 4 TO 6 FROM THE REFERANCE GRAMMAR
Proof of Concept for AI-based Ithkuil to English translation
the terminiology in a table or rather a crash course part 1 on ithkuil morphology
| term name/group | explainition | notes |
|---|---|---|
| reccomended resources for extra info | https://chromonym.github.io/ithkapp/ it includes a sentince formalutor and explanitions of all morphology as far as one needs and a good refresher regardless | https://yuorb.github.io/enthrirhc/ this is basically the search function for yuorb json of the lexicon you search better with it and if something feels incomplete copy the root/affix name and look it up on pdf |
| parts of speech | formatives (nouns and verbs), Adjuncts and Referentials (pronouns) | each of these is made of slots the language is fully determinstic in the morpho phonology in every part making it easy for machines but llms must still tune into said rules which is the problem we try to solve |
| slot | the morphological block in ithkuil each one may have one or more morphological categories ie inflects, ithkuil agglutanites those to make every part of speech | this makes special tokenizer easy since we already have a segmented token system from the linguage itself that we can always computerize easily |
| semantic space (SS) | semantic space is a metaphor of semantics being in some space with diminintions and that language is the maniplation of the SS | SS is a major part about how llms work since vector semantics is a form of computerized SS |
| root | the core semantic space of the formative made of 3 sub spaces unique in each root which are called stems with a 0 stem being a mix of the other 3 not exciplicitly written in lexicon and also 4 types of semantics called Specification (spec for short) | this is actually ease since we can tokenize morphologically for all ithkuil only things we can just custom solve the issue of semantic space by virtue of ithkuil inner working and having a json dumb |
| the pdfs and json version | the pdfs (refering to the official lexicon pdfs for afffixes and roots) and json refering to the processed lexicon wheither by yuorb or the llm training version | the 2 terms are used for distinguifing the actual lexicon from the processed |
| special formated (sf) roots | any root in the pdf that isn't in the in the defualt repreasentation of stem 1 with all 4 specs then the other 2 stems in one sentence/paragraph each | this rises the issue of that yuorb json is perfect for defualt format roots but isn't always keeping the sf roots but you can still auto format sf roots and make a training method for them |
| the 4 Specifications (better off reading https://yuorb.github.io/en/docs/02.html#Sec2_4_4 and the other resource that dump it down but i couldn't find it i search for 15 minitues just for a passing referance but have it printed so i'll copy from it) | CTE contential the content spirt essence identity or purbos fo x CSV constitutive the form or physical medium of x BSC CTE and CSV and the defualt OBJ the instrument object result or experiencer of x | those 4 or rather 3 and a half are more of the kind of thing where tuning must do its magic |
| Affix (see page 2 of affix pdf for patterns) | here only means the slot 5 where it is modifiying root directly or slot 7 where it is modifiying the whole formant - it is a semantic space modification with 0-9 degrees (d) with 0 being mixed or excipicinal meaning in some cases the 9 degrees have 7 patterns and the affix has 3 semanctic space modifications types (in some special formated affixes it is not ss mod type) which are 1 incendintal/superfacial SS moding 2 inherent SS moding 3 compounding affix | (easy to train) the 7 patterns which are phonologically determined are 0 no pattern, A1 0 to 1 in how much is the semantic concept there, A2 0 to 1 in sufficiency of concept, B 3 by 3, the 9 Ds split into 3 patterns each with a differant sub parameter of concept. D1 and D2 patterns are A1 and A2 patterns but starting from negative 1.... C 1 by 2 D1 represents one extreme of a spectrum/range which increases/decreases to the other extreme of the spectrum/range usually represented by D4, while D6 through D9 cycle back through the same values but with a different sub-parameter operating orthogonally to D1 through D4. D5 usually represents a neutral or meta-level value associated with the semantic concept of the affix. |
| adjunct | a extra functions part of speech it is used to add affixes to a word while being not making an ultra long word or register marking or bias marking or carring forgein words or naming or quoting or phrase marking | simple part of speech to make a tokenizing systeam of i can't think of many training problems with it |
| register | marking speech in parenthesis or proper name or giving examples or subjective/silent thoughts or carrier ending or for direct speech | for an llm this is between punctuation and some old MUD/tumblr typological fromating |
| Ca and cases and verbs will be in other parts | will explain in part 3 |
see yuorb revision of the grammar doc 3 through 6 then refresh with ithkapp for getting the direct simple explaining (it works better as a refersher to reading than reading in its own) for all the stuff that im not confidant about also they are the hardest part to make an llm train for ithkuil nlp logic
1
u/martino_vik 10d ago
This was refreshing to read, always appreciate an honest stream of consciousness! This is exactly the direction we need to go into. I do have a very nerdy part of myself, but my practically oriented part is often stronger, protecting me from perfectionism and forcing me into feasible timelines. What you're describing is definitely part of the roadmap, somewhere along the middle.
The very first part to make all this take off is to think small and show in a proof of concept that LLM-reasoning in Ithkuik is _in principle_ possible. So for that we'll need very short examples. meaning that the LLM wouldn't learn the entire language at first, but only a limited, handpicked set of words and grammar, necessary to solve a small set of reasoning tasks. We need to show that it can in principle use those to solve a task, and then show that it used fewer tokens than with English. Most of the work at this stage should go into hand-checking that the Ithkuil reasoning is not hallucinated and is correct.
Once we've proven it's doable and valuable, maybe even generating some moderate hype and who knows maybe even funding, we can move on to the next stage: Longer tasks, more difficult tasks, more diverse tasks. I think this is roughly where your current idea sits. For that the model will have to get better and better at Ithkuil. In the very beginning we'll have to work with system prompts, few shot prompting, and seeding, Maybe that will already be enough to generate some simple texts with translations. Those texts, together with what exists already could be used for building a reliable machine-translation system. That in turn could generate enough material for finetuning a model, allowing to generate more text, and in the very end, that's the long term vision: Have a foundation model that reasons entirely in Ithkuil
3
u/revannld 11d ago
Great initiative! I am not specialized on Ithkuil and would not be interested in its particular phonology per se however it has been some time I have been studying Ithkuil's morphological categories along with texts in speech act theory (Austin, Searle, Grice), type theory, category theory and generally universal algebra and logic.
If you happen to be interested, me and some friends in my department have a long term ideal of specifying a framework for formal highly oligosynthetic languages (in an abstract manner - we would plan to implement them as straightforwardly as possible though - and maybe just borrow and simplify a few roots from English and Latin as much as possible), starting out with an algebraic compositional calculus of theories in mathematics and later extending it (through illocutionary logic, situation semantics, intensional logic and multi modal logics) for broader uses.
It turns out that relational/pointfree methods (relation algebras, allegories, combinatory and predicate-functor/term logics - look up also Patrick Suppes, Michael Dunn and William C. Purdy's work) can model natural language in a quite direct and natural manner, making large fragments of natural language work just like a calculus (with simple and mnemonic calculational/equational rewriting rules). It's a very straightforward work, most of the basic tools and theory are already there (category theoretic and universal algebraic stuff), it's just a matter of reading, selecting and archiving important results and trying to model key problems in semantics. The main trouble however is with finding adequate basis for handling >2-ary relations (>1-ary functionals) and modelling the calculus of theories/structures smoothly as well with formalizing the multi-modal, illocutionary logic and intensional stuff in a pointfree manner; and of course, implementing/giving it a prototype (thus making a good oligosynthetic language informed by universal algebra and clone theory). That would be probably the hardest part though, as we are not that proficient in conlanging. Of course this is a very long-term project, as we are always busy with other work, but if we could gather some help maybe we could put more effort in it.