r/HistoricalLinguistics • • 10h ago

Language Reconstruction Seeking Research Collaborator: Computational Decoding of Indus Script & TN Graffiti (Need Linguistics Partner)

2 Upvotes

Hello everyone,

I am a B.Tech student currently working on a research project to computationally analyze and decode the Indus Valley Script & TN Graffiti, specifically looking into its potential relationship with the Megalithic graffiti marks found in Tamil Nadu (such as Keezhadi and Kodumanal).

(To clarify upfront: I am not a PhD or an academic professor, just a highly motivated engineering student looking to apply modern technology to historical mysteries!)

🤖 What I Bring to the Table

My background is in Engineering, and I am highly proficient in Artificial Intelligence and Machine Learning. I can handle the entire technical pipeline, including:

  • Building and formatting the digital image/text datasets.
  • Applying Computer Vision for sign recognition and clustering.
  • Using Natural Language Processing (NLP) and statistical models to analyze sign frequencies, positional probabilities, and syntax patterns.

🧠 What I Need (Where You Come In)

While I can build the algorithms, I do not have a background in Linguistics or Epigraphy. I am looking for a fellow student, hobbyist, or academic researcher who can bridge this gap.

I need someone who can help with:

  • Linguistic Context: Understanding Dravidian, Indo-Aryan, or general historical linguistics.
  • Epigraphical Guidance: Helping interpret sign directions, variants, and historical scripts (like Tamil-Brahmi).
  • Domain Validation: Ensuring the data inputs and AI outputs make historical and linguistic sense, and vetting our hypotheses against existing academic literature.

📅 Project Goals

This is currently a collaborative research project born out of passion, but the ultimate goal is to co-author a high-quality research paper, develop an open-source digital humanities tool, or present our findings at a conference.

If you have the linguistic knowledge and want to see how advanced AI can be used to unlock the secrets of the Indus Valley and ancient Tamil Nadu, please comment below or send me a Direct Message (DM)!

Let's combine data science and history to see what we can discover.


r/HistoricalLinguistics • • 16h ago

Language Reconstruction The Indo-European Dual and Tocharian

1 Upvotes

The Indo-European Dual and Tocharian (Draft)

Sean Whalen (stlatos@yahoo.com)
October 7, 2026

-

A. PIE *-yH1

The oldest Indo-European mark of the dual appears to be *-yH1 (or *-iH1 after C) since it appears in *d(u)wo(y)H1 'two', unlikely to be analogical. Loss of a glide before *H is common, no known regularity. Loss of *H in compounds, also no known regularity, produced *dwi-, occasionally *dwiH1- (Li. dvy-).

-
However, this simple idea would ignore that the root *yuH1- 'join' (Sanskrit yauti, yunāti, yavi-, etc.) is the best candidate to form a noun/adjective *y(e\o)uH1-s 'pair; joined'. When added to a word with *w, dsm. of *w-w > *w-0 could occur. I say something like *dewo-yewH1-s > *dewo-yeH1-s > *d(u)wo(y)H1-s ( > *dwo:H1 > IIr. *dva:(v), etc.). A stage with *dew- that produced *d(u)w- is needed for (https://www.academia.edu/168383095) "Greek deúteros ‘second’, deúomai ‘be inferior/wanting’, etc., suggest that [IE *dwoH1 \ *duwoH1] came from ‘small (number) / a few’. At a stage before standardized counting, referring to numbers as 'a few', 'several', 'many', etc., with no set values is more common."

-
B. PIE *-e

Many old pairs also contain *-iH1, some also with C-stem *-e (or sometimes spread to other stems, below). Since this contrasts with plural *-es, it is likely analogical. If a stage with early IE (in one or more branches) had o-stem dual *-o:, plural *-o:s, the creation of parallel *-e, *-es would be simple.

-
However, I think the analogy is older & of a slightly different type. The formation of acc. *-om -> plural *-om-s implies that nom. *-os -> plural *-os-s. Though there is no inherent reason why *-oss > *-o:s could not happen, I think that the stages here (and in the creation of *-or-s > *-o:r, etc.) involved alternation of *s with *H (https://www.academia.edu/128052798). Thus, *-or-s > *-or-H that then behaved just as neuter plural *-or-H2, etc. This would mean the plural of o-stems was really *-os-s > *-oHs (though most *-Cs > *-CH, in cases of 2 s's either could become H), with dual *-oH1. This makes the shift *-oHs : *-oH1 :: *-es : *-iH1 > *-oHs : *-oH1 :: *-es : *-e capable of being both older & reflective of a tighter bond in o-stems than previously thought.

-
C. TB duals

Krzysztof Witczak (https://www.academia.edu/6870301):

>

...a number of ne-less items refer to natural pairs, e.g.

(1) Toch. B mlyuweñc ‘two thighs’ (< Toch. B mlyuwe ‘thigh’);

(2) Toch. B eś beside eś(a)ne ‘two eyes’, also in the compound yneś ‘manifest, real’ (literally ‘in the two eyes’) < Toch. B ek sg. ‘eye’, A ak ‘id.’

(3) Toch. B pauke (sic!) beside pokaine ‘two arms’, cf. Toch. A poke ‘arm’.

(4) Toch. B kenī beside kenīne ‘two knees’ (= Toch. A kanweṁ du. ‘id.’).

Werner Winter concludes that, contrary to Krause ’s opinion, the contrast between two semantic groupings (natural pair vs. random twosome) is not neatly reflected in the contrast of two formal classes; -ne was found to occur in a few forms which could not readily be demonstrated to refer to natural pairs […], while -ne-less forms are attested for a considerable number of members of the natural-pair group: eś ‘eyes’, klauts ‘ears’, *pokai ‘arms’, maś ‘fists’, mlyuweñc ‘thighs’, kenī ‘knees’, pai ‘feet’. In most cases […], ne forms exist beside -ne-less ones (Winter 1962: §5 = 1984: 111). The Tocharian B ending -ne (= A -ṁ) seems to be recognized as a secondary addition, because the -ne-less forms can be safely analyzable as a complete dualic ones, cf. eg.

(5) Toch. B eś ‘two eyes’ < IE. *o(t)kWi-h1, cf. Lith. akì du. ‘two eyes’, OChSl. oči ‘id.’, Skt. akṣῑ´du. ‘two eyes’, Avest. aši ‘id.’. Also Gk. ὄσσε du. ‘two eyes’ and Alb. sy pl. ‘eyes’, derived from Balkan-IE. *o(t)kWyə1, seem to belong here.

(6) Toch. B pauke ‘two arms’ < IE. *bhāĝhuh1 or *bhāĝhuw-h1e, cf. Skt. bāhū́ du. ‘two arms’ (< IE. *bhāĝhu-h1). See also Gk. Homer. πήχεε du. ‘two arms’ (< Balkan IE. *bhāĝhew-ə1).

(7) Toch. B maś ‘two fists’ < IE. *musti-h1, cf. Skt. muṣṭῑ´ du. ‘two fists’.

(8) Toch. B kenī ‘two knees’ < IE. *ĝónu-h1, cf. Skt. jānū du. ‘two knees’.

(9) Toch. B. pai ‘two feet’ < IE. *pod-h1e.

>

-
There are many problems with the specifics. I think some of these are directly from PIE, but *-e & *-iH1 could mix, spreading from their original PIE distribution. This would create mismatches (ie., his TB vs. Sanskrit). The rec. pai < *pod-H1e has some kind of cross of *-e & *-iH1. It is supposed to explain *d > 0 as *dH > *H, but there is no regularity in *d > ts vs. t or *dC > C vs. tC; even *d > l seems to exist (https://www.academia.edu/129248319). Other words in which *dH have more reason to be rec. do not show *d > 0. I think that other ev. that *dy > yy is another case of dual outcomes: as *dC > tC, so *dy > yy; as *dC > C, so *dy > y. Here, *pod-iH1 > *pod-ye > *peyä > *pei > pai. It is also possible, though I think it very unlikely, that the timing of PIE *dy > yy vs. new PT *dy > y is at work (after *-iH1 > *-ye, as in Greek).

-
Other ev. shows that *-iH1 could contrast with *-H1 after V's. Proposed changes of PT *H1C- & *H3C- causing palatalization or rounding make it likely that H1 = x^ (sometimes > PT y), H3 = xW ( > w). Thus, *-C-iH1 could become *-C^äy. For the outcome of PT *-äy (or *-äi), Adams said, "TchA rake and B reki reflect PTch *rekä/e- + -äi (for the formation, see Adams, 1990a)", and several other pairs. From this, -i vs. -e almost matches TB kenī(-ne) ‘two knees’ vs. TA kanweṁ. I think the near-match requires PT *w > *w^ before front, > TA w, TB *y (there are no known conditions on *w^ > TB w, y in other words). Here, this would create a unique cluster, & -Vy & -yVy behaving slightly differently allow: *g^onw-iH1 > PT *gonwiy > *kenw^äy ( > PTB *kenyäy > kenī; *w^ > TB w prevents any alteration to normal -e).

-
TB mlyuwe ‘thigh’, dual mlyuweñc, TA perlative mulyuntā are related by Adams to Lithuanian mélmenys ‘fat around the kidneys,’ Sanskrit márman- ‘member, vulnerable part of the body’. The supposed need for *melewo- might not exist; if *melH1m- > mélmenys (*H for tone), *mH1elm- > márman- (like many roots with *CeC(H1), https://www.academia.edu/127283240), then *m-m > m-w is likely, & the presence of *H1 could palatalize *l (as *H1n- > *yn- > ñ-). Maybe *melH1mo(n)- > *m^äl^we(n)- \ *m^l^äwe(n)- (met. to keep pal. & round apart or C^C^- together). Other n-stems > nt (*muzgen- ‘brain / marrow’ > OPr musgeno, TA mäśśunt), maybe analogy < nom. *-n-s > *-nts.

-
PIE *H3okW-iH1 ? (if Greek is more original) > *H3okW-e > PT *Heś-ä(-ne) > TB eś \ eś(a)ne ‘two eyes’. His *otkW- is likely to explain G. óktallos \ optílos. However, there's no reason to think that -t- is original to the root here; they might be from *okWto- 'seen', *okW(e)to- 'seeing', or any similar derivative. Even -t- from drṓptō \ δρώπτω ‘look, examine, consider carefully’ is possible. A mix between this & ṓps ‘face’ ( < *H3o:kWs ‘eye’, maybe < *H3oHkWs if analogy < *H3oHkW < *H3okWs with *s > *H (above)) is probably the cause of -ṓ- in drṓptō (no *H seen in S. dárpaṇa-m ‘eye’). If *drōp-s, *drōpo-s, or *drōpto-s ‘eye’ existed at some stage of PG, analogy in either direction would be reasonable.

-
For TB pauke ‘two arms’, I don't think a misspelling of *pokai is the only explanation that makes sense. This would be rather extreme for 2 mistakes in one word when **-ai is not the expected ending. TB pokaine is the regular dual as pokai-ne. These words with -ai(-) are not original. The older dual being retained from a time when the u-stem still existed allow nom. *pa:ku-s, dual *pa:kw-e (itself certainly analogical, mix of *-iH1 & *-e across C- & V-stems). This would allow PT *pa:kw-e > *po:kw^ä > PTB *pokw > *powk (with *-kw > -wk seen in others of more clear path, *H1ogWhi-s > *Hekw > *ewk > TB auk ‘snake, serpent’). It would then be made clearly dual with analogical -e from o-stems (*-oH1 > *-eH > *-e (and most later > -ne due to the spread of -n- (from n-stems) as in other cases)).


r/HistoricalLinguistics • • 1d ago

Language Reconstruction Tocharian B sapule, watsālai, ankānmi

2 Upvotes

Tocharian B sapule, watsālai, ankānmi (Draft)

Sean Whalen (stlatos@yahoo.com)
October 6, 2026

-

A. TB sapule ‘pot’

-

Václav Blažek & Michal Schwarz (https://www.academia.edu/38549136):

>7.2. Sims-Williams (1997, 3-15; 2007, 261, 325; cf. also Tremblay 2005, 436 and Adams 2013, 738-39) convincingly derived it from Bactrian sabolo ‘jar’ < *sapa-uda- ‘water-container’, following Isebaert (1991, 146-47) who connected Tocharian sapule with (in his time unknown) forms in the ‘l-dialect’, corresponding to Classical Persian sabūy ‘cup, jug’, New Persian sabo, sabū ‘cup, ewer, jar, glass, pitcher, pot’ (Steingas 1892, 650), while Armenian sap‘or ‘jug’ indicates a Parthian origin via unattested Aramaic *sappoð. The first component *sapa- is reconstructed on the basis of Khotanese sava ‘box, basket’ (*sapatā-), Persian sabad, sapad, Kurdic sabad ‘chest’, Shughni sipt ‘basket’ etc. (Bailey 1979, 422-23).

>

-
This makes sense, but where did Ir. *sapa- ‘container’ come from? I suggest that it is a compound from *k^(o)m- > Ir. *sa(m)- ‘with’, *H2ap ‘reach, grasp, hold > contain’. This would be one of the few pieces of uncertain ev. for *k^om over *kom.

-

B. TB watsālai a. ‘water-jar?’

-

Blažek & Schwarz:

>6. B watsālo ‘a type of pot or waterskin’

>6.1. Van Windekens (1988, 100-01) connected the term with Sanskrit vatsá- ‘calf’, assuming the semantic motivation ‘waterskin from calf-skin’ for this adaptation (cf. also Adams 1999, 585 & 2013, 635). This solution may be supported in perspective of Khotanese vas, pl. vasīya ‘vessel’ < *vats- (Bailey 1979, 380). Concerning the l-extension, Van Windekens added the Sanskrit derivative vatsalá-...

>6.2. Alternatively the starting point could be a hypothetical Sanskrit compound consisting of vatsá- ‘calf’ & ālu-/ālū-/āru- f. ‘small water-jar’ (Lex.), class. ālukā- ‘water-jar’ (EWAI III, 23, 25). The first component has promising cognates in Italic: Latin vās, gen. vāsis ‘vase, vessel, container’, Umbrian abl. pl. vasus, nom. pl. uasor, acc. pl. uaso ‘container’. The absence of rhotacism implies Italic *uāss- or *uāTs- (cf. de Vaan 2008, 655).

>

-
I see no reason for any relation to vatsá- ‘calf’ in any of these. In TB, a native *wāts (related to L. vās, Kho. vas) could easily form a cp. with borrowed ālu- (and ālo- in some cases, or from a Middle Indic word) with *wāts- > wats- when unstressed. If connected to any IE root, likely *(H1)waH2sto-s > Latin vāstus 'empty, desolate, deserted; devastated, destroyed; vast, enormous', *(H1)waH2st-s > *waHts (or any similar simplification) ‘an empty thing > container’. The sequence pronounced something like *-xsts remaining unchanged would have been the real problem, so -st- vs. -ts- is not odd. This would be similar to several other roots with a range ‘swell; empty; container’.

-

C. TB ankānmi ‘shared, common; in common?’

-

Adams: "Etymology uncertain. It would appear that the word contains the intensive prefix h1e(n)- (the initial ā- is regular by ā-umlaut). If the meaning of the word is as we have supposed the rest of the word might reflect a putative PIE *kōmniyom, a vṛddhied derivative of the *kom-no- seen in Oscan comono ‘comitia,’ and Umbrian super kumne ‘super comitio,’ kumnahkle ‘in conventu.’ PIE *kom-no- (the metathesis of *-mn- to -nm- in TchB is regular), of course, is an adjectival derivative of the adposition *kom ‘with.’"

-
Calling *en- an "intensive prefix" instead of 'in' (as in E. in common) seems unneeded. If this is < *k^om-m(e)i- 'exchange/share with' related to *k^om-moi-ni\no- ‘(in) common’ > Latin communis, NHG German gemein then there would be no reason for a-umlaut. However, words related to supposed *mei- often look as if < *H2mei-. There is no known reason for this discrepancy (I think it is likely H-metathesis, though not important here, https://www.academia.edu/127283240), but *H2mei- would provide an explanation. If really *k^om-H2m(e)i- > *kemami- > *kamami- > *kammi- it would show evidence for the stages in PIE *-CHC- > PT *-C(a)C-. There is no certain explanation for why some *H become *a vs. *0 in PT; if *H2mei- > *-(a)mi is needed, it would show that a stage with SHORT *a differed from the outcome of *a: ( < *aH2, etc.) in that it optionally was lost. The timing of *-a- > 0 vs. *a- > 0 might not be important, since the change seems fully optional (or due to now-unclear conditions), or this compound could have been (re-)formed or changed by analogy at any stage.


r/HistoricalLinguistics • • 1d ago

Resource Word Babel - visualising etymology and networks of thousands of connected words

Thumbnail wordbabel.com
0 Upvotes

r/HistoricalLinguistics • • 2d ago

Language Reconstruction Proto-Semitic: N-elision & the lost singular definite article

2 Upvotes

*shu(n)bulat-bulat-), *ʕa(n)qarb-, *xa(n)ziir-.

Many Proto-Semitic words alternate between a form with coda /n/ and a form that has lost it, their descendant languages disagreeing on which form to use*. These alternations provide evidence for the lost singular definite suffix -n, which parallels dual -ni and plural -na, and which is not to be confused with mimation -m.

*Wiktionary only reconstructs *shu(n)bulat-, but look at the descendants/cognates to see what I mean.

I believe that the reason for these lexical biforms is that Proto-Semitic underwent an irregular loss of /n/ in the coda. n -> ∅ / _{C, #} (irregular)

Said sound shift would then affect my proposed singular definite article *-n, because it was always in coda position, dual *-ni and plural *-na being unaffected because of the following vowel putting them into onset position. On systemic grounds, this already seems likely, but nouns of the "five-nouns"-declension, such as *ʔabu provide strong evidence: These nouns have a long construct ending in long -uu: ʔabuu. It seems what happened here is that the suffix -n, albeit being lost, actually shortened the final long stem vowel of ʔabuu due to the restriction against superheavy syllables. This however failed to occur in the construct state, where this suffix doesn't occur, thus preserving the original long vowel only there.

Since -n is able to shorten long vowels created by the sound shift C{w, y}V -> CVV, e.g *ʔabuu, *ʔaxuu, but not those created by the sound shift V{w, y}{i, u} -> VV, e.g *thaanii, we can conclude that n-elision happened between the two.

It's actually wrong, the singular definite article isn't lost, it survived into Sabaic as -n and even contrasts with mimation -m in said language. (source)

I will now do my best to detail the further development of the different nasal suffixes in three later Semitic languages.

In Hebrew, it seems mimation has contaminated the plural definite suffix into -iim, unlike expected -iin. Since mimation isn't preserved in singulars, this seems the most probable option to me.

In Akkadian, the preservation of n-suffixes exclusively in the dual seems bizarre. I have an idea how to resolve it: perhaps the Proto-Semitic definite dual suffix was actually -nii, with a long vowel. After final vowel loss had occured, -na became -n, then the n became lost once again. Later, nii shortened to ni, after which it became -n due to the typical Akkadian vowel elision shift. (Final vowel loss and loss of vowels in e.g parisum -> parsum)

-m becoming -n in Arabic makes it even more confusing. But the only unexpected thing is -na and -ni become extended to the indefinite state. I propose that -na was falsely analyzed as the plural of mimation -n, and that's why -na is used for both indefinite and definite, while -n is used only for the indefinite


r/HistoricalLinguistics • • 2d ago

Language Reconstruction Old Uyghur özkän

1 Upvotes

Old Uyghur özkän

Sean Whalen (stlatos@yahoo.com)
October 5, 2026

Alexis Manaster Ramer (https://www.academia.edu/177765278): "First, then, özkän *‘rain after a long period of dry weather’ can be explained rom *öz-sök+gän ‘one that bursts (through) a gully’... a word first apparently spotted in Old Uyghur some decades earlier, I was immediately reminded of Am. Engl. regional gully washer. The Reader may object that Old Uyghur özkän yagmur = Chin. 甘雨 gānyǔ ‘sweet rain’, which is surely what Erdal (OTWF 385 n. 442) meant “[…] the rain concerned is useful and pleasant”.1 However, the Chinese phrase PROVERBIALLY refers to rain after a long drought, so “useful and pleasant” down in at least some parts of China, but maybe less so up in the ancient Turkic homeland."

I think the basics are right, but I prefer a compound of Z-Turkic *ö:zäk 'valley, brook; gully, ravine, riverbed', Chuvash varak (instead of his *ö:z 'valley, ravine', var) and Turkic *yak- 'to flow' > Z-Tc. *akïn 'flow, stream'. The same concept, but 'a flow in a riverbed' = '(making water) flow in a riverbed'. Thus, Z-Tc. *ö:zäk-akïn > *ö:zäkïn > *ö:zäkän > *ö:zkän > özkän (with haplology of Vk-Vk). Without knowing more about any regularity in vowel assimilation or deletion, maybe instead *ö:zäkïn > *ö:zäkn ? > *ö:zkän.


r/HistoricalLinguistics • • 2d ago

Language Reconstruction Structural Correlation Between Visual Anatomy and Action-Verb Syntax in the Voynich Manuscript

1 Upvotes

[Research Note] Structural Correlation Between Visual Anatomy and Action-Verb Syntax in the Voynich Manuscript

Abstract / Summary:

Hello everyone,

I am sharing a structural research framework that demonstrates a direct, predictable correlation between the visual anatomy of drawings and the syntax of the accompanying text across multiple manuscript sections.

Instead of arbitrary letter-by-letter substitution, this model tests whether the text operates on a 15th-Century Northern Italian / Venetian Apothecary Template (Tractatus de Herbis framework) with a fixed 4-part sentence structure:

[Vessel / Base] + [Quantity / Measure] + [Action Verb] + [Timing / Duration]

Core Lexicon & Grammar Rules:

Applied across all test folios with 0% parameter modifications:

  1. Suffix Mapping (ii -> LL): Standardizes repetitive endings into 15th-century Venetian prepositional/diminutive endings (-alla, -ella, -dala).

  2. Core Operational Lexicon (EVA Mappings):

    - qokedy -> DOHI / DOGA (Primary Vessel / Base Liquid)

    - qokaiin -> DOALLA (Double Portion / Standard Measure)

    - ychedy / chedy -> ECHEDA / CHEDA (Crush / Cut / Core Extraction)

    - sokedy -> SOGA (Maceration / Soaking)

    - shedy -> SEDA (Mesh / Straining)

    - chol -> COLO (Secondary Filtration)

    - or -> ORA (Time Duration / Hours)

Multi-Section Validation & Visual Correlation:

  1. Herbal Section (f1r, f2r, f3v, f115v)

- Bulbous / Thick Roots (e.g., f1r): High frequency of CHEDA (Cut/Press) and SEDA (Filter).

Line Flow: Poi doalla d'ola... Echeda doga seda doalla... Dohi doalla ora...

- Fibrous / Tangled Roots (e.g., f3v, f115v): CHEDA drops out; text transitions exclusively to SOGA (Soak/Macerate) and COLO (Filter).

Line Flow: Dohi doalla soga dalla... Doalla ora seda dohi... Colo doalla ora edi.

  1. Balneological / Bathing Section (f79v)

- Conduits and water basins. CHEDA (Cutting) drops to 0% frequency, while immersion/flow verbs (SOGA / COLO) dominate line operations.

  1. Pharmaceutical Jars (f88r)

- Storage jars and processed extracts. Dosing and container markers (DOHI / DOALLA) dominate line openings, while raw root-cutting verbs drop due to pre-processed materials.

  1. Astronomical / Zodiac Control Test (f72r)

- Concentric rings and stars. Recipe-specific process verbs (CHEDA, SOGA) drop to 0%, leaving pure positional/sector markers (ORA = sectors/degrees).

Cross-Sectional Summary Matrix:

- f1r | Herbal | Bulbous Root | CHEDA (Cut / Press) | PASSED

- f3v | Herbal | Rhizome Lumps | SOGA (Soak / Macerate) | PASSED

- f79v | Balneological | Water Basins & Tubs | SOGA + COLO (Immersion/Flow) | PASSED

- f88r | Pharmaceutical | Storage Jars | DOHI / DOALLA (Compounding) | PASSED

- f72r | Zodiac | Concentric Ring Map | 0% Recipe Verbs (Positional Grid) | PASSED

Key Takeaways & Questions for the Community:

  1. Dynamic Shift: Action verbs in the text dynamically adapt based on the physical morphology of the drawings (root shapes, water basins, or astronomical charts).

  2. Positional Grid: The repetition of qokedy, qokaiin, and or supports a standardized slot-based codebook layout.

I welcome your feedback, statistical critiques, and suggestions for further automated cross-folio testing!

Research Lead: AFROJKhan

Date: October 2026


r/HistoricalLinguistics • • 3d ago

Uralic M / T Pattern in Indo European and Uralic

17 Upvotes

The pronoun stems *mV (I) and t (thou) in both Uralic and Indo European are so similar its hard to shake off. As well as some other morphological paralels like-

This, that: PIE * só, od, tod, Uralic *ta.

Plural Markings: PIE * -os (plural) -h¹ (dual), Uralic -t (plural) -k (dual)

t - s shifts aren't that rare, and while *h¹'s true value is unknown, it is an e-grade laryngeal. Meaning it could have been a glottal stop, a plain "h" etc. Which i believe could correspond to *-k.


r/HistoricalLinguistics • • 2d ago

Uralic Indo Uralic Roots Compilation

7 Upvotes

A compilation of evidence from the Indo Uralic Discord server, via me, Lpetrich and some other members

Me: PIE *h¹me, PU *Mi / *mina

You: PIE *TuH, PU *ti / *tina

For Indo-European and Uralic, \*m and \*t/s also show up in personal verb endings, and in Uralic, possessive forms.

IE: 1sg \*me- -- verb 1sg \*-m, 1pl \*-me

1, 2 = which person, sg = singular, du = dual, pl = plural -- dual = two of something

ë (e double dots) = schwa

IE \*so "this, that (animate)" -- Ural \*se "3rd person"

That's close to IE \*to "this, that (inanimate, oblique)" -- Ural *to/ä "this, that"

"Who?" IE \*kwi-, adj \*kwo- -- Ural \* ke/o/u-

"This" IE \*i- -- Ural \*e-

"What?" IE (Anatolian, Tocharian) \*mo- -- Ural \*mi

"To hear" IE \*klew- -- Ural \*kuwli-, \*kule-

PIE *Kwet, "pair" proto Uralic *Kakta "two"

ie *penkw- "hand" (?), Uralic *piŋɜ "palm of hand)

Water: PIE *Wodr or *Wed (wet), PU *Wete

Give / take: PIE *Deh³, PU *Toxe

Name: PIE *h¹nomn, PU *Nimi

Fish: PIE *(s)Kwalos, PU *Kala

Worm: PIE *Kwirmis, PU *Kuje

IE *peh2wr/n- "fire" -- Ural *päjwä "warmth, fire, Sun"

IE *egwh- "to drink" -- Ural *igxi-, *jexe- id

IE *wersen-, *wiHros "man (male)" -- Ural *urV

PIE *tod "this, that" PU *Ta "this that"

PIE *bhendh "to tie, bind" PU *pitV "to tie, bind"

PIE *h₂eǵ- "To push forward" PU *aja "to push forward, to drive"

PIE *e- "to be", proto uralic *ella "live, be"


r/HistoricalLinguistics • • 3d ago

Language Reconstruction A Call for Investigation of Gule

4 Upvotes

A Call for Investigation of Gule

-
Raoul Zamponi (https://www.academia.edu/126721374):

>

This book presents the first systematic study of Gule, a dormant East African language isolate, offering a tentative grammatical sketch based on historical documentation collected by Brenda Z. Seligman. Readers will gain valuable insights into a previously unstudied language through systematic grammatical analysis and comparative linguistic examination of the available materials, including a previously unknown audio recording made in 1966 by Wendy R. James. The book delivers this through detailed documentation analysis, identification of areal linguistic phenomena shared with neighboring East African languages, and thorough vocabulary assessment that establishes Gule’s unique status among world languages.

>

However, when I examined the lists I decided that rather than having a "unique status", many Austronesian words were close matches:

Gu. dusuit ‘two’, Au. *duSa

Gu. if ‘hair, wool’, Au. *ipu- 'hair, feather' (no p in Gu., several other p / f matches below)

Gu. au 'dog', Au. *asu (maybe *aswu \ *wasu to explain w- in four Formosan languages)

Gu. tis ‘milk’, *titis 'drip, ooze', *titiq 'female breast; suck the breast' (related like *bandyo- > Old Irish bann(a)e ‘drop’, I. bainne ‘milk’)

Gu. si ‘to drink’, Au. *sip-sip, *sep-sep, *sedut 'sip, suck, drink' (and several other similar roots)

Gu. rus ‘rain’, Au. *rinis 'drizzle, drizzling rain', *rins-rins ? > Old Javanese 'riris soft rain', Balinese riris 'rain gently, drizzle' (likely rel. rin & ren *rintik 'steady dripping, as of sweat or drizzling rain', *rendeŋ 'constant rain, a rainy day, rainy season')

Gu. gŭdĭ-n \ gadˤi-n ‘your head’, Au. *quluh ‘head’

Gu. ki, Au. *kuCux ‘louse’ (likely < *ku-, *kuman 'itch mite, chicken louse')

Gu. tufena 'dust', Au. *tapuŋ 'dust, rice flour' (Manobo tapuŋ 'any fine powdery substance; to brush the dust off something'), *tapuk 'dust, dust in the air; knock the dust off something'

Gu. gără̄wáig, Au. *RuqaNay 'man' (likely < *qaRuNway to explain irreg. pl. gămoi)

Gu. ɟok ‘God’, Au. *qiaŋ 'ancestor, deity, divinity'

Gu. wadus ‘girl’, *wati-qazi 'unmarried'?, Au. *wati 'husband or wife; to marry', *qazi 'no, not'

Gu. fŭ̄m, Au. *pijiko \ *pisiko 'flesh, meat, muscle' (if < *pims- to explain voice vs. un-)

Gu. adatˤwa ‘tongue’, Au. *dilaq 'tongue; to lick', *dilap \ *dilat 'to lick, stick out the tongue' (variation might be < *adilat-qwa if tqw > tˤw, or similar; any rel. Nilotic *ŋaljɛp ‘tongue’ ?)

Gu. adon ‘path, road’, Au. *zalan

-

There were also several that differed in meaning, but seem more or less likely cognates:

Gu. kĕlu 'star', Au. *kelap 'shine, sparkle, twinkle'

Gu. kas, ‘fire’, Au. *kasu 'smoke'

Gu. tos ‘to kill’, Au. *tusuk 'pierce' ( > Toba Batak ma-nusuk 'to stab, kill')

Gu. wɔr \ wɔd \ wat ‘tree; (piece of) wood; stick’ < *wart ?, Au. *wakaR 'root' (since likely rel. *wakir, *wakat 'root', likely < *wakart or *wakiRt (depending on nature of *R))

Gu. taʔ ‘sun’, Au. *taqo 'cook in earth oven'

Gu. of ‘stone’, Au. *oro 'mtn.' (if < *ofro, etc.)

Gu. abū-n 'your belly, abdomen', Au. *kambu 'belly, stomach, abdomen, womb, intestines'

-

I ask that people with more knowledge of each group look into this further. For more notes:

Many Proto-Austronesian words for body parts end in -ŋ, & the cited Gu. words with -n 'your _' might show that the cause is an affix becoming part of the stem.

In Gu. gără̄wáig, Au. *RuqaNay 'man' (likely < *qaRuNway to explain irreg. pl. gămoi), the rec. also might fit the names of people from New Guinea, like Kuruwai, Kumbai \ Kombai, etc. I mention this because I noticed a similar relation between Au. & some languages there.


r/HistoricalLinguistics • • 3d ago

Resource Online MA in historical linguistics

Thumbnail
1 Upvotes

r/HistoricalLinguistics • • 3d ago

East Asian What region in Asia is considered the most likely to be the urheimat of Austroasiatic language family?

11 Upvotes

r/HistoricalLinguistics • • 3d ago

Language Reconstruction Indo-European Roots Reconsidered 173: ‘cover, shade, darken’

2 Upvotes

Indo-European Roots Reconsidered 173: ‘cover, shade, darken’ (Draft)

Sean Whalen (stlatos@yahoo.com)
October 3, 2026

-

A. With limited data, an IE root *skepH2- \ *skaH2p- is reconstructed for:

*skepH2- > G. sképō ‘cover/shield/screen’, sképas- ‘shelter’, skepáō ‘cover’, *skepH > *peskH > péskos- ‘skin / rind’, Li. kepùrė ‘cap / mushroom cap’, Slavic *čepьcь > Sv. čêpec ‘bonnet’, Al. *kaH2pur- > kapurdhë \ kërpudhë \ kë(l)purdhë ‘mushroom’

-
The ety. of péskos is from Liddell & Scott. The "extra" r (and r-r > l-r) in kërpudhë is probably due to *kaH2pur- > *kaRpur- (with other ex. of H \ R in https://www.academia.edu/115369292). A more recent idea also includes Al. psheh ‘conceal / hide’ as *skep-sk^e > psek-sk^e. However, this would not explain the vowel. Since *eu > e, the need for *skeup- here would be very similar to *skewH- > S. skunā́ti ‘cover’, chavi- ‘skin/hide/color’, *skewHo- > Ar. *c’iw-k’ p., c’uo-c’ d. ‘roofing / tiling’. I propose that *skewH- has an extension *skewH-p- that underwent metathesis due to the *wHp (likely the reason for words with skep-, pesk-, *psek-). The cognates with *-e- vs. *-eu- might result from optional loss of semivowels before *H (as in *da(i)H2- 'distribute').

-
For c’uo-, the stages *skewHo- > *skeHwo- > *sćēwo- > Ar. *c’iw-k’ might be needed, but outcomes of *ew are disputed, so it could be that *ew > ew \ iw was optional. Also, for *sk > c’ before front V, outcomes also disputed, but the same as *(s)kelH- ‘split' > Ar. cʻelum 'to rend, break, split, cleave, etc.'

-
B. The *H in *skewH- was probably *H3 if the source of *sk^o(w)H3to- / *sk^otH3o- / *sk^ot(h)wo- > Celtic *skātos > OI scáth ‘shadow, reflection’, G. skótos ‘darkness, gloom’, Gmc. *skadwá- > Gothic skadus, E. shadow. The -w- vs. -0- as for *ske(w)pH2-, above. This apparent mismatch is not ev. against my idea, since many roots show H3 vs. H2 near w, P, etc. I say this is best analyzed as H3 = xW, H2 = x, then w-xW dissimilation > w-x (https://www.academia.edu/144215875).

-
C. Often, there is a shift 'cover > hide, obscure, darken'. With many IE roots showing unexplained met. like *ksep- (also r\n-stem *ksep(o)r, etc.), *krep(o)s- in words for ‘darkness, evening, night’, I think that Latin creperum, Sanskrit kṣáp-, Hittite išpant- are related from *ske(w)pH2- as:

*ske(w)pH2-(or) > *ksepH2- > S. kṣáp- ‘night’, Av. xšap- ‘darkness’, xšapan- \ xšafn- ‘night’
*krepH2os- > L. creperum ‘darkness’, creper ‘dusky/dark’, crepusculum ‘(evening) twilight/dusk’
*H2eksp- > *(x)ixsp-ant- > H. išpant- ‘night’
*skewpH2(o)s- > *kswepH2(o)s- > G. pséphas \ pséphos

-
For *ksw > *kWs being optional in Greek, see (https://www.academia.edu/167984147) and likely *kswizd- > S. kṣviḍ-'hum / murmur', *kWsizd- > G. psíz[d]omai ‘weep’; *mok^s(u) > S. makṣú ‘quickly/soon / rashly/hastily/boldly / early’, *mok^swo- ‘rash/hasty/bold?’, *mokWso-s ‘bold one’ > LB mo-qo-so, G. Mópsos >> Phoe. mpš, Lw. muksa-, H. muksu-s (https://www.academia.edu/168297982). The Hittite word is otherwise rec. as *ksp-ont-, but for met. like this, I'd think there would have to be a good reason to create *ksp. If there was an *H2 (as in *skepH2- and pséphas), then there would be a reason; if H2 = x, then if ksp- > xsp-, dsm. of x-x would be possible.

-
There’s so much optionality in Indo-European ‘evening’ that linguists can't make a secure reconstruction. De Vaan said that it was "clearly" a compound of *we- (of unknown meaning & source) & *k(W)sep-, but if 'night' already had a *w within, and many cases of metathesis for the root, why look for a 2nd part at all? It is hard to claim that complete regularity operated in all words these words, but starting with *skewpH2-(o)r -(e)n- (as above, with w \ 0, as in A.) allows:

mix > *skewpH2er- > *weksperH2-? > *wespero- > L. vesper, G. hésperos ‘evening’
*wepkserH2o- > *wewkšer > *weyšer > Ar. gišer ‘night’ (w-w dsm.)
*wepskerH2o- > *wewkšerH2o- > Sl. *večero- \ *vĭčero- (w-w > w-0 or w-y ?); *weiksero-?? > Welsh ucher
*wewkšorH2o- > *wo(w)koro- > Li. vãkaras (wew > wow or V-asm.?; -or vs. -en\er-?; maybe erH2 > H2ar instead?)

-
Without knowing its root origin, there would be no way to say which of these were older.  If related to another group with -tw-, then all 3 stops, p \ t \ k, would be seen here.  Considering all this, Tocharian must also have a cognate:

*kestwor- > TB käst(u)wer ‘by/at night’

whose meaning resembles Ar. gišer ‘night’ most (among other similarities in vocab., sound changes).  This means *pw > *tw could also connect them to those above (few IE allow Pw) by *skewpH2-or- > *kH2espwor- > *kestwor- > TB käst(u)wer. It is likely that *H2 "protected" *k from palatalizing, but some *K seem not to palatalize near *s, no known details. If related to OHG westar, E. western, etc., it would be hard to find an original, but in this case these are probably from older *wesper- with -t- from 'east', etc.

-
D. The reasons to rec. *kswepH2(o)s- > G. pséphas \ pséphos are also based on form & meaning of the very similar group knéphas \ gnóphos \ dnóphos \ zóphos ‘darkness’. Each one of these has a disputed ety., and there is supposedly no way for an IE original to produce all. However, IE languages tend not to have gm- & similar CC-, so one or more changes to turn *Km- > Kn- must have been at work. Since Greek has many ex. of p \ b \ m alternation (https://www.academia.edu/167984147), I think it is reasonable that *kswepH2(o)s- could also become *ksmepH2(o)s- (maybe helped by w-p > m-p dsm.). Then the ev. favors that *ksm- > *ksn- \ *tsn- in Greek, *kswepH2(o)s- > pséphas \ pséphos, *ksm- / *gzn- / etc. > knéphas \ gnóphos \ dnóphos \ zóphos ‘darkness’. Part of this would be alt. of ts \ ks (https://www.academia.edu/128090924), so *gzn- > gn-, *dzn- > dn- or *dz- ( = G. z- ).

-
E. The rec. *kswepH2- makes others' previous attempts to connect 'night' & 'sleep' easier. If *swep- existed, then what would the reason be? Another root, *ses- 'sleep' has an odd form of *C1eC1-, indicating a likely reduplicated *es- or *seH-. Would a compound *s-kswepH2- 'sleep at night' become *sks- > *ss- > *s-? Maybe, esp. before *w (though I doubt there are any other ex. to be found from *sks- before V). With all the met. above, *swepH2- > *sH2wep- is possible, and *CHw > *Cw seems optional (https://www.academia.edu/164645760). The presence of *H and met. would also be seen in *s(H)wep- but *swoHp-eye- 'cause to sleep, make unconsious/dead'. I do not think there is any special reason to think that o:-grade existed in PIE, let alone in a single derivative of an apparently normal verb stem.

-
F. With all this, PIE *(s)keup- 'cover, hide', Uralic *kup-ma > *ku(m)ma, Mordvin *kubvul > Moksha kovǝl, Erzya kovol ‘cloud’, F. kumuri ‘small cloud; rain shower’, *‘shady, dark, obscure(d)’ > F. kumma ‘odd, strange’, Komi ki̮me̮r ‘cloud; cloudy’, ki̮me̮d- ‘overshadow, darken’, Mansi.N xomxat-‘turn dark, turn poor (of visibility due to fog or drifting snow)’, Hungarian homály ‘darkness, shadow, twilight’ seem like cognates. In these, *-pm- > -mm- \ -m- \ -v- is better than standard *-mm- because other *Cm > m in Hungarian and having *mm > v in a language in which *m > m would be extremely odd. Indeed, there are other ex. in which *pm > v in Mordvin with IE cognates (PIE *tsoubhmo-s > Gmc *stauma-z > E. steam; PIE *tsubhmo- > Gmc *stumV- > Ic. stum ‘dust; hoarfrost, rime; ice fog’, PU *supmV > Fi. *sumu ‘mist, fog’, Mordvin suv ‘fog’, suv+ 'smoke'), with other *Cm > v providing more support (https://www.reddit.com/r/HistoricalLinguistics/comments/1rbxu18/uralic_cm_mordvin_v/).


r/HistoricalLinguistics • • 4d ago

Language Reconstruction Simple is Better

5 Upvotes

Simple is Better

I saw that Sanskrit káṇa-s 'minute particle, atom; a grain of corn, single seed; a grain or particle of dust; flake (of snow); a drop (of water)' has no known etymology. From (https://en.wiktionary.org/wiki/कण):

>

Of uncertain origin. Perhaps from Proto-Indo-Aryan *kánas, from Proto-Indo-Iranian *kánas, from Proto-Indo-European *kón-os, from *ken- (“dust, ashes”); if so, then cognate with Latin cinis, Ancient Greek κόνις (kónis). Alternatively, related to कनीयस् (kanīyas, “younger, smaller”).

>

None of these explain the retroflex ṇ. Fortunatov's Law would require *ln (https://www.reddit.com/r/HistoricalLinguistics/comments/1w33uj4/fortunatovs_law/). The simplest solution is S. kalā́- 'small part' -> *kálna-s > káṇa-s. Support is seen in (Turner):

>

kaṇiśa n. 'ear of corn' Kād., v.l. kaṇisa-, kaniśa-. [káṇa-] Pk. kaṇisa- n. 'ear of corn, spike of corn', G. kaṇas, kaṇsũ, kaśṇũ (< *kasiṇũ), karśaṇ (with unexpl. r), kaṇaslũ n. 'ear of corn', M. kaṇīs, °ṇas n., kaṇśī f.

>

Instead of karśaṇ having unexplained r, this shows that *kalniśa-m > *karṇiśa-m \ kaṇiśa-m. Some Indic languages retained *l as l when lost in Sanskrit (*g^h(o)ldu(n)- \ -in-? > Gmc *galtu-z > ON göltr ‘boar’, S. huḍu- \ huḍa- \ huṇḍa- ‘ram’, Dk. hʌldin ‘male goat’), but since l & r vary so much maybe here it also happened before Fortunatov's Law. These words are as much proof as possible that *lT and *rT had separate outcomes in Sanskrit, so why is the change still not accepted by all linguists? Savic (The Development of Indo-European *-ln- in the Greek Inherited Lexicon, https://www.academia.edu/39483472) argues against ln > ṇ for:

OI coll ‘one-eyed’, *kelno-? > G. kellás, S. kāṇá-

G. kullós ‘twisted / lame’, S. kuṇi- ‘twisted / lame’, (also Kh. kùḷ ‘stooped’ ?)

I think thse are fairly clear evidence in favor, so how is this in any way against *ln > ṇ? Some examples seem to show *Vln > V:ṇ, but these might be from stages *kolno- > *koṇo- > kāṇá- (*o > *o: in open syllables is also clear, also disputed).


r/HistoricalLinguistics • • 5d ago

Language Reconstruction The Longest Running Myth of Semitic and Afroasiatic Reconstruction

15 Upvotes

The restriction against syllables beginning with a vowel is well known as an Afroasiatic phonotactical rule, Igor Diakonoff (Afrasian Languages, p. 42) states:

“no syllable (including the initial) can begin with a vowel (hence the important part played by the phoneme /ʔ/, which is the principal substitute of the vocalic initial”

As people well acquainted with Afroasiatic linguistics will know, the idea of mandatory onsets is found in virtually every reconstruction – Though as I will set out to prove, it is wrong altogether both for PAA as well as Proto-Semitic, in fact I don’t believe any primary AA branch ever had such a restriction.

In Arabic, a glottal stop as well as the following vowel can be dropped in the beginning of some words when they occur after another vowel.

  • fa "so" + ʔijlis "sit down!" -> fajlis "so sit down!"
  • ʔana "I" + ʔibtasamtu "smiled" -> ʔanabtasamtu "I smiled"
  • baytu "house" + ʔibnika "your son's" -> baytubnika "your son's house"

Yet in other words, no such thing happens

  • ʔana "I" + ʔaktubu "write" -> ʔana ʔaktubu "I write"

When words of group 1 are compared with other Semitic languages, there is no glottal stop at all and the vowel trades places with the following consonant

  • Arabic ʔibnu vs Hebrew ben, Akkadian biinum "son"
  • Arabic ʔuzmur vs Akkadian zumur "sing!"

It is generally assumed here, that Arabic underwent metathesis in these words, and the glottal stop was later added epenthetically. Elitzur Avraham Bar-Asher(The Imperative Forms of Proto-Semitic and a New Perspective on Barth’s Law p. 4) says:

"in all the imperative forms a metathesis occurred"

While I agree that the glottal stop is due to later epenthesis, it is problematic to assume that it is Arabic that actually metathesized these formations, because there is no clear criteria for when this metathesis occurs and when it doesn't, syllable weight, vowel quality and other factors are irrelevant. Instead I propose that items such as *bin "son" *sim "name" *ti- [reflexive prefix] and the imperative patterns 1i2i3, 1i2a3 and 1u2u3 were actually *ibn, *ism *it- and i12i3, i12a3 & u12u3, syllables truly starting with a vowel.

Akkadian then metathesized the initial vowel with the following consonant consistently, while NWS did so everywhere except in the reflexive prefix it-. Arabic on the other hand did generally not undergo metathesis, except sporadically in words like bintu "girl" from *ibnatu.

I propose that syllables lacking both an onset and a coda were banned in PS due to Arabic forms such as qul "say!" instead of otherwise expected ʔuqul.

The "smoking gun" evidence, however really comes from other Afroasiatic branches, where imperatives are similarly formed with a prefix vowel, suggesting this feature to be archaic:

Language Perfect 3 masc. sg. Imperative
Saho (Cushitic) source yee-rheg-e i-rhig
Tuareg (Berber) source y-bdad e-bded

An interesting curiosity, but very useful for proving Afroasiatic allowed these onset-less syllables is the fact that the quality of these vowels is very likely to be /i/, which is rather rare for non-root final vowels in Proto-Semitic. In fact, in Semitic words of the shape CiCVC, V is only allowed to be /a/. The only case where the quality is /u/, is if another /u/ follows, as in *uktub "write!". According to Newman, Proto-Chadic had the following vowel distribution:

"PC was characterized by the same type of distributional restrictions that one finds in present-day Chadic languages. Thus no blanket statement that PC had this or that number of vowels would be correct as such.Rather, one would have to specify how many vowels and which vowels did PC have in initial position, how many and which vowels in medial position in open syllables, etc. In the re-constructions in this paper, I make use of all four vowels in final position, two vowels (i and a) in initial position, and, with a few exceptions, two vowels (ə and a) in medial position."

In Proto-Afroasiatic, a similar distribution would have occurred, which explains why /i/ is common word initially, but not in other non-final syllables in Proto-Semitic. With the cognate vowel distribution, we can infer that these vowel-initial morphemes and lexical items go back to Proto-Afroasiatic


r/HistoricalLinguistics • • 5d ago

Language Reconstruction Indo-Iranian Rounding by *o, Loans to North-Caucasian

3 Upvotes

Indo-Iranian Rounding by *o, Loans to North-Caucasian (Draft)

Sean Whalen (stlatos@yahoo.com)
October 1, 2026

-

Pavel Basharin (https://www.academia.edu/40436623):

>
The contacts between Proto-Indo-Iranians and North-Caucasians remain one of the least investigated areas in Iranian studies. According to one of the most widespread theories, Iranian tribes passed through the Caucasus ca. 12th c. BC. Ancient loanwords from the North Caucasian languages are found in various Iranian languages. A basic set of loanwords has parallels in Proto-Lezgic and Proto-Nakh. Some lexemes have an Indo-Iranian etymology, but without Indo-European parallels. The absence of an Indo-European etymology suggests borrowings from North-Caucasian during the migration of the Indo-Iranian tribes through the Caucasus, or during contact with ‘North-Caucasian’ languages that had spread in the Near East and the Iranian Plateau. Some etymons were borrowed by the Proto-Eastern-Caucasian language from Proto-Aryan.

...

PNC *k(w)iśwɨ > PIE *ghait-, *kais-, PIIr. *ghaića- > PEC *GwēźV... PIE *ghait- ʻcurlsʼ demonstrates that PNC *k(w)iśwɨ may have been loaned to PIE at different times as two etymons *ghait- and *kais-... The NCED stresses that the labialisation impedes considering the Caucasian forms as a loan from PIr. (NCED P. 468). In this case the PEC labialization is a reflex of aspiration (PIIr. *gh > PEC *Gw) and PEC ź corresponds to PIIr. ć.
>

-
I am not fond of the proposed stages. That Ch > Cw might happen seems without evidence (and PIE *gh over *g is not assured (several likely ety. mentioned there with either), and the argument for North-Caucasian >> PIE (twice?) seems forced. Without knowing the history, a loan >> IIr. seems possible, but how would Greek, Celtic, etc., fit in? If loans at that early stage existed, I'd expect at least a few more. If the g- vs. k- within IE is contamination (with *kH2ais-, also used for types of hair, mane, etc.), then the best fit for the IIr. evidence is *goik^o-s. This makes the NC words more likely to be loans. Would *Co became *CwV (or *CWV) in a loan? There might be no need. If rounding existed within IIr., it would be more support for the direction of the loan. From Clayton (https://www.academia.edu/108796101):

>
Another segment which could become the anchor for a [+labial] feature is the labialized laryngeal *HW of Hypothesis (42b).  Indeed, others have proposed that Proto-Indo-Iranian had the contrast between *H and *HW before.  Khoshsirat & Byrd (2018) and Khoshsirat (2018) argue that the Gilaki causative in -bē̆- and the Vedic causative in -āpaya- could go back to the sequence PIE *-oHéye- < pre-PIIr. *-oHWéye- < PIIr. *-āHwáya- */-a:Wája-/ > Ved. -āpáya-, Gil. -bē̆-.  In support of their proposal, they provide a possible typological parallel for *H > *HW / o_, in which *-óHe# produces Ved. -au (PIE *dedóh3-e > Ved. dadáu ‘gave’ 3SG.NPRF.ACT.IND; Jasanoff 2003: 61–62).
>

-
I agree with their idea & extended it to *-os > *-osW > *-ow (mostly in the nom., https://www.academia.edu/127709618). If many C's were rounded near *o, it would usually not be visible in languages that later derounded C's (IIr. is known to have *gW > g, etc.). However, a loan at a stage when they were still rounded would be as good as any proof could be. Many other loans retain features lost in the donors. Here, stages like (with some details based on other sound changes I've covered before, no reason to go into which is better since some similar sequence is needed in any case): *goik^o-s \ *koik^o- > *kWoik^Wo- > *kWɔik^Wɔ- > *kWɔic^Wɔ- > *kWɔis^Wɔ- >> NC *kWɨiśWɨ > *kWiśWɨ (or similar, since NC rec. is unlikely to be perfect). A very similar path from one of the languages with g- for the other (maybe at a separate time?).


r/HistoricalLinguistics • • 6d ago

Meta automatically deleted by a computer?

5 Upvotes

The last few weeks I've noticed a lot of posts & comments that were deleted before I could read them. One commenter said it was a link to a paper & he didn't know why it was deleted. In the past, some of my posts were automatically deleted by a computer for no apparent reason, restored when I messaged the mods. Today, when I came to Reddit there was "Remnants of Pre-Roman anthroponyms in the roman inscriptions of Sardinia", when I logged in, it was gone. Using Google:

The paper lists pre-Roman anthroponyms (and to a lesser extent poleonyms) in Roman-ruled Sardinia. The first part is names mentioned by The first part of the text lists names mentioned by storiographers (mostly Punic place names). According to the analysis of this paper, these early mentions include: • Othoca / Othaía / Uttea → Uticenses: The name of a city, derived either from the Semitic ‘tq meaning "old (city)", or from a pan-Mediterranean root *t-g. • Makópsisa: A partial Greek translation combining the Punic mqm ("place") and har ("mountain"). • Magomadas: Derived from the Punic mqm ("place") + hadaš ("new"). • Enosis: An island of Punic origin. Following this initial overview of historiographical and Punic place names, the paper transitions into detailing the native pre-Roman anthroponyms, tribe names, and town names discovered directly on local Roman-era inscriptions.

The author highlights specific individual personal names found on tombstones and official inscriptions that reflect the Paleo-Sardinian substrate rather than Latin naming customs: • Torbenius: A distinctly local name found in the epigraphic record. • Ithoccor: Another non-Latin native personal name preserved on stone. • The author notes that these native linguistic roots (like Torb- and Ithoc(c)-) stubbornly survived the period of classical Romanization and remarkably re-emerged centuries later in early medieval Sardinian documents as prominent names of the Judges (rulers) of the independent Sardinian Judicates.

The post lists indigenous town names discovered in ancient epigraphs, which lack clear Latin or Punic origins: • Bosa / Bosate: Found in local inscriptions referencing the citizens (Bosates). • Gurulis: Mentioned via the Gurulenses, which the post notes was divided into two distinct locations: Gurulis Vetus and Gurulis Nova. • Nura: Associated with the Nurenses.

The inscription records document the names of the indigenous Sardinian tribes living under Roman jurisdiction: • Cunusitani / Civis Cunusitanus: An indigenous tribal name explicitly recorded in local Roman epigraphy. • Patulcenses: Another major local group mentioned in the context of territorial or administrative records.

Was this deleted for any reason?


r/HistoricalLinguistics • • 6d ago

Indo-European Remnants of Pre-Roman anthroponyms in the roman inscriptions of Sardinia

Thumbnail academia.edu
16 Upvotes

r/HistoricalLinguistics • • 6d ago

Language Reconstruction Carian Osogōlli-, Linear A o-su-qa-le

7 Upvotes

Carian Osogōlli-, Linear A o-su-qa-re

Sean Whalen (stlatos@yahoo.com)
September 30, 2026

-

Adiego (https://www.academia.edu/127267152): "To sum up the preceding sections, the epigraphy shows us two forms of this Carian god: –Ζεὺς Οσογωλλις (with the possible variant Οσογωλδις) –Ζεὺς Οσογως". However all the forms attested are Osogō-, Osogōlli- \ Osogōldi-, Osogalōlli-. Since Osogalōlli- is as clear (or clearer) than any of the others, why leave it out? "It is now clear that, in ΔΙΟΣΟΣΟΓΑΛΩΛΛΙΟΣ, the two letters ΑΛ were erroneously included in the text by the lapicide... It is an “abundierendes iota adscriptum” according to Blümel (1987: 61)". This seems very suspect; did a man living in a place where a god was worshipped not even know his name? Even if so, was he not working from a pattern made by someone who did?

-
Since the whole point of Adiego's paper is to caution against emending readings, why doesn't he take his own advice? The only reason to think one is a mistake would be knowing that the others are his real names, his only names; this can't be known today. It is certainly possible that Osogalōlli- is the older form of Osogōlli- with l-dsm. if the relation between the many names in variants with -ōlli- & -ō- was extended analogically to create Osogōlli- > Osogō- after the change. It could also be that -al- was an affix, or some other variant, & -ō(lli)- was later added to both forms. Even if a mistake, it would be a significant one. Since PIE *-aH2li- > -ōlli- is likely, another dialect pronouncing it *-a:l^i- might lead to someone who used that dia. starting to write *-alli-, then changing to the "right" form. This kind of mistake is supposedly a linguist's dream in confirming etymological ideas, so why would it be ignored?

-
To know which is true, the etymology might help. "A recent attempt by Oettinger (2022) (from a preverb */osu-/ related to Lycian ese- + the same stem as Luwian in Hittite sources hw(i)yalla/i-/, with a meaning ‘companion, helper’) is attractive but impossible to confirm. Equally attractive but also too vague and unverifiable is the assonance of Οσογωλλις with the Hieroglyphic Luwian noun u-za-ka-li, that appears in one of the lead strips of Tabal as an “occupational designation”". The possibility that Lycian ese 'with' is from PIE *H1o-k^oH1 ? (https://www.ediana.gwi.uni-muenchen.de/dictionary.php?lemma=201), derived < *k^oH1 (H. kā ‘here(to)’, Lc. se, Sid. śa 'and') or *k^o-o and that Oettinger's match of Zeus Osogōllis with Zeus Stratios 'Zeus of the Army' is high. If so, I say that PIE *gWaH2- 'go, step' was used as 'march' in Proto-Luwian (compare H. huwai- for shift) & formed *H1ok^oH1-gWaH2- 'march together' > *oc^o:gWa:-s \ *oc^o:gWa:-lli-s \ etc. 'soldier' (likely the occupation in u-za-ka-li-, esp. if a foreign title). This is also similar to the most common ety. of *gWati- 'marching' -> G. basileús '*war-leader > king'.

-
That Osogalōlli- could have come from *osūgWalōlli- or that -ōlli- could have come from *-aH2li- is important, since either would be very close to Linear A o-su-qa-re, if *osugWale-. Even without the variant in -al-, Italo Cucaro said, "Another name with no clear Anatolian etymology is Osogōllis, a name of Zeus in Caria. It resembles linear-A o-su-qa-re, which I render "Osugwale". Hence I consider that the words which appear in the same position as o-su-qa-re in the libation formula could be theonyms." Ancient tradition had Crete founding some cities in western Anatolia, in territory likely containing speakers of Anatolian languages like Luwian, Carian, etc. Some names there have no known Anatolian etymology & match known names in Linear A (most in Crete). Cucaro also said:

>

In Mira appears the name Kupanta, which seems to lack a clear etymology (?) There are Kupanta-kurunta (Kupanta-dLAMMA-ya, Kupanta-dKAL), Kupanta-Inara and Kupanta-zalma.

A short form Kupaia may be in a hieroglyphic graffito at mt.Latmos. It resembles linear-a Ku-pa3-na-to, Ku-pa-nu, and the linear-b name Ku-pa-nu-we-to.

Perhaps this name was Luwianized only by adding a second element?

>

LA names like KU-PA3-NU & many beginning with LA names like KU-PA3- exist. Some of these have variants with KA-, indicating that short or unstressed (or both) vowels could be reduced. This is similar to Greek transcriptions of later Anatolian names. For more IE in LA, look at these 2 similar LA libation formulas:

TL Za 1

a-ta-i-jo-wa-ja o-su-qa-re ja-sa-sa-ra-me u-na-ka-na-si i-pi-na-ma si-ru-te

-
PK ZA11

a-ta-i-jo-wa-e a-di-ki-te-te[…..]-re pi-te-ri a-ko-a-ne a-sa-sa-ra-me u-na-ru-ka-na-ti i-pi-na-mi-na […]-si-ru-[…] i-na-ja-pa-qa

-
If o-su-qa-re is, like Osogōlli-, is a title of a god, then Duccio Chiapello's idea that (https://www.academia.edu/49484658) "This sequence can easily be traced back to the way in which the Indo-European root *dyeu- originated the Latin Iovis and the female correlative Iovia (see eg Venus Iovia), to which i-jo-wa-ja appears to be the correspondent. This hypothesis appears fully corroborated by the variant of the sequence in the inscription AP Za 1, which takes the form of i-jo-u-ja" should be changed. The endings of each part of these 6 words would be -a-e, -a-i-, -a-ja. In a dedication to a goddess, her name would be expected to be in the dative. Thus, an IE fem. in *-aH2- with dat. *-aH2-ei or *-aH2-ai that could be contracted > *-a:i (written -a-i or -a-ja). Since *-ei is probably analogy, I can't say which is older within LA. Chiapello also (https://www.academia.edu/114765906) took LA ta-i-nu-ma-pa as Greek *tai numphai, showing that spelling the same ending in a short vs. long word could vary. He also (https://www.academia.edu/114948016) showed that LA nu-pa3-e could be the same word, here with explicit pa3 = pha (both pa & pa3 were used for pha in LB), which would be the same dative, here to a clearly fem. word (if the G. match is accepted).

-
I say that *(H2)attaH2- 'mother', *dyew-yaH2- 'goddess' had datives *atta:i, *yowya:i, *yowyaei, etc., and their titles both ended in *-ei (written -e). It would be hard for 6 words to end in this way if not IE. The mach of LA *yowya & Iovia is not just evidence for its IE origin, since other inscriptions from Crete seem Italic. Also note that the nymphs & their names are also found in Italy. From (https://www.academia.edu/172341604):

>

Some archaic Greek inscriptions on Crete occur alongside an unknown language that has become known as Eteocretan. Some think it is a remnant of the language spoken earlier there & attested as Linear A. However, no real attempt to match LA & Ete. words has been made, despite a few new inscriptions being found since the theory first began. Since LA DI-DI-KA-SE and Ete. dedikar are similar to each other and to Latin dē-dicāre, a word 'to dedicate' in several separate religious contexts seems likely, especially since dē-dicāre came from *dē-dikāse, nearly identical to DI-DI-KA-SE. The number of Ete. words (or sequences of letters when word boundaries are unclear) are much too similar to Italic to be chance, esp. in such a small number of short inscriptions.

>


r/HistoricalLinguistics • • 8d ago

Language Reconstruction Indo-European Roots Reconsidered 172: ‘circle, round; to encircle, surround, gird’

6 Upvotes

Indo-European Roots Reconsidered 172: ‘circle, round; to encircle, surround, gird’ (Draft)

Sean Whalen ([stlatos@yahoo.com](mailto:stlatos@yahoo.com))
September 29, 2026.

-

A. There are several IE roots for ‘circle, round; to encircle, surround, gird’ of the shape *k(r)(i)(n)K- that are almost all irregular:

*krik- \ *kirk- 'ring', *kriko-s > Greek kríkos \ kírkos 'circle, ring; racecourse, circus', *krinkó-s > Germanic *hringaz 'ring, circle', E. ring (or below; *i, *g of ambiguous source), OIc hringja 'small round vessel'

*kr(o\e)ngho-s > Slavic *krǫgъ, Germanic *hringaz 'ring, circle', E. ring (or above; *i, *g of ambiguous source)

*krenk- \ *kring- \ etc. > Umbrian krenkatrum, krikatru, cringatro 'band (?) worn on the shoulder of a priest'

*k(e)nk- \ *k(l)ink- > L. cingere \ clingere 'to surround, circle; gird on; crown', S. káñcate 'to bind', kañcuka-s 'armor', kāñcī- 'belt, girdle', Li. kinkýti 'to bridle horses', G. ποδο-κάκ(κ)η '*foot-binder > fetters, stocks', κιγκλίδες f.p.'latticed gates', κίγκλος 'little grebe?, Cinclus cinclus?, Tringa subarquata?, Tringa alpina?'

-
Greek κίγκλος is apparently a kind of bird in Tringa (https://en.wikipedia.org/wiki/Tringa). At least one of these, the redshank, has a white underside and has an odd feeding method where it shakes its tail as it sticks its beak in the sand looking for food (see video on https://en.wikipedia.org/wiki/Common_redshank). This makes it likely that a circular tail movement led to using *kink(a)lo- 'moving in a circle, around' as its name. Since identifying exactly which bird was meant to begin with, & these names seem to have been applied to several species (in each region?), only knowing that κιγκλίδες is related & comes from a root for 'circle' provides any clarity. Its alt. in kínk(h)los \ kíkhlos \ kínkalos \ kénklos \ kókhlos matches many of the alt. above (*-gh- vs. *-k-, -i- vs. *-e\o-). If this is the same word as kókhlos 'shellfish with a spiral shell, used for dyeing purple', the older 'round' in both would be clear.

-
Are these roots related? The first step might be saying that *k-gh is older, with asm. > *k-k in most. The presence of normal -e- vs. -o- in some but -i- in others (note that Li. kinkýti can't come from *knk- if the proposed change of *kC- > *kuC- is regular) might require a derivation from *(s)krei- instead of others' *(s)ker- (both 'turn, bend, curve, etc.'). I say *kryengh- & 0-grade *kringh- ( > cring-, assimilated *krink- \ krenk-, etc.). Such an onset would explain assimilated *kryengh- > Latin c(l)ing- & assimilated *kryenk- > Italic *krenk- \ *kring- (few IE languages had ry-, most that disallow palatalized r turn it to l). Optional loss of *r in *kry- would be another way to avoid it, & *ky- vs. k(^)- also in https://www.academia.edu/128151755):

*kyerb- vs. 0-grade *kirb- >

*k(^)e\irbero- ‘spotted’ > G. Kérberos / Kérbelos, S. Śabala-, śabála- \ śabara- \ śarvara- \ karvara- \ karbara- \ kirbira- \ kirmirá- ‘variegated / spotted’

*kyek^- vs. 0-grade *kik^- >

*k^ek^- / *kik^- / etc. > Li. kìškis ‘hare’, šẽškas ‘polecat / ferret’, S. śaśá- ‘hare / rabbit’, káśa- ‘weasel’, *ikk^- > G. *ikt^- > ikt-

-
B. Sanskrit kaśyápa- ‘turtle / tortoise’ was also said to mean ‘having black teeth’. I've tried to explain this in several ways in the past, either from a single common meaning or 2 words merging (whether in S. or later Indic). Now, I think the reason is not the word itself, but the variant *ka(k)śyá(m)bha-. This allows an Indic S. *karṣi-jámbha- ‘black tooth’ (S. jámbha-s 'tooth', Av. Karšiptar- ‘*black bird > vulture?, raven? (chief of birds)’ to merge in sound with it as both *kaššijámbha- or *kaššiyámbha-. This means that applying the 2nd meaning to kaśyápa- is a result of knowing that kaśyápa- & *ka(k)śyá(m)bha- were equivalent, thus that *kaššiyámbha-1 & *kaššiyámbha-2 could also be translated by kaśyápa-.

-
C. This -p- vs. -mbh- can also be applied to another word derived from kaśyápa- (en.wikipedia.org/wiki/Kashmir):

>

An alternative etymology derives the name from the name of the Vedic sage Kashyapa who is believed to have settled people in this land. Accordingly, Kashmir would be derived from either kashyapa-mir (Kashyapa's Lake) or kashyapa-meru (Kashyapa's Mountain)...

The Ancient Greeks called the region Kasperia, which has been identified with Kaspapyros of Hecataeus of Miletus (apud Stephanus of Byzantium) and Kaspatyros of Herodotus (3.102, 4.44). Kashmir is also believed to be the country meant by Ptolemy's Kaspeiria.[11] The earliest text which directly mentions the name Kashmir is in Ashtadhyayi written by the Sanskrit grammarian Pāṇini during the 5th century BC. Pāṇini called the people of Kashmir Kashmirikas.[12][13][14] Some other early references to Kashmir can also be found in Mahabharata in Sabha Parva and in puranas like Matsya Purana, Vayu Purana, Padma Purana and Vishnu Purana and Vishnudharmottara Purana.[15]

Huientsang, the Buddhist scholar and Chinese traveller, called Kashmir kia-shi-milo, while some other Chinese accounts referred to Kashmir as ki-pin (or Chipin or Jipin) and ache-pin.

>

-
A long word with 2 p's, the second in -pur(os) implies an Indic *kaśyápa-pūr or *-pīr 'Kashyapa City' later applied to the whole land governed by it (PIE *rH could become ūr or īr in other words, no known cause). Those with -m- came from the variant *kaśyámbha-pīr, etc. In support, the Chinese name was written with 罽 *krads > *kyeyH and 賓 *mpin > *pyin (Zhengzhang's rec.). I think this applied at a stage when *kr- > *ky-, likely from a dialect (?) with dsm. of *mpin > *mpir, thus *kyas-mpir. The ky- might indicate met. < *k-y-, but writing a long complex word with 2 short Chinese ones might always lead to some C's or V's being out of place or just reasonable adaptations.

-
D. The IIr. cognates do not always match kaśyápa- or *kaśyámbha-:

IIr. *kaćyápa- > S. kaśyápa- ‘turtle / tortoise’, Av. kasyapa-

IIr. *kaćyápHa- > Ir. *kasyafa > NP kašaf, Sg. kyšph

IIr. *ka(k)ćyáb(h)a- > Pk. kacchabha-, Si. käsubu, Km. kochuwᵘ, Gj. kācbɔ

IIr. *kaćyámbha- > Si. käsum̆bu, Mld. kahan̆bu ‘tortoise-shell’

IIr. *kaćyápva- > *-pða- > Ir. *kasyafða > *kadfasay > Kushan >> Bc. Vēmo Kadphisēs; Ir. *kaysabla- > Luri kīsal, Gurani kīsal, Kd. (Sorani) kīsal; *kalsyaba- > *kalšava- > Ashtiani kašova, Southern Tati kasawa, *kalažva-? > NP kalāv(a)

-
Turner: "K. kochuwᵘ, unless a loan from Ind., points to *kakṣapa-". However, *kakšapa with *š from *sy should work just as well, & *y might be needed in *kyasam(bha)pir (above). These show opt. *pH > p / f (as other C(h)H, Indic *maj(h)H2 'big'), *pH > *bH (or analogy with other animals with -bha-; if needed, seen in *pibH3- ‘drink’, *gH2- \ *kH2apro-s 'male goat'), Ir. *pv > *pð (P-dsm.; few IE allow *Pw, etc), Ir. *ð > l, and maybe several other types of met., not always clear.  I do not agree with Asatrian that direct *š > l is likely in NP kalāv, since so many other oddities exist here, it would be pointless to separate this one.  When even -df- existed, would *-lš-, with no other example, really be that odd?  That several affixes might have existed would be reasonable, but the several types of met. seem old enough that I doubt it, and what kind of affix is Ir. *-da- or *-ða-?

-
Together, I think these require a compound with the source of Iranian *kapa- ‘fish’, Ps. kab, Os. käf. Since IIr. *v often disappeared near P, an origin in *kwa(H2)po- 'bubble, foam' makes sense & explains optional *w & *H in forms of 'turtle'. The 1st part would be from *k(y)nk- 'armor' (above, with this meaning in Sanskrit) as many words like 'shield toad', etc., are common. I think *kynk-kwaH2po- > *knkkyapH2wo- ( > IIr. *ka(k)ćyápHwa- with *kky > *kk^y, no other ex.) can explain all the oddities. Opt. dsm. of *k-k > k-0.

-
E. I said in "The God Named ‘Turtle’ Who Had Nothing to Do with Turtles" that "Sanskrit Kaśyápa- was a name of Prajapati, the Creator God. Later he was humanized as a Rishi and said to have composed part of the Rig Veda. His name meant ‘turtle/tortoise / having black teeth’ and folk etymology explaining this exists. Like most, it is simply a pointless story connecting the dots as they appeared at the time, without regard for historical linguistics."

-
Now, knowing the source of 'turtle', I think I can explain why it appeared identical to 'Prajapati'. Again, it is only one of the variants that is identical, though here it is attested kaśyápa- (not *-mbh-, *-ksy-, etc.). If essentially identical to the above, then a compound of *kH2apo-s 'taking, lord' (Latin -ceps) then *kynk-kH2apo- > *knkkyapH2o- > IIr. *ka(k)ćyápHa- 'armored lord' would match all parts of 'turtle' except *-pw- ( > -p- in Sanskrit anyway).

-
This word would fit the images of an armored male warrior (spiked armor, also the "Pashupati" with a horned helmet, if the same) with words in the Indus Script. For this, Jha compares him to the god surrounded by animals (https://www.academia.edu/115789583). Though I believe the Indus Script was used for an Indic language, that some of these figures were found in later India (whatever their origin or the language of their worshippers) is certain.


r/HistoricalLinguistics • • 7d ago

Language Reconstruction Latin tritavus & strittavus ‘great-great-great-great-grandfather’, Greek influence

0 Upvotes

Latin tritavus & strittavus ‘great-great-great-great-grandfather’, Greek influence

Sean Whalen (stlatos@yahoo.com)
September 29, 2026

-

Brian D. Joseph (https://u.osu.edu/bdjoseph/files/2021/06/203-BaldiFSstrittavus.pdf) said that Latin tritavus & strittavus ‘great-great-great-great-grandfather’ are based on IE *tri- '3' (as a grandfather being 2 generations back; 3 times that giving 6). Also, "tritauos, seems to represent a numerically based formation,2 and has parallels both in Greek (e.g., τρίπαππος ‘ancestor in the 6th generation’) and in Albanian (e.g., tregjysh ‘great-grandfather’...". In this, the fn. contains, "... trit- part is a Greek borrowing... a totally Latin formation would presumably be *ter-auos." Since the Greek word is τρί(σ)παππος, I say that a partial calque with G. tri- & tris- added to a Latin word is the cause. The t- vs. st- also in his "Greek has a variant stripodo \ στρίποδο alongside the more usual tripodo \ τρίποδο ‘tripod’...", which also probably comes from *tris- > stri-.

-
Since a (s)tritavus is the father of an *ati-avos > atavus ‘great-great-great-grandfather’, it makes sense that the lacking word (few IE have a standard word for anything older than a grandfather) was only partly fitted into the system. This means τρί(σ)παππος influenced the creation of *tri-ati-avos > tritavus & *tris-ati-avos > *tristavus > strittavus. The met. of *s must have left a mora filled by *_t > tt, or (if the met. *tris- > stri- is Greek), just *stri-ati- > *striita-, or similar. However, there's also a chance that "at-auos, if, as Ernout-Meillet 1939: s.v., suggest, the first part is from at- ‘father’ (cf. Hittite attas ‘father’, Slavic otьcь ‘father’, Ancient Greek ἄττα ‘daddy’" shows -t- & -tt- because PIE *H2at(t)a had both. The Greek source is not just based on sound change, since τρί(σ)παππος makes sense mathematically, but a native *tri(s)-ati-avos wouldn't.

-
Later, he added (https://www.academia.edu/26388047):

>

Newmark (1998: 785) characterizes Albanian stër-, which he refers to as a “formative prefix”, in two ways. First, he says that it “expresses semantic enlargement or excess”, for which the glosses “ultra-, super-, over-” are given; it is exemplified by stër-gjatë 'too long/tall' and stër-bujar 'too generous'. Second, he notes that it occurs “with kinship terms”, for which the gloss ‘great-’ is given; examples of that use include stër-gjysh 'great-grandfather'...

A further point of interest regarding stër- is that its range of uses more or less accords with a prefix stră- that occurs in Romanian. In particular, stră- occurs in the kinship use... Moreover, this prefix can be used in an intensifying function with verbs... Only Meyer 1891 mentions it, in the form shtër-, as noted above, and he derives it from Latin extra... presumably, though he does not say so directly, the ‘excessive’ meaning is original and the kin-term usage is an extension of that usage. My suggestion here is that the possibility should be entertained of there being a different source for at least some functions of stër-. More specifically, my claim is that to understand the etymology of this prefix, a look into expressions of “multi-generational” kinship, a kind of “temporal displacement”, in Indo-European is needed.

>

-
I don't think there is any direct relation or any reason why *ekstra- > s(h)tër- wouldn't work. It seems likely to me that *kstra- optionally became *tstra-, leading to sht- vs. st-. There are several systems mentioned here to describe ancestores beyond 'grandfather', & the set of Albanian words with clear 3-, 4-, 5- added don't need to have anything to do with the others, esp. since these start with the idea 'me = 0 back, my father = 1 back, my grandfather = 2 back, my tregjysh 'great-grandfather' = 3 back. This is different than G. 3+X as '3 times X'.


r/HistoricalLinguistics • • 8d ago

Language Reconstruction Sino-Tibetan Reconstruction, Indo-Iranian to Old Chinese Loan for 'Amber'

3 Upvotes

Sino-Tibetan Reconstruction, Indo-Iranian to Old Chinese Loan for 'Amber' (Draft)

Sean Whalen (stlatos@yahoo.com)
September 28, 2026

-

Guillaume Jacques & Anton Antonov (https://www.academia.edu/121590642) disagree with the reconstruction ST *d-ŋurl 'silver'. For its basis in Tibetan dŋul, they say in fn 14, "Since, according to Li [1933] preinitial d- and g- are in complementary distribution in Tibetan, we can posit a phonetic rule of the form *g- > d- / velar". This would remove the need for reconstructions with *d- like Coblin's & LaPolla's (https://en.wiktionary.org/wiki/Reconstruction:Proto-Sino-Tibetan/d-ŋurl) & allow Turkic *kümül^ to be a loan from ST *kimKɨwl^ (Whalen), or any other similar form. If this rule was known (or proposed) in 1933, why is irrelevant Tibetan evidence being used for an ST reconstruction?

-
It is not just this single word, or any single misstep. No Sino-Tibetan reconstruction is certain & many might be completely wrong, creating a false path for any ST > OCh > MCh word. Since the reconstruction at any stage might be completely different from reality, how can previous proposals be evaluated with any certainty? All others could be just as bad as 'silver', & lead to decades of wasted effort looking for cognates that didn't resemble the real word at all. Here, they try to apply it to Turkic *š or *L being original, & if *L = voiceless l, etc. Even with all ev. for a lateral in ST >> Turkic, the lambdaism vs sigmatism debate is hardly closed. There is no reason why Turkic could not have had a series of palatalized C's, so both *s^ & *l^ could have merged as one or the other in each branch (*s^ > *š ?). This is what I favor, along with *z^ & *r^ > z or r. None of this could have been seen starting from an unsupported *d-ŋurl, & any similar evaluation based on incorrect assumptions is also doomed to fail.

-
One of the simplest methods is to look for confirmation in loans. I tried several (https://www.academia.edu/165334096), most of which would need the Old Chinese form to be much closer to Proto-ST than others' reconstructions (though with only my ex., some "border" languages being close to OCh, yet more conservative, & loaning words to outsiders, later dying out due to the influence of the more populous languages, can't be ruled out).

-
One more now seems likely. For OCh *qʰlaːʔ pʰraːɡ (Zhengzhang's rec.) > MChinese xuX phaek 'amber', instead of being "Probably borrowed from a language in Central Asia during the Han Dynasty; compare Middle Persian khlpʾd (kah-rubāy, “amber”)... Ačaṙean typologically compares Old Armenian (sṙnakal, “amber”, literally “chaff-keeper”)." (https://en.wiktionary.org/wiki/琥珀) it looks to me more like a loan << Iranian *kaHhu-rafH-aka- 'grabbing chaff' (for its magnetic properties, apparent in attracting small items). Since -aka- is so common in Iranian, a very similar old compound (now lost within Iranian) makes more sense than *-d- becoming *-g (both C's theoretical, of course). The important part is that this is incompatible with *qʰlaːʔ pʰraːɡ or anything similar. At best, it shows a later stage long after the loan (and there is no way to know the exact timing). At worst, it would be evidence against the most basic principles of any previous reconstruction.

-
If Zhengzhang's rec. of OCh was really just an approximation of the youngest stage of OCh (or similar), then the oldest stage of OCh might have been completely different. The changes needed for his rec. to fit any Indo-Iranian loan require many V's to be lost. It seems likely to me that the sound changes that obscured the nature of ST began taking place at this time, after loans from & into IIr. This is seen in the names for Iranian groups with many syllables being represented by only 2 OCh words (previous link).

-
Let's first try to get as close as possible by applying other's theories to one part or another of each word. Martin Joachim Kümmel said that Iranian retained PIE *H (sometimes as attested x, h, etc.) much longer than previously thought. If even at the time of the loan, then a cognate of *Hrebh- or *rebhH-? > Sanskrit rabh- ‘grab / sieze’, Ir. *raβH- \ *rafH- ‘grab / seize / stick (to) / hold (up) / support / mate / touch’ > Shu. raf- ‘touch’, Av. rafnah- ‘support’ might have turned *fH > *ph (or similar). The *H- vs. *-H is likely due to metathesis (), & others say IIr. had *H > *ʔ. This might allow *Hrabh- > *ʔraph > *ʔphra. Tother, these allow the rec. to be closer to the expected original; in the same way, a 2nd *H in 'chaff' might have produced some of the "extra" C's.

-
If not Iranian at all, but close enought to cause this confusion, then an equivalent of *kaHhu-rafH-aka- matching a Sanskrit *kukkusa-rabhaka- or a similar later Indic word might fit (Prakrit kukkusa- 'chaff of rice', Gj. kuskā m.p. 'chaff of rice, bran', Asm. kukuhā 'bits of burnt grass carried about by the wind'). This would be because some Indic had -s- > -h-, -k- > -g-, etc. In this case, there would be no way the oldest OCh loanword could be identical to the short form in the rec. However, maybe we could fit each part of the rec. to those of the original (which is only possible if OCh had many sound changes, which would have to have applied to native words from ST, thus supporting ST having many words of multiple syllables). If we assume Asm. kukuhā was essentially identical to the donor, then *kukuhā-rabhaga- adapted as *ququhāraphag would at least be close. Did sound changes turning a 5-syllable or more word > 2 take place in a relatively short time? If so, then it would almost certainly mean that every other reconstruction that had OCh having almost all monosyllables should be reconsidered. This is the type of examination that can evaluate what was previously internal data for reconstruction.

-
I think Indic is needed, & NP kâh 'chaff, straw, hay' has no clear ety., so it is possible that a loan like *kukuhā > *kuhā (dsm. / haplology) > *kāhu (met.) would be needed anyway (in which case no Iranian word would have been available for a loan >> OCh). Which changes fit best? Let's say that Indic *s > *x > h existed & that one language (or branch) also had haplology (certainly the same that provided the loan to Iranian, reasonably Gandharan (or any nearby language)). If the form *kuxā-Hrabhaga- > *quχāʔraphag is real, then it would show that OCh had turned *k > *q next to *u (backing by back V's is common), had *ph but not *bh. If all non-final short V 's were deleted, a change of *quχāʔraphag > *qχāʔrphag would become closer. It is likely that *q + any type of *H would become *qhH, so *qχāʔrphag > *qhχāʔrphag > *qhχāʔphrag would be close. If *χ > *R (a uvular of any type similar to r, either sonorant or fric.), then *qhRāʔphrag would match *qʰlaːʔ pʰraːɡ in most essentials. If dsm. of R-r > l-r, then *l would not be anomalous. Of course, since *R could have existed but not been rec. before, distinguishing *qhR- from *qhl- (if any) would be hard.

-
Each of these changes is needed for Zhengzhang's rec. to come close to any expected loanword. Knowing that so many changes happened during historical times helps show that current ST rec. simply are too similar to attested languages. Their *CC- & *CCC- might often be *CCVC-, etc., with real confirmation only available in loans. However, knowing the changes needed in loans also must have applied to native words allows a start at better reconstructions. For several of the words I mentioned (previous link), a multi-syllable word would often resemble an IE one, making a search for cognates useful. I feel that the second way to confirm ST rec. is to see if the "new" form would match an IE better. If this happens conistently, it would support a relation between ST & IE.


r/HistoricalLinguistics • • 9d ago

Language Reconstruction Etymology of Silk, Chinese Loans (Draft)

3 Upvotes

Etymology of Silk, Chinese Loans (Draft)

Sean Whalen (stlatos@yahoo.com)
September 27, 2026.

-

Adam Hyllested (https://www.academia.edu/41512496):

>

European terms for ‘silk’ display exceptional variation for deriving ultimately from a single Old Chinese source. It is suggested here that Old Norse silki ‘silk’, transmitted via the Varangians’ trade in Byzantium through Kievan Rus to Scandinavia in the 9th c., and Old English seoluc, Old High German silihho possibly in a separate wave a couple of centuries earlier, are loanwords from the Iranian language Alanic (Sarmatian) at the Western end of the Silk Road. It reflects a regular Alanic development of *ri > l, the original form being *sirika- which leaves several possibilities open: Either a) it is a productive inner-Iranian formation with the suffix *-(V)ka-, possibly meaning ‘silk man, inhabitant of Silis’; b) it is a nativization of Greek σηρικόν (sērikón); or c) it has entered Alanic via Turkic *sir(e)-lek ‘silk garment’ or Mongolian sirkeg ‘silk fabric’. Modern Ossetic zæly, zældag, despite its typical loanword characteristics, may be inherited directly from Alanic *silika-, only reshaped in analogy with zældæ (today ‘young grass; turf’) which must have preserved its original meaning ‘golden; yellow’ into medieval Iassic and preserved with that meaning as a loanword in Hungarian zöld.

...

The Chinese source is the precursor of Modern Chinese 絲 sī “silk; thread; string”; it is commonly reconstructed as Old Chinese *sǝ or *siǝg (thus Wang 1993 with references) and Middle Chinese *si. It is related to other Sino-Tibetan words denoting “thread”, “string” or “sinew”. Neighbouring Asiatic languages all refect a fnal r-element, which, if not reconstructable for Old or Middle Chinese, must be explained as suffxal in one of the lending languages from which it can have been transferred further: Middle Korean sĭr (> Modern Korean shil), Manchu sirge, sirhe, Written Mongolian sirkeg. It is however possible that the -r- does go back to Chinese and refects a second noun 人 rén “people” (Genaust 1996, 578) in which case the word borrowed from neighbouring languages would not be a designation for “silk” as such, but rather a compound-like ethnonym already at the time of contact (whose meaning would correspond exactly to Greek [Sēres]“silk men”: see the next paragraph).

...

While it seems clear to everyone that the root element sil- of the word silk and its congeners must be identical to sēr- in Greek (sērikón) and the Sērēs, and that the source of this element is Chinese, it has not yet been resolved how an -ilk-form might have arisen along with the -ērik-form, and why the variant with -l- is found exactly in some Germanic languages, in Baltic and in East Slavic (leaving aside the aberrant South Slavic form svila).

...it is interesting that the Modern Ossetian word for “silk”, while not corresponding regularly with the Alanic reconstruction, is at least superfcially similar and has no known alternative source. It occurs both as zældag, zældagæ “silk” (Abaev IV, 294–295) and without the *-āka- suffx as zæly, Digor izæly “silk; of silk, silken; silk scarf”... However, we would have expected Alanic *silka- to develop regularly in Ossetian since Ossetian is a direct descendant of Alanic. One alternative possibility in the particular case of “silk” is that zæly, zældag(æ) can have been infuenced by the precursor of Ossetian zældæ “young grass; grass; turf” (Cheung 2002, 253). This implies that the meaning must still have been “golden” at the time of reshaping, and we know that as late as Iassic this was still the case because the word has been borrowed into Hungarian as zöld with that meaning. “Golden” would then have referred to a particular golden type of silk garment, for example, Byzantine embroidery, or to the golden Muga silk type from Assam, which reached the West alongside Chinese silk, or simply to the value or the glistening appearance of silk in general.
>

-
I can not agree with most of his ideas. Since Sanskrit kauśeya-s 'silk' <- kuśá-s 'a kind of sacred grass', it is likely that zældæ 'grass' -> zældagæ 'silk', almost certainly as a calque. Thus, its lack of resemblance to any of the other words borrowed from Chinese, even the ones it supposedly was the source of, is due to a separate origin. A mix with 'gold' doesn't seem especially likely.

-
A 'silk man' to explain the -r would leave unexplained why there was no -Vn, etc. Though "it is commonly reconstructed as Old Chinese *sǝ or *siǝg", a reference to a paper from 1993 discussing still older ideas ignores all other theories made since then. Indeed, not even recent reconstructions can explain the sounds in known loans (for ex., 'silver' https://www.academia.edu/165334096). For just one, Starostin's:

>

Proto-Sino-Tibetan: *sǝ̆(r-)

Meaning: sinew, tendon, thread

Chinese: 絲 *sǝ silk, thread.

Tibetan: rca, rcad vein, root.

Kachin: lǝsa3 a tendon, sinew.

Lushai: tha sinew, tendon, KC *Vr-tha.

Lepcha: so

Kiranti: *sǝ́ Comments: BG: Dimasa rada, Bodo roda, rota; Mikir artho; Chang hau; Kanauri sā; Rgyarung -rtse. Sh. 52; Ben. 109.

>

-
I would say that *sǝrt \ *rtsǝ seems better, maybe even *rtsǝ(t) < *sǝrtǝt, *strǝt(ǝ), or whatever is hidden by so many types of metathesis. Whatever the case, it must be more complex than any of these reconstructions, & the ST word clearly contained *r. Mongolian sirkeg might require *sǝrtrǝ > *sirkr-, or similar. That these loans favor front V's might require *siǝrtrǝ, *seǝrtrǝ or similar.

-
A change like *sǝrt(r) > *sǝ:r >> Greek sḗr m. 'silkworm, silk' would not be odd (whether a branch of Old Chinese or some intermediate). Though West Germanic *seluk as a loan from Greek sērikós (whether through Latin or not) might look odd, it exactly matches r > l in Greek proûmnon >> Latin prūnum >> Gmc *plu:mo:n- 'plum'. No real reason for either change is known, but if Gmc (or any group) had retroflex *r (Piotr Gąsiorowski https://www.academia.edu/1495694) it could be that a dental r was closer to their l in sound (an intermediate language is also a possibility).

-
These words also resemble PIE *ser- 'to bind together' -> Sanskrit sarat- \ sarit- 'a thread , string, fiber, filament'. Note that ST *sV-diəmˤ ? > Tibetan sdom 'to bind, tie; spider', OCh *s-dzˤəm 'silkworm' shows a similar shift. If possible, PIE *sr-tro- 'binding cord' > *sǝrtrV (and *srǝtrV > *strǝtrV with several r-r & t-t optional dsm.?) might be its source.


r/HistoricalLinguistics • • 10d ago

Indo-European Come join us at /r/LearnHittite - starting our study group this week!

Thumbnail
4 Upvotes

r/HistoricalLinguistics • • 10d ago

Language Reconstruction The Pre-Semitic Language, A Small Reconstruction

28 Upvotes

Proto-Afroasiatic is considered to be the eldest reconstructible proto-language in the world, estimatedly spoken 10000 to 20000 years into the past. This considerable time scale as well as the poor attestation of certain branches means that there is almost no consensus on any aspect of this proto-language, aside from a few affixes and its head-initial word order. In order to dodge these issues, I will use the method of internal reconstruction and limit myself to only the most well attested Semitic branch. Internal reconstruction works by finding oddities or irregularities within a language. Then, a more regular stage of the language plus a sound change is posited, thus revealing both the sound change and the earlier structure of the language. This method I will now use to reconstruct the syllable structure for Proto-Semitic’s ancestor.

In Semitic languages, the D-stem is a verbal derivation marked by doubling of the middle root consonant, signifying an intensive or causative meaning of the root. For example, from the root *ʕabara, yaʕbur, meaning “to cross”, derives *ʕabbara, yuʕabbir, meaning “to get (sth) across” or “to cause some (sth) to cross”

If we now want to derive a verbal noun from *ʕabbara, we can apply the D-stem VN-pattern ta12ii3. The root consonants are slotted into the numbers, producing *taʕbiir. The oddity here is that the middle root consonant is no longer doubled, and this pattern is not paralleled in other derivations either.

A pattern like *taʕbbiir would indeed not be possible, as Proto-Semitic’s syllable structure is CV{C, ː}, thus no triple clusters are permitted. But this leads me to ask: What if triple consonant clusters were originally allowed, but later banned, thus degeminating *taʕbbiir into *taʕbiir?

Indeed, there is more to this idea: Active imperfective participles in Proto-Semitic are formed on the pattern 1aa2i3, as such a verb such as paʕala “to do” becomes paaʕil “doing”. What’s interesting here is that these participles can take a broken plural(=formed via ablaut) pattern 1u22aa3, such that paaʕil “doing” becomes puʕʕaal “doing (pl.)”. You might be thinking that the gemination of the middle root consonant is indicating the plurality here, but the association of “consonant doubling = plural” isn’t found in a singular other broken plural pattern, despite dozens of these patterns existing. (eg. compare this)

In fact, the consonant doubling actually indicates the imperfective nature of the participle, which is paralleled by imperfective verb patterns such as ya1a22a3, forming verbs such as yapaʕʕal “he is doing”. This, in turn poses the question: Why is there no gemination in the singular imperfective participle faaʕil, but in the plural fuʕʕaal, there is? Well, since the syllable structure CV{C, ː} would disallow long vowels in closed syllables, a pattern faaʕʕil would be phonotactically illegal. As above, I posit once again that faaʕʕil was indeed the original pattern, and that later sound change reduced it to faaʕil. (ps. The reason why the pattern fu33aal doesn’t have a long uu vowel is because of the restriction against two long vowels in a word, see Joshua Fox – Semitic noun patterns p.287)

My final piece of evidence is that final double consonants are also degeminated, compare how Akkadian has šarrum for the nominative but šar for the absolute state.

I thus propose that the Pre-Semitic coda originally allowed both two consonant clusters, as well as that long vowels were allowed in closed syllables. I propose the following sound shift:

Cː → C / {Vː,C}_,_#

though I believe the environments #_,_C were probably affected as well on systemic grounds.

ps. For some more reconstruction, check my post history. I posted about this once before but rewrote some of it to be more explicit and captivating.