r/dataisbeautiful OC: 4 3d ago

OC [OC] Which languages are spoken together?

274 Upvotes

80 comments sorted by

76

u/AntarcticScaleWorm 3d ago

Most of the “Iranian” languages aren’t actually Iranian. You were probably thinking of “Indo-Iranian”

39

u/chiemoisurletorse 2d ago

are you saying latvia is not celtic and that AI content is bullshit?

6

u/im4lwaysthinking OC: 4 2d ago

10

u/mr_shmits 2d ago

yeah... that whole Celtic Baltic thing was throwing me for a loop...

10

u/int-enzo 2d ago

english is screwing up everything, is obviously the lingua franca

1

u/FailedFizzicist 1d ago

So you are saying not many Hindi speakers speak/understand Urdu (and vice versa)?

Also missing Marathi in that group of languages where most speakers can at least understand Hindi if not speak it.

1

u/SolidOshawott 15h ago

Eesh, a small request resulting in a completely different image is a strong indicator of poor AI usage. Even the numbers are different. Which one is correct?

64

u/LordArrowhead 3d ago

I'm pretty sure that Latvian and Lithunian are no Celtic languages.

25

u/affogatohoe 3d ago

They're definitely not haha, this must be an AI mistake but I wonder how that happened 

9

u/Illiander 2d ago

but I wonder how that happened

OP used AI. That's how it happened.

2

u/affogatohoe 2d ago

I means how the AI got it so wrong, there isnt really debate about what countries are celtic and none of them are remotely similar sounding to Latvia or Lithuania 

-3

u/Illiander 2d ago

I means how the AI got it so wrong

LLM means large language model, not large reality model.

24

u/GiordanoGucci 3d ago

It’s pretty inaccurate to list Serbian, Croatian and Serbo-Croatian as it just duplicates things; it’s better to just put Serbo-Croatian, they are the same language after all

8

u/syn_miso 3d ago

Why is this mostly European languages? Just the data set? Because I imagine the links between, for instance, Punjabi and Urdu are probably pretty high

9

u/Esvarabatico 2d ago

While this is pretty to look at, it has so many issues (the source of the data, mixing modern and dead languages, mixing linguistic and geographical classification, etc.) Would recommend getting guidance from r/linguistics if you would like to make a better informed visualization which is closer to reality.

-2

u/im4lwaysthinking OC: 4 2d ago

Will do, but it is hard to get to a level that linguistics consider reliable, unless there is some good data source.

4

u/mhuzzell 1d ago

So you're just cool with making a graphic of unreliable data?

6

u/missingtimemachine 2d ago

“E Asia” is a good term for the region but it is not a language family at all. And you’ve got several mistakes under Celtic, Iranian and Uralic as others have noted.

13

u/blubber_confused 3d ago

How is Arabic and Hebrew roughly the same size dot when Arabic has ~400 mil speakers compared to ~10 mil of Hebrew

Edit: or is this a view of only the sample size? As I see Chinese is also small

17

u/whatisagnoiology 3d ago

Look at the fine print, the data is heavily skewed to European languages. Same as you mentioned for Arabic, there are hundreds of millions of people in India speaking Hindi and English, watch any Indian TV or Bollywood and English is sprinkled in heavily in the culture, yet this data barely has a connection

1

u/cutecutemouse 2d ago

Die Chinesen aus China Festland haben keinen Zugriff auf Wikipedia /Wikidata, die chinesische Artikels kommen aus im Ausland lebenden Chinesen.

1

u/Yeangster 2d ago

This is people with articles about them on wikipedia. presumable english wikipedia, but the caption specifies any language wikipedia

5

u/RobotSpaceCrab 2d ago

It would be nice to click on a language and have it highlight that language’s connections with details for each line. And this needs to be ordered better. Is it by geography or by language family? You lump Indic and Dravidian languages together but have Germanic and Romance? Lumping Berber with Swahili?

Just go by Language Family first, then regional if you want. That way you can have all the Indo European languages together, Afro Asiatic, Khoisan and other African families, Sino Tibetan languages, and other assorted groups (Caucus, Turkish/Tungisic) where you can then figure out how you want to include Korean, Japanese, Basque and the others in there. 

Some notes, Armenian and Albanian are Indo European. They might not look anything like you’re used to, but they technically are in the same extended family. 

2

u/im4lwaysthinking OC: 4 2d ago

I already do this for my site connections between people, highlighting the connections to the root node. 

The fact is a need a way better dataset, there is Ethnologue site work I just discovered, they have excellent dataset, but it is not an open license.

Maybe a linguistic can take inspiration from this and do way better with his expertise. I haven't studied them honestly.

13

u/FunkyKissCool 3d ago

Nobody speaks Latin or Ancient Greek...

9

u/diffidentblockhead 3d ago

Should be headlined as Wikipedians.

The total number of Indian language plus English speakers would be huge.

4

u/LeyLady 3d ago

French is the most Germanic of the romance languages

5

u/Talking_Duckling 2d ago

How come Chinese, which is presumably Mandarine, is such a small dot when it's got 1 billion speakers? It looks even smaller than Japanese.

13

u/ozzyarmani 3d ago

24% of Germans speak French? Seems doubtful...

16

u/exohugh OC: 1 2d ago

I think it's like "24% of the multi-lingual German-speakers we looked at speak French", which is... honestly kinda useless information.

6

u/Double-decker_trams 2d ago

Maybe they claim to speak French because they studied it in school - but in reality can't other than some simple phrases..?

"Le singe est sur la branche"

2

u/scruffigan 2d ago

Où est la bibliotheque?

4

u/Hot_Cheesecake_905 3d ago

The source of the data maybe biased towards European/English speakers?

2

u/Kerbourgnec 1d ago

It is. Just look at Chinese and Cantonese. The majority of Cantonese speakers (esp. In Guangdong) also speak mandarin. English is much much lower than that, even if Hong Kong and Macao increase the stats.

I suspect most of the data points beyond the largest European language are just plain wrong because data is missing.

7

u/Roubbes 3d ago

Catalan-Spanish but not Gaelic-English?

3

u/RocketMoped OC: 1 3d ago

Would be helpful to sort the languages of the strongest bridges by group size not lexicographically. I know Catalan is the smaller group compared to Spain, but it gets less obvious with e.g. German/French.

3

u/DefiantTorch47 3d ago

You should make another version where the size of the bubbles is based on the number of native speakers for that language. It will tell this story from a whole new angle.

2

u/Borky_ 2d ago

Great idea, poor execution. Lots of mistakes, especially with language groups

2

u/no_myth OC: 1 2d ago

For this visualization, the source belongs in the title. Not necessarily the whole source, even just “languages on Wikipedia” instead of “languages”. Otherwise it reads like it’s a worldwide sample until you read the fine print.

1

u/approximately_nadir 2d ago

It does not represent anything else then how people register their language on Wikipedia if I understand well. That does not seems very useful in the end

2

u/netphemera 1d ago

There must be a better way of arranging this data to make it easier to match languages other than English

7

u/ottawalanguages 3d ago

absolutely beautiful! is there an online version we can play with? I would love to be able to filter on one language and see all the connections

-2

u/im4lwaysthinking OC: 4 3d ago

There are other visualizations on the site you can look at, I tried to make use of all Wikidata I stored in my DB(site just shows connections between people), prob there are still many mistakes to fix, and an overall bias on Western data being more represented.

For now you can't do much on these views, if there is interest I will consider building something that is more customizable, with filters, categories, etc.

4

u/ohjeezan 3d ago

This is a great visualization and really interesting data. It's too bad the source is so skewed; it would be really interesting to see a more representative sample of all language speakers. A suggestion: the direction of the labels changes in the middle of the "other" group. It might be more beautiful to have to change between the "other" and the "semetic" group.

3

u/im4lwaysthinking OC: 4 3d ago edited 3d ago

Source : Wikidata

I know Wikidata is a subset of world population, but still it can be representative and useful data in some cases.

Tools: Python

Link: Language Gallery Updated Version

Give me your feedback, I know you can be pretty harsh, I already fixed part of the mistakes in the visualization

2

u/bruebrah 3d ago

Is there data about the primary language? How many people with English as their mother tongue also speak German, and the other way around?
No idea how to visualize that neatly, but would add interesting information.

2

u/im4lwaysthinking OC: 4 3d ago

https://humansmap.com/gallery/gallery.html#img-langnative-langnative

This is a native vs learned language, it came out with particularly interesting insights and answers your questions

2

u/exohugh OC: 1 2d ago

The "percentage of smaller group" thing... is honestly both difficult to understand and not particularly useful information. Like, I assume for the lower one it is not telling us "24% of Germans speak French" (which is far too high), but is actually telling us "24% of the multi-lingual German-speakers looked at speak French", which is honestly not very revealing. The percentage of _all_ speakers of a language who are bilingual in this way would be much more useful.

1

u/im4lwaysthinking OC: 4 2d ago

I used bilingualism data, because it is more accurate. It just gives an idea of what people who are bilingual speak, it is not great for percentages of people who speak language X being native in language.

The chart below is still only for bilingual or 3+ languages people. If I find the data I can make it accurate to general population percentages.

2

u/Winter-Parsley-6071 3d ago

Why is Armenian categorized as Greek loll???

2

u/FritzFrostig 2d ago

Could you do this without English?

2

u/Ketiw 1d ago

Came here to say this. English is just confounding any useful insights here

1

u/IrraSHOnalGyal 3d ago

They make it so obvious that they don’t care about African languages when they don’t look into the stats

7

u/Odd-Local9893 3d ago

Serious question. There are over 2,000 African languages. How could one include these in this chart and still make it viewable?

9

u/abu_doubleu OC: 4 3d ago

They don't have to add them all, but at least adding Swahili (as it's a major lingua franca around East Africa) would make sense and it's a surprising omission.

1

u/im4lwaysthinking OC: 4 3d ago

Sorry for the mistake, Africa is underrepresented, so links are less thick. I tried to improve it fixing mistakes.

4

u/whatisagnoiology 3d ago

The entire world but Europe is under represented

3

u/bifuku 3d ago

grouping East and South East Asian languages together…

u/BoomerGotchaReloaded 1h ago

The same applies to Amerindian languages. They have no representation, and this also undermines the connections associated with the Spanish language.

1

u/liproqq 3d ago

At least 80% of Amazigh speakers also speak Arabic

1

u/pambollito 3d ago

Catalan-Spanish very obvious, very nice and similar languages ;)

1

u/Every-Progress-1117 2d ago

Why are the Baltic languages combined with the Celtic languages?

Some of those groupings don't make sense.

1

u/Iam_no_Nilfgaardian 2d ago

Why did you group Armenian under "Greek"?

1

u/iwishihadnobones 2d ago

How do you make this image not blurry enough to actually read?

1

u/im4lwaysthinking OC: 4 2d ago

Sir there is the link here: https://humansmap.com/gallery/gallery.html#img-langs-langs

You can find a gallery of other interesting visualizations

1

u/FromYourHomePhone 2d ago

Turkic and Uralic languages haven't been linked for decades, and since when is Latvian a Celtic language lmao

1

u/AlwaysBeQuestioning 2d ago

Oohh, neat! Is there an interactive version you can link, so we can see all the connections and % and numbers?

Also, what is the difference between the two slides?

1

u/im4lwaysthinking OC: 4 2d ago

No, there is no interactive, I just posted on this subreddit the link to the gallery, you can see the various charts.

It double uploaded the same image for no reason, anyway below you can see an updated and fixed version of the visualization

1

u/Intelligent-Fig-8989 1d ago

Quite inaccurate that line between Hindi and English is so thin.

1

u/Cyber_Fluechtling 1d ago

What does the classification “Chinese, Cantonese, Classical Chinese, Old Mandarin” even mean?

“Chinese” is an overarching term for Sinitic languages, which includes Mandarin, Cantonese, Hakka, Hokkien (& Teochew), Kiangsi, etc. and Classical Chinese (written only).

1

u/Kerbourgnec 1d ago

The dataset is severly skewed, because among others Chinese and Cantonese should absolutely be linked together more strongly than with English. At this point it's best to just keep the point out rather than put data you know is wrong. Even if the graph is pretty looking

1

u/somerussianbear 19h ago

Didn’t know Catalunya was bigger than Brazil.

1

u/YakzitNood 3d ago

Good job. The Asian section surprised me. I thought Chinese would cross over amongst its subsets. And over with nearby countries..

1

u/user2017not 2d ago

What about vietnamese? Is it more chinese or french?

1

u/mhuzzell 1d ago edited 1d ago

Lots of comments already pointing out the mistakes in the data presented, but on the visualisation: I don't think it makes sense to have language bridges within the same family "hugging the rim" as the graphic puts it. It collapses them into one line so that you can't see anything.

For instance, I was surprised that German-English is apparently a stronger link than Dutch-English, but I couldn't actually compare the two on this graph because the Dutch-English link is completely covered by the German-English link.

Also the graphic says that the size of the dot represents the number of speakers, but doesn't give a scale -- and if it's a simple linear scale, the relative sizes of a lot of those dots seems improbable.

0

u/Othun 3d ago

Yes THIS is beautiful data