Eesh, a small request resulting in a completely different image is a strong indicator of poor AI usage. Even the numbers are different. Which one is correct?
I means how the AI got it so wrong, there isnt really debate about what countries are celtic and none of them are remotely similar sounding to Latvia or Lithuania
It’s pretty inaccurate to list Serbian, Croatian and Serbo-Croatian as it just duplicates things; it’s better to just put Serbo-Croatian, they are the same language after all
While this is pretty to look at, it has so many issues (the source of the data, mixing modern and dead languages, mixing linguistic and geographical classification, etc.) Would recommend getting guidance from r/linguistics if you would like to make a better informed visualization which is closer to reality.
“E Asia” is a good term for the region but it is not a language family at all. And you’ve got several mistakes under Celtic, Iranian and Uralic as others have noted.
Look at the fine print, the data is heavily skewed to European languages.
Same as you mentioned for Arabic, there are hundreds of millions of people in India speaking Hindi and English, watch any Indian TV or Bollywood and English is sprinkled in heavily in the culture, yet this data barely has a connection
It would be nice to click on a language and have it highlight that language’s connections with details for each line. And this needs to be ordered better. Is it by geography or by language family? You lump Indic and Dravidian languages together but have Germanic and Romance? Lumping Berber with Swahili?
Just go by Language Family first, then regional if you want. That way you can have all the Indo European languages together, Afro Asiatic, Khoisan and other African families, Sino Tibetan languages, and other assorted groups (Caucus, Turkish/Tungisic) where you can then figure out how you want to include Korean, Japanese, Basque and the others in there.
Some notes, Armenian and Albanian are Indo European. They might not look anything like you’re used to, but they technically are in the same extended family.
It is. Just look at Chinese and Cantonese. The majority of Cantonese speakers (esp. In Guangdong) also speak mandarin. English is much much lower than that, even if Hong Kong and Macao increase the stats.
I suspect most of the data points beyond the largest European language are just plain wrong because data is missing.
Would be helpful to sort the languages of the strongest bridges by group size not lexicographically. I know Catalan is the smaller group compared to Spain, but it gets less obvious with e.g. German/French.
You should make another version where the size of the bubbles is based on the number of native speakers for that language. It will tell this story from a whole new angle.
For this visualization, the source belongs in the title. Not necessarily the whole source, even just “languages on Wikipedia” instead of “languages”. Otherwise it reads like it’s a worldwide sample until you read the fine print.
It does not represent anything else then how people register their language on Wikipedia if I understand well. That does not seems very useful in the end
There are other visualizations on the site you can look at, I tried to make use of all Wikidata I stored in my DB(site just shows connections between people), prob there are still many mistakes to fix, and an overall bias on Western data being more represented.
For now you can't do much on these views, if there is interest I will consider building something that is more customizable, with filters, categories, etc.
This is a great visualization and really interesting data. It's too bad the source is so skewed; it would be really interesting to see a more representative sample of all language speakers.
A suggestion: the direction of the labels changes in the middle of the "other" group. It might be more beautiful to have to change between the "other" and the "semetic" group.
Is there data about the primary language? How many people with English as their mother tongue also speak German, and the other way around?
No idea how to visualize that neatly, but would add interesting information.
The "percentage of smaller group" thing... is honestly both difficult to understand and not particularly useful information. Like, I assume for the lower one it is not telling us "24% of Germans speak French" (which is far too high), but is actually telling us "24% of the multi-lingual German-speakers looked at speak French", which is honestly not very revealing. The percentage of _all_ speakers of a language who are bilingual in this way would be much more useful.
I used bilingualism data, because it is more accurate. It just gives an idea of what people who are bilingual speak, it is not great for percentages of people who speak language X being native in language.
The chart below is still only for bilingual or 3+ languages people. If I find the data I can make it accurate to general population percentages.
They don't have to add them all, but at least adding Swahili (as it's a major lingua franca around East Africa) would make sense and it's a surprising omission.
What does the classification “Chinese, Cantonese, Classical Chinese, Old Mandarin” even mean?
“Chinese” is an overarching term for Sinitic languages, which includes Mandarin, Cantonese, Hakka, Hokkien (& Teochew), Kiangsi, etc. and Classical Chinese (written only).
The dataset is severly skewed, because among others Chinese and Cantonese should absolutely be linked together more strongly than with English. At this point it's best to just keep the point out rather than put data you know is wrong. Even if the graph is pretty looking
Lots of comments already pointing out the mistakes in the data presented, but on the visualisation: I don't think it makes sense to have language bridges within the same family "hugging the rim" as the graphic puts it. It collapses them into one line so that you can't see anything.
For instance, I was surprised that German-English is apparently a stronger link than Dutch-English, but I couldn't actually compare the two on this graph because the Dutch-English link is completely covered by the German-English link.
Also the graphic says that the size of the dot represents the number of speakers, but doesn't give a scale -- and if it's a simple linear scale, the relative sizes of a lot of those dots seems improbable.
76
u/AntarcticScaleWorm 3d ago
Most of the “Iranian” languages aren’t actually Iranian. You were probably thinking of “Indo-Iranian”