What is it about?

The Bantu language family is one of the largest in the world, and is spoken from Cameroon in the northwest to the African Great Lakes region in the northeast and to the Republic of South Africa in the south. The Bantu languages are said to form a “language family” because they are all related to each other and descend from a common ancestral language. As the Bantu family is spread across vast distances, but originally stems from a single linguistic community, it must have expanded dramatically over time, diversifying and separating into individual languages in the process. But how did this big linguistic expansion happen? Historical linguists use linguistic evidence to try to find in which exact groupings languages splitted away from each other, forming separate branches on the Bantu language family tree. They also try to discover the exact order of those splits. Together, this information can be summarized in what is called a language-family tree, similarly to how the historical relations between different biological species can be represented by a species tree or a tree of life. But finding out the true Bantu language family tree has been very difficult! Classical historical-linguistic methods, developed in their essence already in the 19th century but still remaining the gold standard for historical-linguistic work, could not deliver the answer so far. The problem seems to be that even after the splits and diversifications, Bantu languages continued to exchange not only words, but also new grammatical features. Exchanges of both words and of grammar is called “language contact”. For finding the true language-family tree, we need to tease apart the signals pointing to the earlier splits vs. the effects of subsequent language contact. Imagine you find an interesting and unusual grammatical feature that is present in a number of Bantu languages, but not in all of them. If this feature could only be inherited, but not spread via language contact, its presence will give us precious information about the tree: the languages having this feature *must* have had a common ancestor. But if the feature could also be “borrowed” between languages, and not just inherited, some languages may have it because they were in contact with Bantu languages already having that feature. In this case, the presence of that feature does not tell us that all languages that have it had a common ancestor which already had that feature. So teasing apart inheritance and contact is necessary for determining the language-family tree. But for the Bantu languages, this appears to be very, very difficult: many grammatical features seem to have spread via language contact within that language family. Is there another way? In the 21th century, historical linguists started to use some tools from computational genetics adapting them to infer language-family trees. Usually these researchers use lexical evidence as data. In particular, they need words descending from the same ancestral word, called “cognates”. These data are then processed computationally, and a set of likely family trees is statistically inferred. This type of computational statistics functions reasonably well for many language families, so when such trees were inferred for the Bantu languages, many scholars in the field thought we have solved the big part of the puzzle. However, there is one little problem... The methods used to infer those Bantu trees were based on the assumption that the Bantu languages did not have appreciable language contact after they first separated from each other. But wait a minute, wasn’t the presence of widespread language contact precisely the problem for classical historical-linguistic methods?.. In this paper, we developed a computational, statistical analysis that does not assume that language contact is negligible. It embraces contact, one can say! For this, we used a different set of tools developed in mathematical genetics and statistics than the ones used in historical linguistics before. We wanted to see what the analysis would tell us about the early history of the Bantu languages if we do not ignore language contact, but include it into the model. The answer of our model was: the signal of the early splits into separate branches has been mostly erased from our lexical data! In other words, we cannot tell anymore how exactly the Bantu language family started to diversify; at least not from the linguistic data that scholars have at the moment. How exactly were we able to find out? We looked at whether the original major splits could be recovered under different conditions. We *simulated* 600 thousand pseudo-histories for the Bantu family: what the linguistic history could have been like under certain circumstances. The properties of those simulations are shown on the graph below; the circles are the individual simulations, and the eight-ray star is the actual Bantu dataset. We can look at which of those simulations were closest to the actual Bantu data, and this allows us to infer some of the properties of the historical process that the Bantu language family went through. Crucially, we can also look at those simulations and try to infer which scenario for the early Bantu history they were generated with: we know this because for each simulation, we ourselves told our model to use this or that possible historical scenario. It turned out that with the levels of lexical change and of language contact inferred for the actual Bantu data, the signal of the original historical scenario is largely erased. With such amounts of change and contact, it is impossible to infer the initial scenario of splits from our simulated data — and by extension, also from the real Bantu data. So on the negative side, we cannot at the moment say what exactly happened in the early period of the Bantu linguistic expansion. But on the positive side, we have also shown why: there has simply been too much lexical innovation and language contact! As a field, we still only possess very limited knowledge of individual Bantu languages. There is also not enough research into the linguistic histories of small subgroups of the Bantu languages: such research is being conducted and has increased substantially over the last half a century, but there is still an enormous amount of work to be done. After more such work is invested, perhaps we will be able to find out how the Bantu linguistic expansion really started! For the moment, however, we can say that one thing is certain: the high levels of language contact within the Bantu language family. This importance of contact has already identified using classical historical-linguistic methods in the 20th century. Our study has now confirmed this on different data and with an entirely different, computational methodology based on the mathematics originally developed for genetic evolution.

Featured Image

Why is it important?

Languages are part of human cultures. Linguistic history is thus part of the overall history of our societies, so learning more about the past of the Bantu languages we will by definition learn something valuable about the past of the Bantu societies. At the same time, to understand the human past holistically, we will also need to combine the linguistic history with other aspects of history: how the human societies in this part of the world adopted and adapted smart innovations (e.g., ironworking); how they interacted with each other; how they resettled across vast landscapes (language spreads are not necessarily demographic spreads, but for the Bantu languages we know from genetic data that there were also demographic spreads); what their precise demographic histories were: who inter-married with whom, with which patterns, and what cultural processes might have accompanied such marital systems… To obtain a holistic picture of the past, we will need all the strands of the evidence about these different phenomena, so studying the linguistic history of the Bantu language family is also important as a contribution to the overall human history of this part of Africa.

Perspectives

Working on this project has been one of the most rewarding experiences of my PhD. I joined the project during the second year of my PhD after Silvia Ghirotto and Andrea Benazzo invited me to help develop and test a coalescent-theoretic framework for reconstructing language prehistory. This project allowed me to explore how computational methods from population genetics can be adapted to investigate linguistic evolution. It also gave me the opportunity to spend time at the University of Tübingen, where I worked closely with Igor Yanovich and other researchers from different disciplines in the Center for Advanced Study “Words, Bones, Genes, Tools”. Coming from a population genetics background, I had never thought that languages could be studied using concepts and models similar to those applied to DNA sequences. It was fascinating to see how ideas such as migration, traditionally used to describe gene flow between human populations, could be adapted to model language contact and better understand the history of language families. Developing a framework that treats language contact as a fundamental component of linguistic evolution showed me the potential of transferring approaches across disciplines and encouraged me to think beyond traditional research boundaries. This experience reinforced my belief that interdisciplinary collaborations are becoming increasingly important for understanding the complexity of human history. By combining approaches from genetics, linguistics, archaeology, and computational modeling, we can integrate different types of evidence and gain a more comprehensive view of how populations and cultures have changed through time. This project showed me that innovative perspectives often emerge when methods and concepts are shared across fields that are traditionally considered separate.

Patrícia Santos
Universita degli Studi di Ferrara

Read the Original

This page is a summary of: Treating language contact as normal: Coalescent-theoretic modeling of the prehistory of the Bantu language family, Proceedings of the National Academy of Sciences, July 2026, Proceedings of the National Academy of Sciences,
DOI: 10.1073/pnas.2608670123.
You can read the full text:

Read

Contributors

The following have contributed to this page