Lexicography in Four Dimensions:
Database Structure for a Matrix of Human Expression Across Time and Space
Martin Benjamin, Jérôme Bâton, Greg McKeen
This paper documents the lexicographic structure that Kamusi has implemented in a Neo4j graph database to accommodate complex linguistic data for all known human languages. (For a complete primer on Neo4j, see Bâton and Van Bruggen 2017). We call our approach “molecular lexicography”, seeing terms as the basic particles that interconnect in a variety of ways to form language. We have designed the Kamusi database, Kam4D, to account for the many elements that form the way we express our thoughts, in a way intended both for human reference and for use as a data source for advanced language technologies. The result is a matrix for linguistic data - words that have been digitized in a way that can be used within computer processes - that can document what people say from one geographical location to another, with a temporal dimension that can show how today’s words emerged historically. Collecting all that data is a separate task that we discuss elsewhere (Benjamin and Radetzky 2014; Benjamin 2015b; 2015a; 2014a); this paper lays out the Kam4D architecture in which unlimited multilingual data can comfortably reside.
Our data structure has evolved through years of experimentation with diverse multilingual data. The structure combines established aspects of lexicography with a number of features that are new to the field. Our innovations arose to address a challenge that lexicography has not previously encountered (Colowick 2008): interlinking linguistic data coherently across more than two languages. Lexicography began as a monolingual exercise, where it was possible to adorn each word in a language with as much information as seemed relevant, without too much worry about how those words connected beyond their alphabetical order. Bilingual lexicography traditionally included less information about the words in a language, in favor of showing how surface forms of those words connected to the words of another language (Adamska-Sałaciak 2013; Piotrowski 1994), on the premise that a monolingual dictionary might be available on either side as a supplement. Bilingual lexicography thus rests on the implicit assumption that the reader has recourse to enough knowledge of one or the other language to tease out the meanings of words, and to differentiate senses when the same combination of letters can produce more than one meaning (Parent 2009). Further, bilingual lexicography pretends that languages have neat equivalents from one to the next, what (Zgusta 1971) calls “anisomorphism”.
These assumptions fall apart in the effort to connect three or more languages, in ways that we elaborate below. As we set out to create a multilingual model that could maintain meaning across more than 7000 languages (more on that estimate at http://kamu.si/7000-languages), we were faced with a stark truth: if it were easy, it would have already been done. In fact, the interplay among languages introduces so many nefarious complexities that no true multilingual dictionary has ever been attempted, beyond lining up words in some languages and praying for the best.
This article is written with three different audiences in mind. First are academic linguists and lexicographers, who know a lot about how language works but may not have given deep consideration to how to structure linguistic information as actionable data. Second are technologists involved with data science and natural language processing, who know a lot about how computers work, but may not have deep knowledge of lexicographical issues. Third are citizen linguists: people with an interest in one or more languages who do not have any specialized training, but would like to contribute their knowledge to the growth of data in Kamusi, and thus want to understand how and why the system was designed as it is. The writing is therefore aimed at a common middle ground, using language that we hope will satisfy specialists with a range of different specialities, while at the same time being approachable by the general public.
A word about words: we do not necessarily adhere to precedent regarding the terms that describe our data. We say “party terms” where the literature speaks of “multiword expressions”, for example, and “costumes” where tradition goes with “inflections”. We do this so that non-specialists can better understand the underlying concepts, as well as people working on the technical side who might not have a linguistics background, and readers who might not have English at an advanced level. We also introduce some terms that are new to lexicography, such as “smurfs” and “ducks”. We will define any terms that are new or stray outside convention. In addition, we tend to talk about “terms” rather than “words” when we are talking about combinations of letters or signs that form a particular signification. A “word” in most written languages is a combination of letters without spaces that may or may not be independently meaningful (Chinese that does not use spaces and Vietnamese that puts spaces between every syllable are prominent exceptions). A “term” can be a single word, such as “state”, or a combination of words such as “United States of America” where the individual words do not signify anything on their own. Also, knowing that 50% of readers will quibble no matter which side we fall on, we use “data” as a singular and beg forgiveness from the other half.
Just as molecules are composed of even smaller bits (atoms, which are themselves made of protons, neutrons, and electrons, which are further built of even smaller stuff), the terms of a language bond together several elements: shape, sound, meaning, time, place, and relationships. As our star graphic hints, each of those elements are composed of smaller parts. The Kam4D data structure is intended to account for as much of that data as we are able to collect for any term. The various elements of a term play together in subtle ways - for example, a term may have a single shape, like “really”, but a change in sound may induce a change in meaning (what the kids call “prosody”). Really? Really. The following sections explore the Kam4D chemical composition, with a description of the sub-components for each element. Because all of the elements interact, the ordering of the sections is somewhat arbitrary, with frequent cross-referencing back and forth. Buckle your seatbelt, the flight through our graph might get bumpy.
Before take-off, the in-flight safety card highlights two features that power the craft: smurfs and ducks.
It is crucial to note that machine translation services such as Google Translate generally guess at vocabulary equivalencies starting from spelling correspondences (lemurs, stay tuned) inferred from parallel corpus data (basically, big bunches of human-translated texts[1]), whereas the Kamusi method deploys human knowledge[2] as the basis for a term to enter a duck and begin its association across languages. A duck is a number that is associated with a rather ethereal concept - duck-ness - that exists outside of any sole linguistic embodiment, while a smurf is a number that pinpoints a specific manifestation of that concept within a language.
In the beginning is the word. In alphabetic languages, a word is usually represented with a sequence of one or more letters (although symbols are possible, such as & and %, or the ! that indicates a click sound in Khoisan languages). Some words have a fixed shape: in English, you can have “information”, but you cannot have plural “informations”. Many words can change shape, though, depending on context. In English, most nouns have both singular and plural forms, such as «star/ stars», and verbs have as many as five forms,[3] such as «see/ sees/ saw/ seeing/ seen». These shapes do not change the underlying meaning of a word - whether you “see” a movie or you “saw” a movie, you visit the same dictionary entry to read the crystalizing definition.
Traditionally, dictionaries are organized around what is variously referred to as the lemmatic, canonical, or dictionary form. This is the form where users are expected to open up their dictionaries and find their answer, such as “see” in English, or the infinitive “divisar” in Portuguese. We refer to this form as a lemur, short for “lemmatic unit reference”, because it sounds like “lemma” but encompasses a wider range of items than included in some scholarly definitions of that term. Also, we can represent lemurs with a cute graphic that will be more memorable in our explanations to the public than a specialist term like “lemma”.
A lemur can be a single word, or something larger, or something smaller. A party term like “out of order” is a lemur, because that string of letters and spaces is the container for one or more meanings; in this case, the lemur is the reference shape for smurfs that mean “broken”, “misarranged”, or “improper”. Conversely, “ki” in Swahili is also a lemur, although it only ever appears as a fragment within a longer word, such as “Kiswahili” (ki = language), “ukiwaona” (if you see them; ki = if), or “walikiona” (they saw it; ki = it).
Lemurs are not useful as vessels for meaning. For meaning, we have smurfs; one polysemous lemur, such as “see”, can be wed to many smurfs, like a polygamous man can be wed to many wives. Like the wives in a multiple marriage, those smurfs usually do their own thing when they interact with the world, but they all cohabit under their lemur’s umbrella for certain tasks. For example, the ancestor-spawn relationship between “operate” and “operator” is probably lemur-to-lemur, not smurf-to-smurf, because we have no way of saying which specific sense of the former gave birth to which specific sense of the latter. For wardrobes and costumes, all the action occurs relative to the lemur; since shapes have nothing to do with meaning, it makes no sense to repeat inflectional information for every smurf.
saw (noun) wardrobe 1 has costumes = saw, saws | saw (verb) wardrobe 2 has costumes = saw, saws, sawed, sawn, sawing | see (verb) costumes = see, sees, saw, seen, seeing |
From a data perspective, we need to track whether “saw” pertains to seeing a movie versus seeing a romantic partner versus a tool for cutting, so that we can perform processing tasks such as finding the correct term for translation. Linguists call these different shapes “inflections”, but we call them “costumes” as a friendly way to remind users that a single word, like the little girl in the picture, can dress up in a variety of clothing without a core change in essence. Costumes are stored in a “wardrobe”, such that all of the appropriate senses of the verb “see” can select their outfit from a wardrobe that consists of «see/ sees/ saw/ seeing/ seen». When a software process encounters “saw”, it can interrogate the database for all the wardrobes that contain that costume (in this case, the costumes «saw/ saws/ sawed/ sawing/ sawn» and «saw/ saws» would also be options), and thereby locate all the senses that access those wardrobes.
Most languages have costumes (some do not, making declarations like “she past see” instead of conjugations like “she saw”), but there is no consistency among languages. For instance, English has three adjective forms (e.g., bright/ brighter/ brightest), whereas Swahili has 16. English has five verb forms, French has 96, and the Bantu language Kinyarwanda has north of 900,000,000 - far too immense for a database. Kam4D allows each part of speech for each language to specify the number of hangers it needs for its costumes, and for labels like “past participle” to be assigned to those hangers. The 96 hangers for the French verb approach the outer limit of what is practical to store in a database. Beyond that size, the rules often become regular enough that a linguist and a computer scientist could together instruct software to generate or parse inflections with high accuracy, following a consistent algorithm that fluent speakers keep in their heads. We have done this for the approximately 18,576,000 potential forms[4] for most Swahili verbs, another of the roughly 400 agglutinative Bantu languages that make up over 5% of the world’s total. We posit that people and computers can embed a set of rules (a few hundred in the Swahili case) in their processors that look funky but are internally consistent, or they can embed a set of irregular words (such as a few hundred verbs in English), but no person and no dataset needs to store 18 million variations. Kam4D does not attempt to catalog the infinite potential agglutinations in a language like German, which can stick nouns together endlessly (like gluten glues bread molecules together); rather Kamusi software can spin through an ornate cape like Rindfleischetikettierungsüberwachungsaufgabenübertragungsgesetz (beef labeling supervision duties delegation law, in this case, which is a real thing) and find its disguised parts in the database. We have not begun to engage polysynthetic indigenous languages of the Americas that have similar complexities, but when we do, we expect that a manageable set of rules (mental algorithms) will govern a large portion of costumes we observe or can reasonably predict.
Part of speech (PoS) can affect shape, such as adding an “s” to make a plural for nouns in many European languages. Not all languages have the same parts of speech (we have thus far inscribed 60 PoS designations that linguists have named from 34 languages), PoS do not act the same way in all languages, and PoS obviously have different names in different languages. Kam4D has a general boolean field in which a duck can be labeled as describing a pos. When a language is configured in Kamusi, the relevant PoS duck can be assigned, such that a language with prepositions will offer preposition-ness as one of its options, with the term for “preposition” displayed in the user’s language, be it “réamhfhocal” in Irish or “nakalazi” in Luganda.
PoS can also have grammatical attributes, for example a verb can be transitive or intransitive, which may or may not affect the shape. Kam4D collects the universe of grammatical attributes as ducks, and makes grammar_attribute available to be assigned to their parent PoS when relevant for a language. This knowledge will be important for building language models that can recognize and preserve relevant features during translation.
Shape is complicated by the possibility of alternate spellings and alternate alphabets. This is really just a matter of packaging, with everything else about the term remaining consistent - not much more consequential than putting identical chocolates inside different color wrappers, except that Kam4D must be able to lead people or software to the contents no matter their orthographic predilections. In English, for example, “color” and “colour” are the same word, with different orthography for different places. Therefore, Kam4D has a field for alternate_spelling, and includes the possibility of marking alternates according to criteria such as geography (either a named location, or geo-coordinates) or dialect. The database also handles digraphia, where words are written in two or more scripts. In Japanese, for example, a word can be represented in as many as four scripts, such as the word for “goldfish” that is variously 金魚 (kanji), きんぎょ(hiragana), キンギョ (katakana), or kingyo (romanji). The number of scripts for a language is specified and labels are assigned, the way the number and names for hangers are treated for wardrobes.
Another complexity of shape occurs in “party terms”. Professional linguists call these multiword expressions or MWEs, but “party terms” is a more user-friendly way of communicating their spirit. Party terms are collections of words that form meaningful phrases when they dance together but cannot be understood by looking at their parts, unlike “wallflowers” that are single words that keep shyly to themselves.[5] There is no need to lexicalize “chicken sandwich” because it is clearly a sandwich that features chicken, two wallflowers with their own stories to tell who will only dance when asked. No person or machine, though, could process a party term like “shit sandwich” or “South Sandwich Islands” by referencing the component words individually. Identifying and deploying party terms is a major challenge for both language learners and computational linguists (Sag et al. 2001; Rayson et al. 2010).
While a party term like “United States of America” has a fixed form, many others can change, in two ways exemplified by the term “drive up the wall”. First, the term can be inflected: “drove up the wall”, “driving up the wall”, etc. Second, the term can be separated: “drives me up the wall”, “drove everybody at the office up the wall”, etc., with intervening strings that could be one word or forty. Kam4D stores party terms as both lemurs and lexicalized smurfs, such that “drive up the wall” can be found as an operable entity. The graph takes notice that the term contains the words «drive¿ up¿ the¿ wall», and allows wardrobes to be assigned to those words; in this case, the wardrobe for “drive” would be enabled, but the wardrobe that includes the plural “walls” would not. Further, party terms can be marked for separability, such as “drive || up the wall”, so that separated terms can be reassembled in Kamusi’s SlowBrew source-side predisambiguation translation software - for example, when SlowBrew encounters “drive” it can paddle any distance downstream in the sentence to see if it encounters the additional parts of any party term that starts with that word in Kam4D, and if it sees “up the wall” floating along, it can act on the unified entity.
Because a party term is handled as a smurf, it is easily aligned in a duck with other smurfs across languages. This approach overcomes debilitating limitations of current “neural machine translation” (NMT) platforms such as Google Translate. NMT generally fails to discern party terms at all, especially idioms and colloquial expressions. This is because party terms are excessively difficult to discover automatically by computational analysis of available monolingual corpuses (Lauder 2010), and orders of magnitude more pernicious to find translations within bilingual parallel corpuses. For the rare party terms NMT can identify in their base form, the technique is destroyed by inflections and separation. In Kam4D, by contrast, “drive up the wall” pairs with “rendre fou” (make crazy) and “rendre chèvre” (make a goat) in French, “tirar do sério” (remove the seriousness) in Portuguese, and “~을 미치게 하다” (make insane) in Korean, none of which has anything to do with driving or walls.
Consider the party term “reckless endangerment”. This term, just two unseparated words, does appear in some officially-translated parallel documents, so Google and DeepL pick up French and Spanish translations that they propose in some sentences but not others, and are valid in some contexts. For legal purposes, though, a translation that is flat-out wrong will be rendered at least some of the time in every jurisdiction where courts use French or Spanish. The term does not appear in parallel data between English and Swahili, so Google makes a guess that back-translates as “to endanger with danger”, making it wrong for that language 100% of the time. In fact, arbitrary guesses are the modus operandi for this party term and most others across the 108 languages Google treats, and not even guesswork is available within NMT for the other 98.5% of languages. Kam4D, on the other hand, stores equivalents that are validated by knowledgeable humans, with unlimited languages providing confirmed equivalents for a concept within a duck. Moreover, geotagging makes possible the relegation of specific translations to identified locations, such that the crime of “reckless endangerment” can be rendered in Spanish with the right smurf for the jurisdictional legal equivalent in Florida, Venezuela, or the Dominican Republic. Current Kamusi data is still far from the standard where you will get the right translation in the right location for every party term, but the database is designed to support that task as an end goal.
Kam4D has a boolean field for whether an item is a party term. Often, this can be discerned automatically by the presence of spaces. However, in languages like German and Chinese where party terms may exist without spaces, or Vietnamese where spaces separate the syllables of individual words, participants will need to manually mark the divisions between words. The database has a field for party_members for languages where the component words are not obvious. A party term that is marked True with the separable boolean acquires the party_split field that shows the location where the separation is possible or mandatory.
The flip side of party terms are abbreviations and acronyms. “USA” is a shortened shape of “United States of America”, but its meaning, history, and relations are the same, so it belongs in the same smurf. Kam4D therefore has an abbreviation field, which can be multiple (e.g. “US” as well as “USA”). Contractions, by contrast, are handled as a part of speech, with each part of the contraction shown in the ancestor field; “didn’t” is a smurf, its PoS is “contraction”, and “did” and “not” are its ancestors.
Software processes English possessive ’s forms externally to the database. That is, costumes such as “the girl’s dress” or “Mary’s dress” could pop up anywhere, so it is easier to code a rule than to store “girl’s” and “Mary’s”, etc, in hard form. Similar situations in other languages will be treated similarly, when we are aware of them.
“Don’t worry. Be happy.” This song worried us a lot at the onset of the Kamusi Project, giving us problems in both English and Swahili. In Swahili, the verb for “be happy” is “-furahi”, which should be displayed with the hyphen, but sorted and searched without. (The hyphen indicates that one or more grammatical affixes can occupy that space.) In English, “be happy” works as a phrasal verb, but there is no verb “to happy”.
To solve this problem, we originally used a field called “SortBy”, to indicate the place you would look for the term in a standard print dictionary. The SortBy for “-furahi” was “furahi” and the SortBy for “be happy” was happy. We later changed this field to “headword” to sound more lexicographyish. We came to realize, however, that designating a single headword had reached its sell-by date in the era of electronic dictionaries.
Giving useful results for party terms (aka multiword expressions) can be tricky. For a test case, consider “home security system” as a thing that might, with 15,300,000 Google hits, merit its own entry. We could take the kit-and-caboodle approach, where searching any of the component terms will yield the item as a result, but that would flood the sorting task for a word like “system” - solar system, Dewey decimal system, system operator, etc… Imagine the mess if all the terms that need lexicalization in English as explanatory phrasal verbs in relation to Bantu languages had to be manually sorted under “be”: “be educated”, “be green”, “be early” and thousands more. For our test case, we probably want to bring users to the item by default when they search for the entire party term, or by intent when they search for the midword “security”, but keep them out of the soup for “home” and “system” unless they opt to view messy catch-all results.
A term like “tight end” would be filed under the full default string, or via the tailword “end” and grouped with other types of “end” in American football. “United States of America” should be found from the default and from “United States” and “America” (“the States” is a separate smurf with a shared meaning), as well as from its abbreviations “US”, “U.S.”, “USA”, and “U.S.A.”. Clearly, “US” is not a headword.
We finally settled on “find agent” as the term to describe what we are looking for. FindAgent allows users to list the components of a party term that should return results, in addition to abbreviations and costumes. For example, English needs an entry for “be said” as an explanation for Swahili “-semwa”, which itself is a costume of “-sema”. So, the English FindAgents are “said”, and “say”, and the Swahili FindAgents are “semwa” and “sema”, in addition to the defaults. FindAgent is more of a computational convenience than a lexicographic edict - we can consider it an e-dict edict that users should find the information they are looking for in the places they are likely to look.
Costume presents some unresolved challenges for FindAgent. Kam4D easily brings the known costume “tight ends” into the set of search terms that will, along with “tight end” and “end”, bring back “tight end” and link to its translations. However, “ends” is neither a costume nor a FindAgent. Listing FindAgents for costumes would take us down a treacherous path, demanding excessive labor from volunteers and getting tangled up with other search terms. For example, “ran up” is a costume for “run up”, and can swiftly lead a user to the groupings for the several senses of that phrasal verb. Isolating “ran” as a FindAgent, though, would place “run up” in the thick of hundreds of other ran-dom terms, such as “run down”, “run ragged”, and “run of the mill” - which would create nightmares for grouping and ranking. We will continue to ponder an efficient way to bring elements of multiword costumes into FindAgent, but for the moment they will only appear when users explicitly conduct unsorted kit-and-caboodle searches.
Not all languages use letters to represent their words.
Emoji_set is different from ducks, because one emoji can map to many different concepts in different languages. As ducks, we can say that English “chick” maps to French “poussin”, and English “egg” maps to French “œuf”. The 🐣 emoji, though, maps to all four, though eggs and chicks are clearly different concepts. Therefore, emoji_set shows items that are joined among languages solely by their reference emoji, but these are not assumed to have a close semantic relationship. Many emojis have been linked to the relevant Kamusi duck through a shared English meaning. 🐣, for example, has been separately joined to “chick” and “egg”, and the relationship propagates separately through those ducks to French “poussin” and “œuf”, and to all other languages that have terms in those ducks. Terms might be linked across languages through emojis but not through ducks, or through ducks but not the data cultivated from CLDR,[7] or through both. It is also possible within Kam4D to cement a direct link between a non-English term and an emoji, such as “poussin” and 🐣, but this is activity for future gameplay.
Most languages use speech more than writing, but conventional dictionaries do a horrible job of representing the sounds that words make. Consider the word “thorough”:
The one point of agreement among a few of these sources is that one pronunciation in the UK is “/ˈθʌɹə/”, which is composed of four characters that 99.999% of dictionary readers, probably including yourself, cannot read. Huh?
Attempts to represent speech using Latin letters are doomed because there is no standard association between letters and sounds. For example, the “th” in “thin” is a different sound, with the tongue behind the teeth, than the “th” in “then” where the tongue is under the teeth. Different variants of English have different pronunciations of many words - for example, people in the UK pronounce the “th” in “clothes” while Americans do not. And different languages that use the Latin character set have different ways of rendering the same sounds, such as the “th” in French “théorie” that is pronounced with a hard “t” tongue tap to the hard palate. The 26 letters we use in English, some diacritic marks such as “ț” in Romanian, and maybe a few supplemental letters such as “þ” in Icelandic, are just not enough to produce a convincing pronunciation system for all the phonemes that languages expel from the human mouth. Amharic solves the problem via the Ge’ez script, with 231 characters that represent all possible consonant-vowel combinations in the language (as seen in the ancient manuscript in the photo), but we do not foresee the rest of the world adopting Ge’ez at any point this millennium.
The International Phonetic Alphabet (IPA) is a linguists’ solution to the problem of representing sound, with 107 letters to represent consonant and vowel sounds, 31 diacritics to modify those sounds in some subtle way, and 19 additional signs for features such as length and tone. If your mouth can produce a sound, IPA can represent it, including phonemes such as the clicks found in Khoisan languages. There’s just one catch: only linguists with special training can actually read IPA, which leaves it as scrawl on a page for most dictionary readers.
IPA does have an important role to play, so Kam4D handles IPA spelling within the architecture. IPA is available for all costumes of a term. Any smurf or costume can have multiple IPA spellings, each of which can be given a geographical label, or geotagged with map coordinates. The intent is not so much to print IPA to a page, but:
You like potato and I like potahto
You like tomato and I like tomahto
Potato, potahto, tomato, tomahto!
Let's call the whole thing off!
The systems used to represent sound on a page run smack against the fact that people speaking the same large language do not pronounce words consistently across its geographic range. English speakers can often identify whether a speaker on the radio is from Texas or South Africa or Ireland, or is African-American or African, based on pronunciation alone. Francophones might tell you the canton in Switzerland where a speaker learned French, and a coastal Tanzanian can point to the inland region where a visitor mastered Swahili. The one or two pronunciation options that dictionaries usually present are polite fictions - economies of space that transmit a poverty of data.
Many online dictionaries address the pronunciation conundrum with audio clips. Audio is a good solution. The best current implementation is Forvo.com, which geotags the locations of the contributing speakers, but those recordings cannot be used as data for other applications. Kam4D enables speakers to geotag the location with which they identify, which can be used to learn or reproduce the accents of specific areas - the English of Manchester, England, in comparison to Manchester, New Hampshire, say.
Pronunciations are attached to costumes and included in wardrobes, so a pronunciation produced under the auspices of “wound” as the past tense for the verb “wind” meaning “wrap around a coil” will be accessible to “wound” as the past tense for “move along a twisted course”, but can be segregated from the differently-pronounced “wound”, “an injury to living tissue”. This is a significant advance over the way other dictionaries handle pronunciations, in two ways:
Over time, Kam4D can result in localized sound clouds, which can solve problems such as the inability of voice recognition software to decipher regional accents. We have not yet addressed how to chronical differences in pronunciation that sometimes (not always!) adhere to characteristics other than geography, such as racial or ethnic background or sexual orientation, and might be heard within meters of each other on the same checkout line or subway car - collecting detailed socio-demographic information from our contributors is technically feasible, but is a big ask that touches on sensitive privacy concerns.
Tone is an essential feature of perhaps 70% of the world’s languages, which lexicography typically handles even worse than pronunciation. Languages often use a variety of rising and falling intonations to distinguish different words, or grammatical inflections for a single word. When it comes to showing those tones in writing, though, we run into a problem created by the printing press in the 15th century, when text was flattened to the simplest possible set of uniform-size characters printed in a uniform color. When printers added pronunciation diacritics to their arsenal, such as åçčêñțś and ʈìḷȡēṧ and üᵯłãūṯş, they actually reduced the visual possibilities for representing tone in a confined space. Recent attempts to document tone have resulted in a tangle of systems that are inconsistent and difficult to learn; the Wikipedia section on tone notation clocks in at about 3000 words, and is worth a peek to see how pernicious the problem is.[10]
Kamusi has returned to the fifth century to find a solution for marking tone. Before the printing press, books were written by hand, which meant that the scribes of illuminated manuscripts could change their colors or character sizes at will. Kam4D will contain space for tone, to be transcribed in a proposed markup language (to be implemented pending funding for the experiment). Let's pretend the word "beautiful" could have different tones on the middle syllable, giving different meanings. The way the markup would work will be something like:
The number of indicators needed for a consistent system across languages is fairly low. The result will be data that software can use to render tone with visual features, such as size and color:
Using color to represent tone is such an ingenious idea that the creators of the Pleco for iPhone Chinese Dictionary have already had it, and implemented it, in the years that Kamusi has been whispering to the wind with our partners at CERDOTOLA in Cameroon about making it happen. Where we might claim innovation is the universal markup method that can be used across languages, and having the data open to inspection, user contributions, and easy export. Through this system, an ordinary reader of any tonal language will be able to interpret tone without ever having to master tone notation. The big catch will be to work with large companies to incorporate the output within their systems for rendering text.
Kam4D also supports more complex audio for two future features. First, long passages from other Kam4D fields, such as definitions and usage examples, can be recorded, geotagged, and stored. This will be helpful to language learners and people with visual impairments, but the main goal is to build a repository of natural speech in multiple languages that is inherently aligned to written text, and can therefore be used for speech recognition and speech synthesis technologies. Second, we have software in the pipeline for talking dictionaries for unwritten endangered languages. When ready, field researchers will be able to elicit recordings from native speakers not just of a word like “star”, but short explanatory vignettes, such as “a star is a bright shining object in the sky”. By collecting these oral definitions alongside the terms they define, Kam4D will enable a new dimension for the documentation and preservation of the human heritage that is embedded in many languages that might otherwise disappear from memory. The database also contains a field for the transcription of oral vignettes, as well as fields for the translation of those vignettes to other languages, and, at risk of a recursive loop, the recording of those translations for the purposes discussed above regarding other long passages. These fields have been created with the long game in mind, as preparation for as yet unfunded data collection activities.
Having established how terms are represented, we can now discuss what they represent. Shape and sound are arbitrary adornment for the thoughts that words are intended to convey - a rose by any other name (waridi, 玫瑰, τριαντάφυλλο) would smell as sweet. Rose-ness is a free-floating concept that can be expressed 7000 ways across 7000 languages, with only a few languages like Nicaraguan Sign Language approaching an intuitive link between their representation of “flower” and the essence of the object.[11] In most cases, learning the meaning of a term comes from some combination of an explanation, awareness of how it is used in the real world, matching it to a known equivalent term in the same language, or matching it to a known equivalent in another language. As data, these elements are respectively definitions, usage examples, synonyms, and translations. A term may also have a meaning as the name of an entity such as a person, place, or organization. Additionally, the meaning of many terms can be conveyed by seeing an image of the item in question (Lew 2010); these could be moving images, which are extremely effective for action verbs (Moneglia et al. 2014). Kamusi has a unique approach to each, all of which are handled in Kam4D in a more complicated and effective way than in any previous language resource.
A term can be explained in many ways that may all be correct. (Or incorrect - see Adamska-Sałaciak 2012.) Regard four definitions for “sound”. All are correct ways of describing the same concept, pitched at different audiences:
Kam4D supports multiple definitions for any smurf, with future work to enable users to score definitions for level of difficulty. All definitions that have been brought in from Wordnet are stored as Wordnet_def; for historical purposes they all need to be kept due to their associations with other language projects, but some, like “a woman policeman” to define “policewoman”, are truly terrible. Our task list includes fields for citations for definitions that enter Kamusi from other sources, with the master reference available as a data item that can be called from any smurf; we hope to cull our sources through the Z39.50 protocol used across libraries, but we have not yet explored this bibliographic tangent to the main mission. Definitions can be recorded as audio, and those recordings can be individually geotagged.
Definitions can also have translations. Looking at the fourth definition of “sound” above, an Estonian speaker learning English might well want to read the explanation in Estonian, much as an English speaker learning basic Estonian might want an explanation of the Estonian word “heli” (meaning “sound”) in English, rather than the Estonian definition, “Keele (pillikeele), õhusamba, kile (trummikile), plaadi, häälekurdude jne ühtlasel võnkumisel põhinev akustikanähtus.” So, a definition for a word in any language can be translated in the definition_translation field by Kamusi participants to any other language, with the potential for the translation to be given source citations, recordings, and recording locations. We must stress that the Estonian definition_translation of an English definition for “sound” begins life as a different entity than an Estonian definition for the Estonian word “heli”, though they can serve dual duty if joined in the graph. The term as conceived in Language A might be subtly different from the term given as its equivalent in Language B, so, as a matter of lexicographic principle, we do not want to elide meanings across languages in the database.
People often learn words when they come across their use in context. For example, a then-eight-year-old we know explained, in the middle of a lake on a hot summer day, that her goggles had fallen off a dinghy because the boat had become “disequilibrated”. Obviously she had not gone to a dictionary to look up the definition. Instead, she had picked up the meaning of the term from hearing the ordinary discourse of her mother, a physicist. Similarly, some readers of this paragraph may have been unfamiliar with the term “dinghy”, but learned that it was a type of boat by seeing the word in a sentence where its meaning became evident. (Kamusi has the definition, the term’s translation to, e.g., Greek, and the Greek term’s definition in Greek, but not yet usage examples, at http://kamu.si/dinghy-eng-ell.) Good natural usage examples are surprisingly difficult to come by. “I might buy you a rubber dinghy for your birthday” is a poor example, while “Anyone who is responsible for a vessel at sea, from the smallest dinghy to an ocean going supertanker, must be familiar with International Colregs” paints an exquisite picture of what a dinghy is all about (both are from Twitter). Examples are complicated data features that Kam4D handles in a unique way.
For starters, examples are associated with smurfs, not terms. The muddle demonstrated by examples for “sound” in Wordnik,[16] where a lot of different definitions are shown on the left of the page, some examples are given on the right, and the reader has no indication of which example pertains to “sound” as noise versus “sound” as fitness, is solved in Kam4D. Or consider Wordnet, which provides the following examples for the synset «sound/ wakeless/ heavy/ profound»: a heavy sleep. fell into a profound sleep. a sound sleeper. deep wakeless sleep; not only are these poor examples of the concept, but each of the examples applies to only one member of the synset. In Kam4D (once we program how to process Wordnet data where the example appears in a costume) “a sound sleeper” will join to “sound”, while “deep wakeless sleep” will join to “wakeless”. Wordnik shows what happens when multiple meanings are lumped to a single shape, and Wordnet shows what happens when multiple shapes are lumped to a single meaning. Tying examples to smurfs, which are the intersection of a spelling and a meaning, solves both problems.
Examples might exemplify more than one term. For example, the tweet “riding on the dinghy through salt water mangroves (possible only at high tide) from the Bank into the Sound (Atlantic Ocean)” exemplifies both “dinghy” and the notion of “sound” as a body of water. Kam4D makes it efficient to link multiple smurfs to the same example. The links must be curated manually, though, lest a computer algorithm guess that the above sentence is useful for “salt” or “water” or “possible”.
Examples might also demonstrate a term wearing a costume, rather than its strict dictionary form. In the tweet above, for instance, “riding” is a costume of “ride”. In a language like Swahili, verbs in action will always have a different shape than the root form where the term is filed, except if they are commands to an individual. (“Kaa!, mbwa alimwagiziwa, na mbwa akakaa matakoni” - “Sit!, the dog was ordered by him, and the dog sat on its buttocks” - would be a rare example where the verb “kaa” (sit) is not transformed.) When possible, examples in Kam4D are pegged to their costumes, and discovered by the other costumes through this connection, such that the tweet above would be associated with “riding” but would also turn up in queries for the same sense of “ride”, “rides”, “rode”, and “ridden”. This is not always possible, though; for Swahili, for instance, verb examples that don’t happen to contain a second person singular command must be associated with the stem form, because we don't keep all 18 million+ legal forms of each verb as stored data, and even “matakoni” (on his buttocks) should probably be filed under the base noun “tako”.
The Swahili example also shows that usage examples need translations. Most readers of this article probably do not know Swahili at an advanced level, and even intermediate students would have a problem with a verb conjugation like “alimwagiziwa”. (The stem is “agiza”, equivalent to “order”, playing “Where’s Waldo” between three grammatical transformations on the front end and two on the rear.) However, with a translation of the example in English, an English-speaking learner of Swahili can interrogate the data as a study aid. If the example gets translated to, let’s say Yoruba, that will help Swahili students in, in this case, Nigeria. If 10,000 examples get translated to Yoruba, we begin to get human-quality parallel training material for machine learning that could lead to machine translation between two of Africa’s most economically important languages, spoken by a combined 150,000,000 people and currently bereft of any resource between them. (Technically, one could select “Yoruba” on one side in Google Translate and “Swahili” on the other, pass poorly through English, and receive words on the target side, but the result would read like the literary creation of a drunken chimpanzee.) Kam4D has fields for example_translation, the language of the translation, one or more audio recordings of the translation, and geotagging of the recording.
Examples also need source information, and not just because citation is the polite thing to do. Readers need to know the provenance of a usage example in order to assess whether the context resembles their needs - does it come from a science textbook or the lyrics to a hip-hop song? Kamusi is developing methods for the crowd to help harvest usage examples from a variety of open sources, such as Twitter and Wikiquotes. The ability to associate multiple examples with a smurf resolves the lexicographical dilemma of targeting users with disparate skills in a language (Humblé 1998). Further, the date of the quotation provides important historical information - there is a big difference between a term that was first used in 1589 and last seen in 1702, versus one that shows its first use in 2007. With dates as data, we can introduce a brand new tool, a time cloud, that will become useful for historical linguists of the future. Therefore, Kam4D has fields for the example source source_title (e.g. “New York Times” or “Twitter”), article_title, author, year, and url.
Tint is something that gives a certain extra color to a term, such as "archaic" or "offensive" or "informal". A term could have several tints, but the list of possible tints is limited. When a duck is designated as a tint, and the data has been completed, the term for that tint will appear in the UI in the user’s language, for example in drop-down editing menus.
Sorry, Google, you cannot define “automobile” as “a car”. (Nor can you define Google as a dictionary, though they make that claim on billions of page impressions.) Automobile and car are synonyms, both of which are defined by, well, let’s let Google get it wrong again with its definition for “car” that doesn’t recognize the existence of Tesla: “a road vehicle, typically with four wheels, powered by an internal combustion engine and able to carry a small number of people”. Automobile and car are synonyms - terms that refer to substantially the same concept.
Synonyms are important clues for understanding meaning, so much so that Kam4D lists them twice. First, we have “synset”, for clusters of synonyms produced by Wordnet (Fellbaum 1998), such as «car, auto, automobile, machine, motorcar». Second, we have a home-grown “synonym” field, which allows a lot more nuance.
The English synsets from the Princeton Wordnet (PWN) are both helpful and problematic (Benjamin 2014b; 2016). Kamusi used them as the fiber to weave together data from dozens of wordnets produced for other languages in relation to PWN. Many other projects also hook to PWN through json reference numbers, so it is important to maintain the integrity of that data even as we strive to move beyond it. Wordnet was never really outfitted to be a dictionary, though, which shows upon close inspection (Benjamin 2018). Synsets are one area where questionable decisions were made and now remain more or less locked in stone. Some synsets are too broad, and others overlook important twins. You be the judge of whether these five terms have a fundamental sameness: «estimate, guess, judge, gauge, approximate». At the same time, it’s anybody’s guess why “ship” is not in a synset with “boat”.
The Kam4D synonym field makes it possible to break apart overly generous synsets, document new synonym relations, and describe the subtle ways that near-synonyms differ. Synonyms connect smurf-to-smurf. In cases where they are exactly the same thing, such as car and automobile, they can also share a definition. In cases where the terms are slightly different, Kam4D provides an open text input field similar to the definition field, with all of the same translation and recording options, where contributors can describe the difference: "A boat can be any size, but a ship is always large. A cruise ship is either a boat or a ship ('They returned across the Atlantic by boat instead of flying'), but a dinghy is only a boat." For future work, we envision a differentiator, a grid to distinguish differences among a range of similar items, such as «boat, ship, dinghy, barge, canoe, schooner, raft, yacht», and establish how items match to smurfs in similar grids for other languages, which will be a new feature within lexicography.
Antonym is the antonym of synonym - a term that refers to a substantially opposite concept. The antonym of “hot” is “cold”, of course. Except when it is “stale”. Or “ugly”. Or “legal”. The antonym of the verb “dust”, meaning “cover lightly with a powdery substance”, is “dust”, meaning “remove powdery covering from a surface”. The antonym of “fact”, since the beginning of the Trump administration in 2017, is “alternative fact”, in fact. The notion that antonyms are easy is an alternative fact. The antonym of “easy” is “hard”. The antonym of “hard” is “soft”. Or “non-alcoholic”. There we go again…
We address the complexities of antonyms in two ways in Kamusi.
First, our antonym field is tied to a smurf, not to an overall spelling. Thus, each sense of a term can be joined in an antonym relation to the terms that embody the antithetical idea. The heat sense of “hot” has the antonyms “cold” and “frigid”, while the sexiness sense has an antonym relation with “ugly” and “unattractive”. “Dust” can even be shown as an antonym of itself, since the relationship is charted between two smurfs that just happen to have the same shape. Antonym relations are tracked through ducks, so “Staub wischen” (remove dust) in German can be contrasted with посыпа́ть (sprinkle with powder) in Russian, dusted with a bit of salt to show readers the basis for our computed assertions.
Second, Kam4D has a field for range as well as antonym. Concepts can mark grades on a scale, rather than being polar opposites, such as «scorching - hot - warm - lukewarm - tepid - cool - cold - frigid». The tool for placing smurfs on a range within the Kamusi web and mobile UIs is on the development tasklist at time of writing - one sticking point being how to place concepts on a spectrum when the borders are not clear, such as “fib”, “lie”, and “alternative fact”.
Ok, thousands of words into this article, we have finally laid enough groundwork to discuss translations. Kam4D evolved from a basic bilingual dictionary for translation between Swahili and English, a seemingly simple data matching task. Hah! For starters, one word often does not mean just one thing in one language - the problem of “polysemy”, cured by smurfs. Next, one word often does not translate to just one thing in another language - a double-sided problem of both polysemy and near-synonymy, whereby many terms from Language A could map to one term in Language B, or vice versa. Third, something in Language A might not map exactly to something in Language B - the problem of semantic drift, whereby "mkono" in Swahili is the body part from the shoulder to the fingertips, which in English are two separate body parts, "hand" and "arm". Fourth, something might exist in Language A that does not have a word in Language B - “lexical gaps”, such as an absence of "winter" in most tropical languages. These are the hurdles of transmitting meaning when just two languages are involved.
For a multilingual dictionary, each of these problems is compounded when languages C, D, E, and Z are stirred into the mix, expected to retain their own integrity, and expected to bond correctly in all combinations, such as Language F to Language Q. Because there are so many ways to go wrong, no multilingual dictionary has previously passed the laugh test. Kamusi is also guilty of some humdingers when our data sources are less than ideal (for example, the teams for both the French and Russian wordnets that we used for seed data used error-prone computational compilation methods, so we do not yet trust our own presentations for these languages). For data sets that have been curated by humans, though, the DUCKS system (Data Unified Concept Knowledge Sets) as implemented in Kam4D largely solves the problems of aligning term equivalents across languages.
Placing a smurf within a duck provides a general indication of equivalency among languages; a term in Marathi might be matched to a term in English, and the same term in English is matched to a term in Aymara, so DUCKS posits that the Marathi and Aymara term have roughly the same meaning. This computed link can, in principle, subsequently be validated by a bilingual Marathi/ Aymara volunteer. (That is an unlikely combo, but there are many qualified bilinguals for non-English pairs among European languages, or among diasporic African language speakers.) When a link has not been confirmed, it is displayed with its pathway (usually through English at this point), so that the user can follow the breadcrumbs.
Links from a smurf to a duck, or between individual smurfs within the duck, are refined with one of three types of equivalence label. First are things that are the same for all practical purposes, such as "eye" in English and "ojo" in Spanish, that are labelled parallel. Second are things that overlap but are not exactly the same, such as "mkono" and "arm" mentioned above, which are labelled similar. Third are the aforementioned “lexical gaps”, such as a way to express "winter" in Swahili. The goal for these cases is to find a few pithy words that can transmit the meaning of the concept in the second language, but not to claim that the new construct is an actual term. The relation between these concepts is labelled explained_by on the Language A side and explains for Language B: "winter" (a known concept) is explained by "majira ya baridi" (an invented term that back-translates as "season of cold"), and "majira ya baridi" explains "winter". (Explanatory terms may eventually become standard, such as the Turkish confection “lokum” that marketers in Los Angeles explained by “Turkish delight” when they introduced the product to the US in 1964.) Something labelled explains in Kam4D is not considered a true term in its own language, so it will be excluded from many search processes.
Equivalence is mapped through ducks. That is, if "bra" in French is labelled parallel to "arm" in English, and "arm" is labelled as similar to "mkono", then the similar equivalence will be shown automatically between bra and mkono, unless a human overrides that designation. It is important to note that the pathways for these inferences are always displayed to users, unlike services such as Google Translate that present their wildest stabs in the dark as though they were fact.
As with synonyms, translations have fields to explain difference as open text and to translate and record audio for those explanations. For example, the synonyms “bateau” and “navire” in French have a similar difference as “boat” and “ship” in English, but the boat/ship difference explanation does not migrate to explain bateau/navire, and the boat/navire translation_difference deserves a different explanation, and the bateau/ship translation yet another. Kam4D provides space for this data, but the data itself does not yet exist. Producing collaborative games that will direct students on the path to supplying data is on the task list as future work.
Kam4D does not yet support costume-to-costume bonds across languages, so that we can say with confidence that, for example, a third person past tense passive construction in Language A rings the same bells in Language B. The database is built with a granularity that can support such translations in some situations. For example, we can pretty much ascertain that “he corrido” in Spanish maps to “j’ai corru” in French, and that both can map to “I ran” in English, and that imperfect “yo corría” in Spanish maps to “je corrais” in French that maps to “I ran” in English, but we could not confidently propose how to get the right Spanish or French from “I ran” without further context. Determining and coding the needed models are roughly a Master’s project for each language. With enough confirmed data between some language pairs, machine learning could construct bridges among others (for example, if we had French to Fula and Wolof and Bambara, and Fula to Wolof and Bambara, we could make a reasonable start with machine inferences for Wolof to Bambara) - work for another decade.
Sometimes, meaning is enhanced by an explanation that does not fall easily into a lexicographic category. For example, a term might be considered pejorative if it is uttered by a member of one group about another, a mild example being “nerd”, but seen as fine if it is exchanged within the affected group. This uncategorizable extra information can be placed in a “notes” field, which can further be translated or recorded. Kam4D has four categories of note: usage, cultural, historical (which touches on time more than meaning), and special, which is a cop-out label to catch anything else.
The words on a page often refer to one specific thing, a “named entity” (NE). This is a person, place, organization, or other category of thing that has a name, might appear in print (for example a news article), but does not usually have a place in a dictionary. “Mount Anthony Union High School” is an example that is composed of five words, one of which is even a person’s name in its own right, but in combination refers to a specific educational institution in Bennington, Vermont.
Many NEs are available in open source databases, from a gazette of the name of every town in China to a list of every man who has ever played professional baseball. Knowing that something is a name is extremely important within Natural Language Processing, for instance determining whether to translate “mount” and “union” and “high” and “school” as individual words, vs. leaving the entire entity name intact. Some names cannot be translated at all, such as “Boston”, but they can be transliterated to other scripts. Conversely, more than 400 variations of “Muammar Ghadaffi” have appeared in European news publications over the years.
The Kam4D schema has several special features for NEs. A smurf has a Boolean tag to indicate whether it is an NE. Smurfs that are identified as NEs can be further categorized according to the three level “Sekine” hierarchy,[17] such as Location/ Geological Region/ River. Many NE databases include translations to a number of languages, so each of those terms become smurfs that are associated as ducks, and smurfs from other languages can be added to an NE duck at any time. Transliterations will be automated through models to be developed as future work, and stored in Kam4D to assist NLP when matching strings are encountered in text, in addition to transliterations such as “Muhammar Gaddafi” that can be determined from data sources. NEs that are associated with a place can also be geotagged, which will help distinguish the smurfs for “Boston” from among a city in the US, a town in the UK, a street in Lausanne, or one of three villages in Kyrgyzstan.
Harvesting NEs from open sources can be a remarkably straightforward student project: find a dataset, mark the common attributes (e.g., person/ athlete/ baseball player; location: USA), run the import process. NEs do not initially require definitions, usage examples, relations, or other lexical information, though such details could optionally be added later; audio recordings of personal and place names by local speakers of the language will be especially helpful, e.g. for locations such as Ynysybwl in Wales, Slough in England, or Worcester in the US. Kam4D is configured to include all relevant NE information in the graph for as many NEs as we can reasonably inhale, without making a fetish of obtaining encyclopedic information for each entity. The more NEs that are included, the more complete the matrix of human expression, and the better we can prepare machines to interpret the texts they encounter, within and between numerous languages.
Flora and fauna have names in local languages, and they also have Latin or Latinesque scientific names. For example, the common duck, which Kamusi shows as “بَطُخ ” in Kashmiri and “pidipidi” in Setswana, has the universal classifier “Anas platyrhynchos” that scientists and birders use to distinguish it from other species. These taxonomic designations are linguistic oddities, in that they are not actual words in an actual language that people speak among themselves. When the 19th century botanist Charles James Meller “discovered” a new species of duck in eastern Madagascar (already well known by its local name “harki” in Malagasy,[18] the sort of data that we seek to track down to add to the globally-accessible linguistic knowledge set that names the bird in 22 European and Asian languages), the Zoological Society of London invented the Latin-sounding “Anas melleri” to cement its identity, and named it “Meller’s duck” in English. “Anas melleri” was subsequently bequeathed with the name “Madagaszkári réce” in Hungarian, “マダガスカルガモ” in Japanese, etc., always using the Latinesque name as a hook so people could be clear it was the same creature. Kam4D supports a taxonomy label that can be associated with a duck, so that the name for a plant or animal species as rendered in any language can be pinpointed to its universal scientific designation.
Language names are a special type of NE within Kam4D. Each language has a name for itself, and quite likely has names for many other languages - for example, Zulu indicates that a word is a language name by appending the prefix “isi”, and thus can name 7000 languages, starting with isiZulu, without batting an eye. In addition to names for all known languages in English, Kam4D has the names for hundreds of languages as rendered in hundreds of languages, as produced by the Common Locales Data Repository and mapped through ISO-639-3 codes. These language names are then integrated in a language selector that we have engineered, so that someone using the Chinese UI who is looking for a term in Yiddish will be presented with the Chinese word for “Yiddish” within their search parameter options.
A picture is worth a thousand words, and pictures have a place within many dictionaries. In the image of the Google result for “define automobile” above, most readers probably understood the car-ness of the concept instantaneously upon seeing the photo, with the written definition as a secondary confirmation. Kam4D supports images in three ways, two of which are unique.
The conventional aspect of images in Kam4D is to associate an image directly with a term. An image is uploaded, along with a variety of boring meta-information such as the contributor details and copyright permission. After the image is vetted as safe by other users, it can be linked manually to any relevant smurf; a picture of a beach might be joined to “beach”, “ocean”, “seaside”, or equivalents in any language. This is the method that Wikipedia uses, for example, to enable editors to embed images from their “Wikimedia Commons” stock within any appropriate article. Kam4D will eventually also enable embedded links to Wikimedia images, with the premise that the data will remain stable, since there is no need for us to seek and host redundant images.
Our first innovation is to propagate images carefully within ducks. An image that has been associated with English “boat” will appear for the linked French “navire”, along with the important caveat that the image has been grabbed from some other smurf.[19] When programming is complete, users will have the option to vote for whether or not the image applies. (This article will not discuss the way the database deals with user objects or voting records.) In this way a picture of a dinghy can be disassociated from navire, while a picture of a cruise ship can be retained.
Our second unique innovation is “KamuSee”, a visual dictionary that is far along in the development pipeline at the time of writing. Elements of a picture can be tagged for lexical content, the way group photos on Facebook can be tagged with the names of the people shown. Just as Facebook has users link to the right “John Smith” in the photo rather than choosing a random “John Smith” from its database, KamuSee contributors link to the correct smurfs for their tags. When users search for terms that are within a duck shared by one of the tags, they are shown the master image with all the tags in their preferred language, as shown in the photo for the Fongbé language of Benin. As above, users will have the option to vote on whether a tag applies to a particular smurf that has been generated within a duck. Kam4D tracks associations between images, tags, bounding areas (don’t worry about it), smurfs, and ducks.
Time is the fourth dimension, the 4 in Kam4D. Languages evolve over time. That evolutionary path is important for understanding how words relate within contemporary languages, and how languages relate to each other. Most of the words that people have ever expressed have been lost to the winds the moment they were spoken. A certain portion of our language genomes, though, have been retained in the fossil record as written text. This historical information can be captured within Kam4D as linguistic data (defined for our purposes, remember, as “words that are digitized in a way that can be used within computer processes”), and subsequently interrogated in ways not previously possible for historical and comparative linguistics.
We should be clear: historical data is limited. Only some languages have left traces; we have data for Old English, but not for Old Cherokee. The traces are only to words in the form that they have been written down; if a language had a formal register for writing and an informal register for speaking, we only have records of the former. The vocabulary is limited to the subjects of interest to the scribes, so we might know a lot of terms about wars and very little about, say, raising children or raspberries. We can only speculate, at best, about the sounds that words made. The written record only goes back a few thousand years, while language has been around for tens of thousands. Thus, even when Kamusi succeeds in the long-term goal to document every word preserved from each language that has left footprints, the temporal dimension of our matrix will be woefully incomplete.
Historical languages are treated within Kam4D as any other language. Languages that have an ISO-636-3 code (34 “Old” languages, such as Old High German and Old Tamil, and 8 “Ancient” languages such as Ancient Aramaic and Ancient Zapotec) are included in the language table, enabling data to be added immediately across the full range of Kam4D fields. Hundreds of languages that have lost their last speakers in recent years, but been documented by linguists, also fall into this category. Other non-ISO historical languages can be added if scholars arrive with data. As with living languages, further effort to configure the database for such features as part_of_speech and grammar_attribute would make the data more robust, when language specialists can provide such determinations. As with contemporary languages, every term is treated as its own smurf, with inflected costumes hooked to the smurfs they belong to.
Dates are data, when it comes to tracking language over time. For languages such as Old English and Middle English, the date when a source text was written is known information. Therefore, a date field is available in association with source citations, often gathered within the context of usage examples (discussed above). By seeing first and last known uses, for example, one could discover when “ruthless” and “ruthful” entered Middle English from “ruth” (pity or compassion) in the early 14th century, and when “ruthful” and “ruth” dropped away three hundred years later.
Historical information is charted in two ways. First, Kam4D has a historical_note field that is an open text opportunity to expound on known information about a term’s origins. This information might be attached to a specific smurf, or to all instances of a spelling form, depending on whether it pertains to a split in meaning. Second, terms can have ancestor and spawn relationships, or be linked as members of a family when we don’t have evidence for which is the chicken and which is the egg. These relationships can occur within a language, such that “operation” is spawn and “operate” the ancestor within English. Or they can occur across languages, such that “safari” in Spanish entered as the spawn of its ancestor “safari” in English, which is the spawn of its ancestor “safari” (meaning “trip”) in Swahili. When the dates are introduced, etymologies can be traced backwards and forwards through time. For instance, the word “virus” can be traced back in time through Latin and Proto-Italic to Proto-Indo-European, which can be traced forward through Old Church Slavonic “višnja” (cherry) to the contemporary Romanian word “vișinata” for cherry liquor.
Some Kam4D features fall by the wayside in the documentation of historical languages. For example, there is no call to write definitions of Ancient Zapotec in that language, because no spirit will rise from a Mexican grave to take advantage of that information.[20] Nor can we make audio recordings of the examples we have from the written record. However, as a fifty year digital humanities goal, the database is designed to collate a large range of genomic linguistic data, resulting in dictionaries and other language technology resources that traverse the fourth dimension.
Search Google for “number of languages in Papua New Guinea” and you will get the answer, “At least 2”. More credible sources place the count somewhere between 820 and 851. This linguistic diversity is largely the result of geography, with each valley evolving its own way of speaking over the course of centuries. We can witness how geography causes languages to branch apart, with substantial differences between Portuguese as spoken in Brazil and Portugal, and the evolution of Afrikaans as a separate (but largely mutually intelligible) language from Dutch over a few hundred years. Imperial languages of violent conquest and subjugation, such as English, French, and Spanish, are expressed with different words, sounds, and shapes throughout their colonial wake. Small languages like Swiss Romansh remain central to the identities of the people where they are spoken, though migration and economic pressures endanger many such languages. Whether place is a unifying feature of a language or a cause of its variability, it plays an important role in how people express themselves.
“Dialect” is the murky term that people usually deploy in regard to geographic variation, but nobody can really say what a dialect is. A Londoner could walk into a pub with an American, a South African, and an Australian and get every joke, but be hard pressed to understand the English of her downstairs neighbor. Chinese likes to call many of its mutually unintelligible languages “dialects” when they mostly just share a writing system, while Burundi and Rwanda like to call Kirundi and Kinyarwanda separate languages when they mostly just differ somewhat by spelling. Some dialects actually are cut and dry, such as Kiamu being the dialect of Swahili spoken on the island of Lamu. Kam4D has a field to list the named dialects of a language and attach them to smurfs and translation blocks, but we are unenthusiastic about what it offers intellectually.
Our enthusiasm instead lies in a novel approach to language variation, through geotagging. Kam4D has a field for a smurf to be geotagged as pertaining to one or more locations. For example, “robot”, in the sense of traffic light, can be geotagged for South Africa and other nearby countries where it is used, and joined to the duck that contains traffic light, traffic signal, stoplight, Croatian “semafor”, Thai “ไฟเขียวไฟแดง”, Finnish “liikennemerkki”, etc. A Thai driving in Cape Town, or even an American, will thus be able to make sense of an instruction such as “turn left at the robot”.
This data solution was developed after a conference in Algeria exposed the fraught politics of North African lexicography, with even the names “Berber” and “Amazigh” being the subject of hot debate. Rather than trying to describe one or several varieties, geotagging will result in word clouds that can be consulted as they pertain to a given area. Some terms might turn out to be universal within the full range of a macrolanguage like “Hindi”, some to apply in Area A only, some in A and B but not C, and some just in B and C. The solution is equally applicable to the Arabic varieties spoken from Morocco to Tajikistan, the Fulani spoken in pockets across West and Central Africa, and many other situations.
Accents are another feature of place, as shown in this video of two Scots encountering speech recognition technology that is trained on a more dominant version of English: kamu.si/11th-floor. Various fields in Kam4D support audio recording, as discussed previously, and those recordings are geotagged with the location from where the speaker identifies their upbringing. These recordings could be used in future speech recognition and speech synthesis technologies, including as training material for machine learning.
Kam4D currently stores the names of every country in English and in about three hundred other languages, as provided by the Common Locales Data Repository. These country names are available in the user’s UI options in their preferred language, such as when they select their country in their user profile.
In principle, it is possible to define regions and towns as well, create ducks for known translations (e.g. Geneva, Genève, Genf, Ginevra, Ginebra), and geotag those ducks to a map. We’re not there yet. The section above about named entities discusses the route we intend to take. Location documentation remains in its infancy at Kamusi, but has been hardwired into Kam4D in anticipation of volunteers who will want to help develop language resources pinpointed to their localities.
Until now, we have been largely discussing terms as discrete entities - the premise of smurfs is that we can isolate a single intersection of a shape and a meaning that operates independently within a language. We have discussed semantic associations such as translations and synonymy under the category of meaning, but other relationships affect how readers and machines can access and interpret terms as functional data.
Terms are often connected in ways that are not evidently chronological, spatial, or semantic, yet we know they are somehow related. For example, "operation" and "operator" obviously share a relationship, but neither is the ancestor of the spawn of the other, neither is especially associated with one variant of English, and an operator is not usually the person who performs an operation. When we know that two terms have some sort of kinship, but can say little else about the relationship, Kam4D has a field for a reciprocal "family" relationship. This relationship can be smurf-to-smurf, smurf-to-wardrobe, or wardrobe-to-wardrobe. It can also cross languages - for example, we know that "música" in Portuguese is related to "موسيقى" (musiqaa) in Arabic, but we do not know the ancient path through which the languages acquired the common term. However, family relationships are not propagated into a duck, because we normally have no basis to posit genetic connections between terms that are linked through rather than shape.
Dictionary entries are often presented in an order that appears scientific, but is really pretty arbitrary (Lew 2013). With a term like "light", there is no logical reason for an ordering in a monolingual dictionary such as:
1. Not dark
2. Not heavy
3. Not fattening
4. Not serious
1. Not heavy
2. Not dark
3. Not serious
4. Not fattening
Nor is there an inherent reason to list all of the translations of a term in a bilingual dictionary in alphabetical order, such as these English Wordnet results for the Slovene term "lahek":
1. calorie free
2. light
3. lite
4. low-cal
When Kamusi was still exclusively an online Swahili dictionary, we frequently received the valid criticism from Swahili instructors that student essays were peppered with less preferable synonyms that they had gotten from our site - on the order of bovine in favor of cow. Students would look up a word and take the first result, assuming that the top listing was the most preferable. In fact, we were just returning results according to their ID in the database, which generally adhered to the alphabetical order followed by our student research assistants as they typed in entries from a source print dictionary, which would handle bovine many pages before it gets to cow.
The solution was a system for grouping and ranking spelling-based search results. A user would see the results in the then-current order, and have the option to slide them up or down the chart. Furthermore, they could add a dummy “break” line that they could position anywhere on the list. The software converted the ordering into numbers in a priority field, and used the break lines to designate a sense_group number for the items that were clustered together. Items within a sense group share some sort of meaningful relationship as well as some common aspect of shape, such as “street light” and “ceiling light”; different sense groups share a common aspect of shape but no semantic relationship, such as a “street light” that provides illumination and a “pilot light” that ignites a gas stove that are only related by the coincidental shape “light”.
Grouping and ranking numbers are hidden from the user, because they are not the result of some scientific process such as corpus frequency counts. What is important is that a knowledgeable person has spent some time thinking about the entry, decided that terms for “light” related to brightness are probably more relevant to searchers than terms related to caloric content, and decided that a logical arrangement of the caloric content results would be: light/ lite/ low-cal/ calorie free. A subsequent user could always tweek the arrangement, but the underlying lexical data is in no danger of malicious destruction. The stakes at play in ranking one smurf higher than another are pretty low, and future programming will make it possible to roll back any changes from a user determined to be malicious.
The 2005 version of the tool is shown in the image. Subsequent programming made a more user-friendly version for our years using Drupal, using sliders, that we do not seem to have ever screenshotted. Resurrecting the tool in our newest node.js/ angular platform remains on the task list as of this writing.
Grouping and ranking information is portable within a duck. Whether the search term is Polish “nietłusty” or Romanian “dietetic” or Finnish “vähäkalorinen”, the same order is given for results versus English that match to the food-energy sense of “light”, and “nietłusty” will similarly have results for Finnish that adhere to a Finnish speaker’s ranking of kaloriton/ kevyt/ vähäkalorinen. Groups are from searches that gather both exact matches and find_agent results (discussed below), such that “sabre-toothed tiger” will appear in the list for “tiger”, but not for “sabre” or “tooth”.
The graph architecture enables Kam4D to follow costumes in the presentation of groups and ranks. The search term “wound”, for example, will return the arrangements pertaining to the noun for a bruise or cut, as well as the verb “wind” for which “wound” is the past tense form. However, results for “wind” pertaining to air movement will not blow in, because those are not connected to any smurfs associated with costume “wound”. Non-specialists should not spend too much time trying to decipher how grouping and ranking winds through the multilingual data, but are encouraged to participate in improving the arrangements for each language when the tool is back online.
Terms can be related to each other through a variety of conceptual links. For example, a “wing” can be part_of a “bird”, or part_of a building (meronym↔︎holonym), or a kind_of appetizer (hyponym↔︎hypernym). These ontological relationships can foster connections that are not intuitive through words alone. In e-commerce, for example, one might look for a ‘belt’, find that it is a kind of ‘accessory’, and explore other kinds of accessories to find a matching ‘purse’.
Conducting such a search based on spelling could erroneously connect the ‘belt’ of a car engine with an ‘accessory’ to a crime and the payout ‘purse’ of a horse race. Using smurfs and ducks, though, Kam4D can ensure the right senses are linked. This example gets you from a Hebrew toy to the Chinese for Lego, through the real-world Amazon ontology:
סביבון ↱ Spinning Tops ↱ Baby & Toddler Toys ↱ Toys & Games ↳ Building Toys ↳ Building Sets ↳ 乐高
Conveniently, Wordnet is built around an ontology that describes both part_of and kind_of relations. These are mapped to synsets, to which many Kamusi ducks are also mapped. Thus, Kam4D is able to use the Wordnet data as a starting point for a logical dictionary organized around semantic relations, with the integration of additional ontologies, especially from medicine and the life sciences, remaining as future work.
These ontological pathways enable a major innovation in Chinese lexicography in particular.
As of this writing, about 77,000 Chinese smurfs connect to English and other languages through 42,000 ducks. The typical literate person in China can recognize about 3,000 or 4,000 characters from a catalogue of 106,230 that has no logical ordering and no relation between sound and shape. Before now, finding a Chinese term for which one does not already know its character-shape, along with the skillsets of how to compose characters using a keyboard (happily being rendered less necessary with handwriting recognition on touch screens), or the term’s sound and how to represent that using Roman letters, has been a search for El Dorado. The ubiquity of WeChat in China shows that the average Chinese mobile phone user can communicate using the characters in their quiver, but the 100,000 outside their repertoire remain constantly elusive.
Kam4D wires the Chinese terms to the English-based ontology, though a Sinocentric ontology is a long-term aspiration. A user inputs a Chinese term they know to be a part of a bird, such as 翼 (“wing”), and they can follow the chain to other bird parts they do not already know how to write, such as 利爪 (“talon”). An entry for “利爪” that is fully treated within Kam4D would be attached to a slew of data elements for the user to ascertain that they have found the right word, as discussed in the sections above. At the moment, our Chinese data only contains a term’s lemmatic form and part of speech, but the database is structured to support a rich and complex resource for the language. When mature, this logical dictionary will offer Chinese speakers an innovative way to access a language that hundreds of millions can speak with art, but have difficulty finding the words for reading and writing with anything approaching its full vocabulary,
A terminology is a set of vocabulary that has precise meanings within a domain of activity or knowledge. In architecture, a “ceiling” is a surface above your head. In aviation, a “ceiling” is the highest altitude a plane can fly. In architecture, a “window” is a glass pane you look through. In computers, a “window” is a viewing area on a display screen. Sports from soccer to horse racing to chess each have their own terminologies, such as “tight end” and other positions in American football. Medical specialties from pediatrics to gerontology have their own terminologies. The Catholic Church has its terminology, as does Islam, with both unique and overlapping concepts between the two. The function of a terminology is to leave no doubt about the meanings of the items under discussion; when air traffic control radios “caution wake turbulence” or “low-level wind shear on approach”, it is imperative that the pilot have no question about the terminology employed.
In lexicography, nuance is a feature. In terminology, nuance is a bug. Lexicographers seek to chronicle, for example, that the term “sabilli” in the Mampruli language of northern Ghana encompasses the range of colors from green to black. Terminologists want a clear differentiation between “green” and “black”, and other definable terms such as “olive” or “emerald” to pinpoint items more specifically. The complex lexicographic method within Kam4D of distinguishing shades of meaning among languages, discussed in the section on definitions above, can be inappropriately nuanced for terminology, where a thing is what it is, across languages.
Any smurf in Kam4D can be associated with one or more members of the domain_top field, such as “architecture” or “aviation”. These top-level domain labels are themselves smurfs that have definitions, translations, etc., so a Laotian speaker would see the Laotian label for aviation terms. Once a term is labeled as terminology, through the domain_member field, its definition in its own language is also transferred as the definition in that language for the translation of that term in other languages.
If a terminology item in Language A does not have an exact equivalent in Language B, it behooves the speakers of Language B to agree on a term that will pinpoint that item. Terminology creation is often a flawed process whereby one or a few experts proclaim that a particular term applies to a concept, then that term is inscribed in an obscure list that few people ever see or use. For example, the English-educated doctor and translator who teamed up to produce a Swahili medical dictionary claimed the term “inflamesheni” as the official equivalent of “inflammation”, which is incomprehensible to almost any potential patient who would readily understand the term “uvimbe”. Kamusi developed a pseudo-democratic system that produced ICT terminologies for 12 African languages, in which members of the public can propose and discuss potential terms, but experts make the final decision. This method results in terminologies that make sense within the language, and become available on the devices billions of people have on hand, the moment a term is approved - unlike the expensive Swahili medical dictionary that almost none of the hundred million speakers of the language could ever consult. In Kam4D, the usual option to label items as “similar” or “explains/ explained_by” between languages is removed for terminology ducks. Instead, a translation is listed as “proposed” until a designated overlord signs off on it, at which point it becomes “confirmed”. The database also supports features associated with the participants in terminology development, including interactive discussion threads, that are out of the scope of this article.
When we have the data, users are invited to generate custom lists, such as pediatrics and gerontology terminologies displayed in Laotian, Khmer, and Thai. These Neo4j queries are trivial within Kam4D, and easily accomplished through a web or mobile interface.
Hovering somewhere in the atmosphere above Kamusi is the need to localize our user interface (UI) into the thousands of languages through which people could wish to access the data. Kam4D treats localization similarly to terminology. A UI string such as “update profile” can be designated as localization. A localization item should have contextual information, such as a definition that explains the function of the string, usage examples when applicable, and perhaps a screenshot of the situation in which it appears. Like with terminology, members of the public can suggest and discuss potential translations. Like with dictionary entries and terminology, approved translations join in ducks. Unlike terminology, localization terms can be adopted by majority vote. Unlike dictionary entries, localization terms are not available through standard search, since they are not part of the lexicon per se. However, approved localization strings will be incorporated immediately into the Kamusi web and mobile UIs, and, as a special bonus, will be available via API for websites worldwide.
Kam4D is a unique graph database that is intended as a central repository for as much linguistic data as we can gather and codify. This article has described many intertwined elements that Kam4D assembles within the framework of “molecular lexicography”, a multidimensional approach to linguistic data as a composite of shape, sound, meaning, time, place, and relationships. Much of the described database structure now undergirds kamusi.org, with some of the promised elements (noted in the text above) still in the pipeline. We have certainly overlooked some features that should be included, and welcome suggestions for what to add. The big deficit now is data; it is time to populate Kam4D with a world of smurfs, and connect them through flocks and flocks of multilingual ducks. The project has much left to do - there is no conclusion.
Adamska-Sałaciak, Arleta. 2012. “Dictionary Definitions: Problems and Solutions.” Studia Linguistica Universitatis Iagellonicae Cracoviensis 129 (4): 323–39. https://doi.org/10.4467/20834624SL.12.020.0804.
———. 2013. “Issues in Compiling Bilingual Dictionaries.” In The Bloomsbury Companion to Lexicography. London: Bloomsbury.
Benjamin, Martin. 2014a. “Collaboration in the Production of a Massively Multilingual Lexicon.” In LREC 2014 Proceedings. Reykjavik, Iceland. https://infoscience.epfl.ch/record/200376.
———. 2014b. “Elephant Beer and Shinto Gates: Managing Similar Concepts in a Multilingual Database.” In Proceedings of the Seventh Global Wordnet Conference, 201–5. https://infoscience.epfl.ch/record/200381.
———. 2015a. “Crowdsourcing Microdata for Cost-Effective and Reliable Lexicography.” In Proceedings of AsiaLex 2015 Hong Kong, 213–21. Hong Kong. https://infoscience.epfl.ch/record/215062.
———. 2016. “Problems and Procedures to Make Wordnet Data (Retro)Fit for a Multilingual Dictionary.” In Proceedings of the Eighth Global WordNet Conference. Bucharest, Romania. https://infoscience.epfl.ch/record/221046.
———. 2018. “Inside Baseball: Coverage, Quality, and Culture in the Global WordNet.” Cognitive Studies | Études Cognitives 0 (18). https://doi.org/10.11649/cs.1712.
Colowick, Susan. 2008. “Systems for Multilingual Interaction.” Utilika Foundation.
Fellbaum, Christiane. 1998. WordNet: An Electronic Lexical Database. Cambridge, MA: MIT Press. http://mitpress.mit.edu/books/wordnet.
Lauder, Allan F. 2010. “Data for Lexicography The Central Role of the Corpus.” Wacana Journal of the Humanities of Indonesia 12 (2). https://doi.org/10.17510/WJHI.V12I2.116.
Lew, Robert. 2010. “Multimodal Lexicography: The Representation of Meaning in Electronic Dictionaries.” Lexicos 20: 290–306.
———. 2013. “Identifying, Ordering and Defining Senses.” In The Bloomsbury Companion to Lexicography. London: Bloomsbury.
Moneglia, Massimo, Susan Brown, Francesca Frontini, Gloria Gagliardi, Fahad Khan, Monica Monachini, and Alessandro Panunzi. 2014. “The IMAGACT Visual Ontology. An Extendable Multilingual Infrastructure for the Representation of Lexical Encoding of Action.” In Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC’14), 3425–3432. Reykjavik, Iceland: European Language Resources Association (ELRA). http://www.lrec-conf.org/proceedings/lrec2014/pdf/318_Paper.pdf.
Piotrowski, Tadeusz. 1994. Problems in Bilingual Lexicography. Wydawn. Uniwersytetu Wrocławskiego.
Rayson, Paul, Scott Piao, Serge Sharoff, Stefan Evert, and Begoña Villada Moirón. 2010. “Multiword Expressions: Hard Going or Plain Sailing?” Language Resources and Evaluation 44 (1/2): 1–5. https://www.jstor.org/stable/40666345.
Sag, Ivan A., Timothy Baldwin, Francis Bond, Ann Copestake, and Dan Flickinger. 2001. “Multiword Expressions: A Pain in the Neck for NLP.” In In Proc. of the 3rd International Conference on Intelligent Text Processing and Computational Linguistics (CICLing-2002), 1–15.
Zgusta, Ladislav. 1971. Manual of Lexicography. The Hague: Mouton.
[1] We can learn a lot when we are sure that a text means exactly the same thing between languages, sentence by sentence. Good parallel corpus resources exist between English and several other lucrative languages, particularly official translations of European Union documents. Some useful resources exist between English and a few dozen languages a little down the scale, but there isn’t much for non-English pairs at the top end, and nothing at all for over 98% of the world’s languages.
[2] You can compare a single quotation from Nelson Mandela, “Lead from the back - and let others believe they are in front” - as translated by human neurons, versus Google Translate’s “neural network”, for a large portion of the languages they cover, at http://kamu.si/mandela-lead. The association between which languages score well on this simple translation in Google and which fail is highly correlated with whether the language has good corpus resources in relation to English.
[3] The verb “to be” has eight forms, just for fun: be, am, is, are, was, were, being, been
[4] Verbs can have grammatical markers in as many as six slots. For cases where a relative pronoun marker can appear internally, the possibilities are combinations of 50 subject prefixes x 16 object infixes x 11 relative infixes x 27 tense markers x 20 extensions including plural command variations x 3 suffixes = 14,256,000 forms. Where the relative marker appears at the end of the verb, the possibilities are 50 subject prefixes x 16 object infixes x 27 tense markers x 20 extensions including plural command variations x 10 suffixes = 4,320,000 forms. This calculation does not include static, contactive, and inceptive suffixes that do not apply to all verbs, but will increase the total possibilities. On the other hand, we may have overlooked a few scenarios where the occurrence of X in one position obviates the occurrence of Y in another, which could slightly overstate the possibilities. Also, some things are possible linguistically but not in practice, e.g. a person cannot be split and a board cannot cry, so the associated forms cannot be constructed in people’s mental models. We welcome any Swahili linguist to help polish this calculation.
[5] https://twitter.com/suphannahrucker/status/1368954731058638852
[6] We are looking for a volunteer to do the legwork needed to get Adinkra into Unicode. Could that be you?
[7] UNICODE CLDR Emoji Names and Keywords, http://cldr.unicode.org/translation/characters-emoji-symbols/short-names-and-keywords
[8] According to Ethnologue, “The exact number of unwritten languages is hard to determine. Ethnologue (23rd edition) has data to indicate that of the currently listed 7,117 living languages, 3,982 have a developed writing system. We don't always know, however, if the existing writing systems are widely used. That is, while an alphabet may exist, there may not be very many people who are literate and actually using the alphabet. The remaining 3,135 are likely unwritten.” Eberhard, David M., Gary F. Simons, and Charles D. Fennig (eds.). 2020. Ethnologue: Languages of the World. Twenty-third edition. Dallas, Texas: SIL International. https://www.ethnologue.com/enterprise-faq/how-many-languages-world-are-unwritten-0
[9] MP3 files for “wound” wrapped around a coil are available for download at https://forvo.com/word/wound_%28past_participle%29/#en, and for “wound” injury at https://forvo.com/word/wound/#en. Even knowing that, there is no path for extracting that information, or similar recordings from other sources, for automatized downstream use.
[11] https://youtu.be/H9iO87IeQLE
[18] A speaker of contemporary Malagasy, from a different part of Madagascar than Meller spotted his duck, informs us that the name of the bird is “draki”. Either the name has changed over time, or the duck has different names across the island’s geography, or Meller got the local name wrong, or “harki” and “draki” are different birds. This is a small example of the sort of research question across space and time, touching here on colonial history and natural science in addition to comparative and historical linguistics, that Kam4D can facilitate.
[19] BabelNet attempts something similar, but generated via automated navigation that generates many false results. BabelNet grabs images that match to English labels in ImageNet or Wikimedia Commons, and then maps those images across Wordnet synsets. When this article was first published to the web, the results for “navire de croisière” (“cruise ship” in French) had 2 pictures of cruise ships, followed by 32 pictures of national flags, and then a bunch of pictures that are mostly nautical and sometimes related to cruise ships. This is an “I’m feeling lucky” approach to finding images, not a data-centric curatorial approach. When the BabelNet page was revisited in August 2021, all 40 images pertained to cruise ships, with no indication of whether a new automated process has been incorporated, or whether the page was manually changed.
[20] Hobbyists might prove us wrong. A version of Wikipedia written in Old English has over 3,000 entries as of July 2020, showing that some people enjoy the challenge of creating content in dead languages. Many of the Old English pages are minimal stubs, though, as with the entire page content for Romansh: “Retoromanisc sprǣc: Sēo retoromanisce sprǣc is Rōmāniscu sprǣc, ȝesprocen fram 35,095 lēodum, mǣst in Graubünden.” https://ang.wikipedia.org/wiki/Syndrig:Random will bring you to a random Old English article.