We take a lot of shortcuts when speaking, it's actually much harder to transcribe phonemes than to transcribe words, even when aware of the language being spoken. Some models have been trained for the task (e.g. look at https://huggingface.co/spaces/KoelLabs/IPA-Transcription-EN ), but the error rate is really high.
There are, broadly, two kinds of audio recordings that linguists want to transcribe. One is native speakers telling traditional stories, where they're speaking naturally and taking the natural shortcuts (such as "wanna" and "gonna" in English). The other is native speakers reading words (or short example sentences) very carefully and distinctly, so that the linguist can listen to the recording over and over to learn how to pronounce the word right. In those recording, they'll say "want to" and "going to" rather than "wanna" and "gonna".
Thanks for the pointer; I'll check out that model and see if it handles the "slowly and carefully" type of recording better than the "natural speaking" type. (And depending on what kinds of errors the model makes, even the recordings where it makes errors can prove useful: for example, a linguist studying regional variations in speech would want the model to produce the IPA for "gonna" rather than "going to").
Dialects degenerate phonemes that would otherwise occupy identity relations between different utterances of the same "word" (word/concept mutable hyperobject as is the standard in any socially-relevant spoken language) which would be a bit of irony in this thought experiment since common knowledge dictates that more data samples must be present in the dataset (not less, as in rarely-spoken languages) to associate separate pronunciations of utterances representing the same underlying concept. However very-rarely-spoken languages probably don't have distinct dialects since so much focus is put on mutual intelligibility with the few members of the group that remain fluent in that language. It's not outside the realm of possibility that small speaking communities nonetheless fractionate into dialectical specialities but that seems increasingly unlikely as the fervor for preserving/recognizing dying languages increases, and global instant communication continues to become more commonplace.
Example: Schwabisch is wild and would be phonetically transcribed very differently from Hochdeutsch which is its ostensible language progenitor (technically more a cousin than an ancestor in the lineage of language evolution), but if the goal is merely to focus the model purely on phonetic transcription then you can add additional post-processing layers which map sounds to core concepts shared across dialects for actual translation. But I like your idea of interacting with the intermediate elements to familiarize yourself at least with the phonetic patterns, we humans are still thinkers enough to infer patterns of grammar and semantics from these building blocks just as we have done for the entire history of the species/lineage before written representations of language came along (relatively late -- evidence of script cropped up only once civilization had centralized to a sufficient degree to make economics non-local and non-trivial).
tl;dr the big words: it's not til you collect enough spoken samples of the dead(ish/dying) language being spoken that the local idiosyncracies are discovered, luckily linguists are smart enough to probably anticipate and certainly post-process language snippets to grasp the common structures for this or that given language.
> ... very-rarely-spoken languages probably don't have distinct dialects ...
That's true if you mean "very rarely spoken" literally, as in even the native speakers don't get to use it very often. But many languages aren't widely spoken (such as only in a certain geographical area, which sometimes is only a single village, or other times a small number of villages). But inside that area, they are frequently spoken. And you might be surprised how many of those small-geographic-area languages still have distinct dialects.
For example: my wife (a linguist) did her master's thesis on the pronunciation of a language with about 7,000 speakers, and identified how many distinct dialects there were. (Which is why I know a little bit about this). She recorded native speakers from all 13 (I think it was 13, but it might have been 14) villages where the language was spoken, and found five different dialects, which she grouped into two "main" dialects. (Think American vs British in the English language, with subdivisions into Midwest, New York, and New England accents and so on, and you'll have the right general idea — though these dialects were closer to each other in sound than Midwest vs New York). I'd have to go reread her thesis to give you any more details. But this was a language that was only spoken in a small geographic area, but it was frequently spoken, because that was the main language of those villages. (The country's official national language is what the kids learned in school, but some of the people, mostly those 60 years old or older, hadn't gone to school, because the first government school in their area was only built 60 years ago -- so they only spoke their minority language, and not the country's language, and their kids had to translate for them if they had to leave their village and go shopping in a major town).
Which suggests another approach to the solution; you have 20Ls and 20Rs which can appear in any order so it’s a permutation with repetition problem and thus 40!/(20!20!)
That approach has the advantage that it’s easily adapted to non square rectangular grids (n+m)!/(n!m!)
Even his examples (Meta, Google, Airbnb, Uber) are "hated and accused" mostly by pundits and activists. Their users and investors are overall rather happy with their services (all have healthy competition) and their employees can quit and work for somebody else at any moment.
It's almost like the people who spend their time thinking and researching the positives and negatives of a particular company/service come to different conclusions than the people who make money off the company or are directly marketed it.
> people who spend their time thinking and researching the positives and negatives
If only we could trust such big hearted people who selflessly donate their time and expertise for thinking and researching - without asking those pesky little questions like "what do they stand to gain? what hidden agenda do they have? what ideology dictates their value system? what axes to grind and biases do they harbor?"
I'd rather trust the people directly involved, whose interests are mostly clear and known.
>I'd rather trust the people directly involved, whose interests are mostly clear and known.
That seems like a wild statement to me. Just so I understand what you're saying. Are you saying you would trust Meta's statements on if their algorithms are harmful to youth/society over independent researchers because those researchers might have some hidden bias?
No, I am not saying I trust Meta's statements at all. I am just saying that I have ZERO trust in those "independent researchers".
I am old enough to remember when "researchers" and "experts" where warning us about the evils of violent FPS gaming, the whole panic over teens and kids playing Doom. The Columbine shooting. The end was nigh! Of course, it was all blown out of proportions...
At least for meta, its employees, and its customers (advertisers), and its users, you can infer easily why they are involved. Researchers have other motivations like currying favor via an opinion or paper with a particular benefactor, or the tenure game, or 1000 other hidden things you cannot reason about without disclosure on their part.
Seems 100% AI generated and automated, the judge also seems suspect - in the first one it's actually GPT-5.5 pro which has the correct email RE: the deepseek one will match a@b.com1 as "a@b.com" while 5.5 will correctly require a word boundary at the end of the email.
I quit after this. No test-cases = useless judge.
11% capacity is not 11% MFU. The first is about actually using the hardware for something in the first place, the latter is about how efficiently you compute. Different things.
The original post of 11% is referring to leaked numbers about _MFU_, which has erroneously been re-reported as fraction of GPUs being used at all. The parent post is trying to correct this misconception.
Using existing enterprise apps probably - this solution is scalable for the vendor and it's easier to sell using existing software as-is than to start out by writing new custom tools.
After doing few experiments, I think that having Agents work on browser for all tasks wouldn't be best due to many factors like token cost, safety, etc. But browser/computer can be a tool that the agent can be alongside MCPs to complete tasks that requires interaction with such modalities.
Yes, I can see the usecase for legacy desktop apps etc, but the web? it's a DOM. WebMCP coming now too, no need for screenshotting or DOM querying then either..
Mid-way I realized this was AI writing (took me a while), then I read a quote in the text about a comment that "The tragedy isn’t that they cheated; it’s that the system was designed to let them thrive for a decade before anyone bothered to look at the data." I didn't find this comment in EJMR, or anywhere on the internet except the OP post, for that matter.
Moshi was an amazing tech demo, building the entire stack from scratch in 6 months with a small team was an amazing show of skill: 7B text LLM data + training, emotive TTS for synth data generation (again model + data collection), synth data pipeline, novel speech codec, rust inference stack for low latency, audio LLM architecture incl. text "thoughts" stream which was novel.
But, this piece is a fluff piece: "underfunded" means a total of around $400 million ($330 million in the initial round, $70 million for Gradium). Compare to Elevenlabs who used a $2 million pre-seed for creating their initial product.
A bunch of other stuff there is disingenuous, like comparing their 7B model to Llama-3 405B (hint: the 7B model is a _lot_ dumber). There's also the outright lie: team of 4 made Moshi, which is corrected _in the same piece_ to 8 if you read enough.
I’m a hands-on engineer who’s spent the last 6 years doing freelance ML + data science, primarily in audio/speech, and before that 10+ years in startups building and scaling production systems.
I’m looking for where research meets real systems: training and/or inference for large models, especially roles that value end-to-end ownership. Open to freelance engagements or full-time roles.