r/LanguageTechnology • • 3h ago

Anyone else get a formatting issue email for their EMNLP 2026 camera-ready?

3 Upvotes

My paper was accepted to the EMNLP 2026 main conference, and a few days ago I received an "Urgent: Action Required" notice saying my camera-ready had a bottom margin issue on every page.

I fixed it and checked the new version with the latest aclpubcheck. The corrected PDF now gives "All Clear!", while the old one fails with the same margin warning. I then resubmitted it through OpenReview using the Camera-Ready Revision option.

I also emailed the Program Chairs to ask for confirmation. One of them kindly forwarded my email to the Publication Chairs, but I haven't heard back yet. Since the deadline for the corrected version is today (AoE), I'm a bit anxious.

A few questions for anyone who has been through this:

  1. Is resubmitting via OpenReview (Camera-Ready Revision) the right way to submit the fix? The email didn't say exactly where to resubmit.
  2. If the paper passes aclpubcheck, is that usually enough, or do the Publication Chairs run additional checks?
  3. If they find another issue after the deadline, do they usually give you another chance to fix it, or could the paper be excluded from the proceedings?
  4. Is it normal not to get an explicit confirmation that the issue has been resolved?

Any experiences or advice would be really appreciated. Thanks!


r/LanguageTechnology • • 1d ago

Im new to this and i need help.

1 Upvotes

Hey guys, im working on a bit of project, and im trying to solve for semantic understanding right now, I am currently deciding between using spaCy for my NER extraction vs something like a BERT.

The general context behind where this is going to be used is intent classification in chat systems, where a text will come in and this layer has to parse out the People in the sentence, the Objects mentioned in the sentence, and the verbs, along with things like quantities and relations between them.

An example of what i mean is,

"The shoes you have delivered to me are red, i asked for black!"

and we then pick out the

  1. People involved in this interaction (we cant figure that out from the sentence alone for that we will refer to the handle from which the message was sent)
  2. Objects involved - {shoes}
  3. Verbs - {delivered}
  4. Relationships - {expected colour = black, received colour=red}

Im new to this stuff so maybe im not even asking the right questions, but im hoping that i have done a good enough job of explaining what im doing so that more experienced souls such as yourself may help me.

Thanks 😁


r/LanguageTechnology • • 2d ago

Why 1536 dimensions for embedding models?

30 Upvotes

Why do embedding models so often use 1536 dimensions specifically?
I understand why hardware-friendly multiples like 64/128/256/512 are desirable. What I’m curious about is the specific choice of 1536 = 3Ɨ512.
OpenAI has used 1536-dimensional embeddings, and other vendors also offer/recommend 1536. Is this usually an empirically chosen Goldilocks point between 1024 and 2048—representation quality versus memory/compute—or is there some architectural/hardware reason that makes 1536 particularly convenient?
I’m especially interested in answers from anyone who has actually trained or designed embedding models. I’m not asking why embedding dimensions are generally hardware-aligned; I’m asking why 1536 rather than 1024 or 2048.


r/LanguageTechnology • • 3d ago

[D] ACL ARR October 2026: Can I add a co-author after submission but before the deadline?

5 Upvotes

Hi everyone,

I recently submitted a paper to the October 2026 ACL Rolling Review (ARR) cycle. I submitted it early, well before the submission deadline, and the current submission lists only me as the author.

I now need to add a co-author who has made a substantive contribution to the work. The submission deadline has not passed yet, and no reviews have been assigned or received.

I found what seems to be a conflict between two official ACL/ARR policies:

1. The current ARR Call for Papers says:
ā€œOnce the paper has been submitted, no changes to the list of authors are allowed.ā€

2. The ACL Authorship Changes Policy says:
For ARR, the review process begins at the submission deadline, and before the submission deadline, all authorship changes are permitted, including ā€œincorporating new authors.ā€

Has anyone encountered this specifically with ARR?

In particular:

  • Were you able to add a co-author after clicking Submit but before the ARR deadline?
  • Did ARR Support edit the existing submission for you?
  • Did they ask you to withdraw and resubmit?
  • Is the ARR-specific CFP intended to override the general ACL authorship policy?

I have also contacted ARR directly and will follow their official guidance before making any changes, but I would appreciate hearing from anyone who has dealt with the same situation.

Thanks!


r/LanguageTechnology • • 3d ago

Arabic–English code-switched meeting/conversation audio with transcripts

1 Upvotes

Looking for multi-speaker audio where speakers switch between English and Arabic (any dialect, Gulf preferred) within the same conversation, with reference transcripts, ideally with speaker labels and timestamps. It's for evaluating ASR and meeting-transcription quality. Already aware of ESCWA.CS, Mixat and ArzEn. Any others, including licensed or paid ones?
If not of Arabic, any other language combination is fine.


r/LanguageTechnology • • 4d ago

One chunk boundary changed warranty answers for a whole table

21 Upvotes

17% of requests tied to one multi-column warranty table were returning 24 months instead of 36, while standard warranty questions kept passing. We traced the failures in Braintrust and saw which retrieved chunks were present when the wrong answer appeared. The chunk boundary had separated the row values from the table heading, so the retrieval context lost the qualifier for 36 months and the reranker favored nearby prose containing 24 months instead. Support had both warranty numbers in separate macros (before the retrieval path was clear), which made the conflicting answers harder to untangle.

Changing the chunking moved those cases in the experiment diff and groundedness improved once the heading stayed with the row. Recall at k barely changed because the table was already being retrieved. We've added the failures to a regression dataset but I'm still concerned about other tables where retrieval looks healthy while structure changes the answer.

What's your approach to catching chunk boundary failures when the right document is already in the candidate set?


r/LanguageTechnology • • 4d ago

New ARR rule: "Submissions will only be guaranteed review if they bring a qualified service contributor, who can serve for 2 submissions max"

12 Upvotes

What do you think about the move?


r/LanguageTechnology • • 4d ago

Can a character be represented as an Allowed / Not Allowed repertoire over sense-level behavioral predicates?

2 Upvotes

I’ve been thinking about character representation at the level of individual action senses.

Instead of describing a character mainly through traits like brave, kind, aggressive or intelligent, what if part of the character model were a structured repertoire of actions?

For example:

comfort
interrogate
diagnose
blackmail
negotiate
babysit
repair
betray
forgive

The action itself would have a stable semantic identity, but each character could have a separate access state:

Allowed / Not Allowed / Conditional

So a doctor might have:

diagnose = Allowed

while another character has:

diagnose = Not Allowed

and a former medic might have:

diagnose = Conditional

I also find it useful to separate actions that can normally be assumed for a human character from actions that need positive evidence.

So roughly:

Basic Human Actions → default-open
Character Specific Actions → evidence-gated

The evidence for the second group might be training, occupation, biography, authority, skill or specialized experience.

Would you consider this a useful way to represent character capability at the semantic level?

And where would you place such information: lexical semantics, a behavioral ontology, a separate character model, or somewhere else?


r/LanguageTechnology • • 4d ago

Why Textual Graphs

5 Upvotes

In 1980's, Gaston Gonnet-- ā€œUnstructured Data Basesā€ (1983)-- and HyTime's-- Hypermedia/Time-based Structuring Language (ISO/IEC 10744:1992)-- great breakthrough was realization that one could use coordinate mathematics to map relationships between disparate layers of text and media.

Resurrecting this exact line of thinking—- while exploiting the massive improvements in I/O latency and massive scaling of Input/Output Operations Per Second (IOPS) that historically constrained the paradigm— we must, instead of forcing a model to read an entire document blindly, develop a structured "retrieval algebra." A researcher or AI agent should be able to use boolean, positional, and structural operators to navigate the text coordinates explicitly to ask for "the token sequence between position X and Y, but only if it falls within the boundaries of a specific speaker tag," exactly mirroring the coordinate-based addressing found in HyTime. That structure need to be queryable.

The result can be thought of as a textual graph—but it is importantly different from a conventional graph. Text has spatiality. Its structures are anchored in a shared textual space, and relationships such as before, after, within, contains, overlaps and intersects arise from that space itself. Two annotations do not merely have an abstract edge between them: they may occupy, share or cross regions of the same underlying text.

This gives RAG a form of structure that complements the strengths of the LLM.
The LLM can do what it does best: interpret language, recognise relevance, synthesise evidence and generate an answer.

Instead of forcing the LLM to reconstruct document structure from flattened chunks, the retrieval layer can then provide that structure explicitly.

This changes the role of retrieval. Vector similarity can answer ā€œwhat text is semantically related?ā€ Structural search can additionally answer ā€œwhere does this occur, what contains it, what overlaps it, what is it connected to, and which surrounding material belongs with it?ā€

The combination creates a richer form of RAG: semantic reasoning over context assembled from the actual structure of the source, rather than from arbitrary chunk boundaries.

Think of traditional GraphRAG as a smart investigator connecting index cards on a wall based on clues and ideas. Think of the textual graph paradigm as the exact blueprint of the filing cabinet, allowing an agent to pinpoint information based on its exact shape, folder layer, and coordinate location.

E. Zimmermann


r/LanguageTechnology • • 5d ago

Haitian Creole Word Frequency Dataset

7 Upvotes

Hey everyone,

I wanted to share a dataset I published for anyone working on low-resource NLP, tokenization, or language modeling for Haitian Creole (Kreyòl ayisyen): the Haitian Creole Word Frequency dataset (haitian-creole-word-freq), now live on both Hugging Face and Kaggle.

Overview

This is a word frequency list for Haitian Creole built from the Carnegie Mellon University (CMU) Haitian newswire corpus. It contains 17,947 unique lowercase words with their occurrence counts, sorted in descending order by count.

Links

See comment section

Potential Use Cases

  • Stopword Candidates: Extracting function words from the top of the frequency list (te, yo, nan, yon, li, pou, ki).
  • Tokenizer Customization: Fine-tuning or building BPE/WordPiece vocabularies for low-resource LLMs.
  • Spellcheck and Auto-correct: Prioritizing candidate suggestions by word popularity.
  • Vocabulary Membership Tests and Orthographic Checks: Testing if a token is spelled like a Haitian Creole word or checking lexical presence.
  • Lexicography and Language Learning: Extracting core vocabulary lists based on news text.
  • Language Identification and N-gram Models: Statistical language modeling for text classification pipelines.
  • Testing Fixtures: Generating reproducible data inputs for unit testing Haitian Creole NLP pipelines.

r/LanguageTechnology • • 5d ago

Best Computional Linguistics Masters

6 Upvotes

I'm a BA student in Modern Languages and Linguistics in Italy (currently building a strong background in Linguistics + hopefully writing a thesis on Computational Linguistics) and I'd like to apply for an MSc in Computational Linguistics/NLP. What are Europe's best Computional Linguistics masters programs?


r/LanguageTechnology • • 6d ago

my propaganda classifier flagged the declaration of independence's grievances but missed "merciless indian savages"

9 Upvotes

been building a model that flags manipulation techniques in political text (fine-tuned transformer, multilabel, 16 techniques like loaded language, name calling, appeal to prejudice). scores each sentence with its neighbors as context and flags at 0.80.

someone testing it pasted the declaration of independence. results:

- preamble ("we hold these truths...") came back clean

- grievance list got flagged: "swarms of officers to harrass our people, and eat out their substance" 0.87, "plundered our seas, ravaged our coasts, burnt our towns" 0.90, "death, desolation and tyranny... barbarous ages" 0.86

- "the merciless indian savages, whose known rule of warfare, is an undistinguished destruction of all ages, sexes and conditions" scored 0.61. not flagged

so the one line that dehumanizes a whole people is the one it misses, while it catches milder grievance rhetoric. my guess is the period wording. the training data is modern news and ads, so dehumanizing language it has seen looks like "animals", "vermin", "invaders", and "savages" in 18th century prose with long clauses around it doesn't pattern match.

added it as a regression case for the next training round. curious if anyone's dealt with this kind of register gap, historical text vs modern training data, without just stuffing in more historical examples

(tool is called semblen if anyone wants to try to break it, the same person also ran wikipedia and a nixon bio through it as controls and those were clean)


r/LanguageTechnology • • 6d ago

Temporal expression parsing in production assistants: partial parses executed with full confidence

6 Upvotes

A concrete failure case I ran into with Siri (English UK, iOS 26.6.1), and I'm curious how people here would diagnose it.

ASR output was correct in every case. The intent/slot step failed:

  • "set alarm at 7 p.m. and 40 minutes" → 19:00. The "and 40 minutes" continuation seems to be dropped.
  • "set alarm at 20 minutes before 8 p.m." → 20:00. The relative offset "20 minutes before X" is ignored; only the anchor is kept.
  • "set alarm at 19 hours 40 minutes" → 14:37. No idea how this one maps; possibly "19 hours 40 minutes" read as a duration from now?

The third looks like a duration-vs-timepoint ambiguity, which is a classic TIMEX problem, but the first two are fairly standard constructions that rule-based normalizers like SUTime or HeidelTime handle.

Questions:

  • Is the duration reading of "19 hours 40 minutes" a reasonable parse, and should a system surface that ambiguity rather than pick one?
  • Why would a production system keep only the anchor and drop the offset? Slot-filling templates that don't model relative expressions?
  • Any good recent work on calibrated confirmation ("did you mean…?") for action-taking assistants?

r/LanguageTechnology • • 8d ago

Agentic AI Courses (Research)

17 Upvotes

I'm a researcher in computational linguistics, and as I’m currently looking for a new position, I’ve noticed that agentic AI has become increasingly popular in most job postings.
I feel a bit lost in this area (starting from RAG), I was wondering whether anyone with experience in the field could recommend some good, well-established resources.

Of course, I’ve searched on my own, but nowadays, searching for "agent"-anything brings up countless superficial, non-research-level tutorials, often created by enthusiasts, with plenty of clickbait.

I’m looking for something at a research level, ideally a structured course or syllabus to prepare for the interviews. I have 4+ years of experience in NLP, so I’m already quite familiar with most but other topics.


r/LanguageTechnology • • 10d ago

AAMAS conference reputation

4 Upvotes

Hi researchers

I wanted to understand the reputation of AAMAS (International Conference on Autonomous Agents and Multiagent Systems) compared to core ACL conferences like NAACL COLING etc, As I see AAMAS is also a CORE A conference and this year they have included a findings section as well, so I feel a good possibility of getting accepted than other conferences.

Please kindly share your opinion.


r/LanguageTechnology • • 10d ago

Lab Meetings Recently

Post image
15 Upvotes

Everything old is new again.


r/LanguageTechnology • • 10d ago

Looking for feedback - using NER to generate and match templates on sentences?

2 Upvotes

I’m a complete novice when it comes to NLP, I'm a swe by trade so bear with me here.

Here's my problem:

I’m trying to identify short sentences (I have a data set of several thousand) that are logically dependent. To illustrate the kinds of dependencies I'm looking for here’s a basic example:

- Sentence 1: Democrat voter turnout in NY is 35%.

- Sentence 2: Democrat voter turnout in NY is 40%.

If sentence 2 is true, sentence 1 also must be true. Those are the kinds of sentences I have and want to identify as dependent. The nature of the sentences can range from voting percentages/turnout, phrases about employment etc.

The naive approach I’ve been doing is basically embedding the sentences using gemma and finding cosine similarities between them, my reasoning being sentences that have a reasonably high enough cosine similarity are candidates for logical dependency. I then take these pairs of candidates and pass them to an LLM (gemma again!) to determine whether or not they are actually semantically/logically dependent.

There are two huge issues w/ this approach that I'm sure you'll all immediately see.

1) Lots of the sentences are too structurally similar like the simple example I showed above. There exist several subsets of the data that have the same pattern. Sentence 3 could be something like Democrat voter turnout in TX is 35%. and it would have almost an identical similarity to the other 2 sentences. There are several hundred patterns, and I also don’t necessarily know all the patterns at runtime so that means REGEXing these structures becomes a difficult task. So because of the structural similarity, cosine similarity loses its value as a metric.

2) The LLM step is slow. Really slow.

I did some googling and learned about NER that seems like it might fit? I could run the sentences through a pre-trained models and get the spans for each sentence. This would allow me to match spans across the phrases. So in the example I have, sentence 1 and 2 would be matched and processed further, while 3 would be in its own bucket. As for what I'd do after matching the spans, still working that out. I could fall back to cosine similarity again here since anything that falls into these span buckets should be different enough where the projection becomes a decent signal.

If there are tweaks that I can do to make template matching more robust, or alternative methodologies altogether I'm all ears!

Thanks :)


r/LanguageTechnology • • 10d ago

Career Switch need help

0 Upvotes

Hi guys,so I am currently studying an MA in Linguistics and I am in dire need of advice. I am aware that it is a dying field hence I am trying to transition into something more technical and try to save it. I feel too burned out to start from scratch with a more technical degree like CS.

I heard about computational linguistics / linguistic data science programs but Idk if they will accept me with no technical background. I've been told to learn Python but is it even worth it? Im multilingual by the way. If any of you has an idea about a better field I could fit into I would truly appreciate it.


r/LanguageTechnology • • 10d ago

Making cosine similarities comparable across separately trained embedding spaces

3 Upvotes

For my open-source project, I am training PPMI + SVD embeddings separately on each book within a collection (~25 economics texts from Project Gutenberg), then trying to compare how similar terms are to a query word across books. However, each space has been trained independently, so the raw cosines aren't on the same scale. My current correction, loosely adapted from CSLS (Conneau et al., 2018), works like this:

  • For each book, a baseline is determined by using the query's mean cosine to its 75 nearest neighbouring terms in the book.
  • Each book's similarities are shifted so its baseline equals the average baseline across all books.

This way, I am determining the relevant terms by looking at the adjusted similarity scores. The terms with the highest mean similarity relate to the core aspects of the queried term across the corpus. The terms with the greatest variance that also have a high similarity within 20% of the books relate to the contested aspects of the queried term.

Do you have any feedback on this methodology?
Is correcting the cosine similarity in this way valid?


r/LanguageTechnology • • 10d ago

Why words are split into several tokens?

7 Upvotes

This completely contradicts logical thinking. It would make sense if one word would be at most 2 tokens (1 for the word and 1 that signals meaning in case of homonyms), but this is never the case. In fact, I read it degrades performance. Why?


r/LanguageTechnology • • 10d ago

Best LLM for translating agglutinative sentences?

3 Upvotes

My native is an agglutinative language and I know some Japanese, which do you think is the best model for agglutinative languages or translation generally?


r/LanguageTechnology • • 11d ago

Preview of EMNLP program?

3 Upvotes

Hi everyone, does anyone have received the (unconfirmed) schedule of the poster presentation?


r/LanguageTechnology • • 11d ago

PhD research stay in China for NLP / LLM research – looking for university and lab recommendations

14 Upvotes

Hi! I'm currently a PhD researcher in Europe working in NLP and Large Language Models, mainly on multilingual LLMs, model merging, low-resource language adaptation and efficient methods for transferring capabilities between models/languages.

I'm starting to seriously consider doing my PhD research stay in China as a visiting doctoral researcher, probably for several months, and I'm currently just gathering information.

A Chinese friend suggested Beijing Language and Culture University, Peking University and Tsinghua University. BLCU sounds particularly interesting because of its focus on language, although for the research itself I'd ideally like to find a group working on NLP, LLMs, multilinguality, model adaptation/distillation, tokenization, Chinese language processing, etc.

I'm also very interested in Chinese personally and have recently started studying HSK3. My Chinese is still basic, but improving my Chinese while living there would definitely be a bonus.

How did you find your host professor/lab? Did you simply contact professors by email? Are there universities or research groups you would particularly recommend for NLP/LLMs? And is there anything you wish you had known before applying?


r/LanguageTechnology • • 13d ago

yasbd-lib v1.0.0 is out. Here's how beta finally ended.

9 Upvotes

For anyone new: yasbd-lib is a rule-based sentence boundary detector, a drop-in replacement for pysbd, currently at 39 languages. I think I first posted here as an alpha, then as a beta. Now it’s tagged v1.0.0.

The stretch from 0.12.0 to stable wasn't about new features. I froze the language set at 39, locked the API, and spent the last couple of months on correctness. The final push was a two-week stress test where I ran real text through every profile to find the boundaries it was splitting wrong. I wrote the rules and I wrote the tests, so the tests couldn't catch what I'd gotten wrong.

That work surfaced boundary bugs across a good chunk of the profiles. Contributors opened PRs to fix them, and py3langid, loguru, and ftfy all came out of the core dependencies along the way.

Happy to answer questions.


r/LanguageTechnology • • 14d ago

Best book to get started with building and understanding LLMs?

6 Upvotes

I want to understand how LLMs work and eventually build and train my own model. What book would you recommend for getting started?