Research
Observing data and models

Natural language processing (NLP) is the field that studies how computers can understand and process the natural language we use every day. In recent years, the common approach in NLP has been end-to-end: prepare a dataset of many instances, and obtain a model that performs the task through machine learning (deep learning). However, this approach is known to often produce models that behave in ways their developers do not expect.
As an example, consider a task whose input is a pair of sentences, a premise sentence and a hypothesis sentence, and whose output is the semantic relation between them (entailment or contradiction). This task is called recognizing textual entailment. One dataset contained a bias that no researcher had noticed: when the hypothesis sentence contained the word “nobody”, the pair was a contradiction 98% of the time. Models trained on this biased dataset achieved high overall accuracy (over 80%), but internally they inferred the relation from the hypothesis sentence alone, without looking at the premise sentence (Tsuchiya 2018).
In other words, simply applying end-to-end methods without critically observing the data and the models is dangerous. For this reason, we are committed to starting from the definition of the task and the analysis of the data, creating data through annotation, and only then building models.
A common workflow
- Choose a task
- Choose an existing dataset
- Build a model
Our workflow
- Define the task and analyze the data
- Annotate (create data)
- Build models, and observe their behavior and internals
Communication during disasters

We examine, at scale, how the way place names are written in disaster-time social media posts (official names or abbreviations) relates to the quality of the information in those posts. We use LLMs to create training data and train classifiers on it.
Recent activity
- March 4, 2026PresentationPresented 2 papers at the 5th Annual Meeting of the Society for Computational Social Science of Japan (CSSJ2026)
Evidence-based judgment: evaluating LLM reliability
We evaluate whether LLMs correctly rely on the evidence they are given (such as Wikipedia citations or court decisions) rather than on their internal knowledge. This work grew out of our research on the "shelf life" of information, which automatically detects whether information on the web has become outdated. Carrying forward the question of whether information is still correct today, we work both on building test sets and on understanding how models process evidence internally.
- Award Hitachi, Ltd. Award, 21st Symposium of the Young Researcher Association for NLP Studies (YANS2026) (2026)
- Award Committee Special Award, 27th Annual Meeting of the Association for Natural Language Processing (2021)
Recent activity
- August 17, 2026PresentationPresented 2 papers at the 21st Symposium of the Young Researcher Association for NLP Studies (YANS2026)
Do LLMs really understand?
Using natural language inference, shogi positions, named entity recognition and code generation as subjects, we diagnose with linear probing and activation patching whether LLMs rely on surface patterns or capture structure. Our goal is to explain the causes of accuracy drops and errors from inside the model.
Recent activity
- March 10, 2026PresentationPresented at NLP2026
- December 2024PresentationPresented at PACLIC38
Finding style, emotion and stance inside language models

We discover directions in the internal representations of LLMs that correspond to attributes such as style, emotion and viewpoint, by combining sparse autoencoders (SAEs) with contrastive learning and in-context learning. Using Japanese social media posts and reviews, we aim to separate an author's habits, the type and intensity of emotions, and the aspect and polarity of evaluations.
Multilingual and multicultural language models, and social media analysis
Building on our analysis of how information is exchanged on social media, we have extended our work beyond Japanese and English. In Thai, Indonesian, Finnish and other languages, we study emotional expressions in social media posts and news headlines, and bias in the cultural knowledge of LLMs. Combining trends in mention counts with emotion analysis reveals, for example, that a topic can be mentioned more and more while the content is negative.
Recent activity