About me
I am a tenured researcher at Inria Lille working on natural language processing and language models.
My research focuses on reasoning, evaluation and data. I study how language models acquire and apply symbolic reasoning capabilities, how these capabilities can be evaluated reliably, and how procedural environments and synthetic data can support model training. I also work on datasets, benchmarks, data provenance and open research infrastructure.
Selected publications and projects
- Reasoning Core: a scalable procedural data generation suite for symbolic pre-training, post-training, evaluation and reinforcement learning. Software
- Humanity’s Last Exam: an expert-level benchmark for assessing advanced AI capabilities, published in Nature.
- Logic Haystacks: a long-context evaluation of logical reasoning without easily identifiable unrelated padding.
- Bridging the Data Provenance Gap Across Text, Speech, and Video: a multimodal extension of the Data Provenance Initiative, published at ICLR 2025.
- Scaling Synthetic Logical Reasoning Datasets with Context-Sensitive Declarative Grammars: the paper introducing the approach implemented in gramforge.
- tasksource: a structured collection of hundreds of interoperable NLP tasks. Software
I also train encoder models for natural language inference and zero-shot classification. ModernBERT-base-nli remains state of the art on several NLI and reasoning benchmarks, while supporting long contexts and efficient inference.
See the research, publications and software pages for more.
