Software and data

I develop open tools, datasets and models for reasoning, evaluation and reusable NLP research.

Reasoning Core

Reasoning Core generates verifiable textual tasks for language-model pre-training, post-training, evaluation and reinforcement learning. It covers formal logic, mathematics, planning, algorithms, syntax and related symbolic domains, with task-native scorers and more than 10 billion tokens of pre-generated data.

NLI and zero-shot encoder models

ModernBERT-base-nli is an efficient encoder trained across a broad collection of natural language inference, logical reasoning and zero-shot classification tasks. It remains state of the art on several NLI and reasoning benchmarks and supports long-context inputs.

gramforge

gramforge is a Python library for synthetic data generation with declarative context-sensitive grammars. It supports parallel generation in multiple representations, depth constraints and inspectable abstract syntax trees.

tasksource

tasksource provides hundreds of curated and standardized NLP datasets with reusable preprocessing functions. It supports large-scale evaluation, multi-task learning, instruction-data construction and NLI model training.

Data Provenance Initiative

The Data Provenance Initiative develops datasets, tools and analyses for understanding the origins, licensing, attribution and evolution of training data used in AI.

MindGames

MindGames is a dynamically generated benchmark based on epistemic modal logic for evaluating theory-of-mind reasoning in language models.