Projects
Projects
PubMed Landscape
A 2D atlas of 21M+ biomedical abstracts built from text embeddings, used to study the emergence of research fields and the dynamics of COVID-19 research.
LLM Excess Vocabulary
Detecting the fingerprints of LLM-assisted writing across 15M+ biomedical abstracts by tracking shifts in word usage over time.
Text Visualizations
An SBERT-based projector trained end-to-end with a contrastive, t-SNE-like InfoNCE loss, producing 2D text embeddings that double as both a good visualization and a good sentence embedding.
Cropping vs. Dropout: Text Embedding Augmentation
Companion code for our TMLR paper comparing self-supervised fine-tuning strategies for text embedding models, showing that cropping-based augmentation outperforms dropout.
LLM Readability
Quantifying how LLM-assisted writing has changed the readability of biomedical abstracts published between 2010 and 2025.
ICLR Dataset
An open dataset of 24k+ ICLR submission abstracts (2017-2024) with decisions and keyword-based labels, used to study trends in machine learning research.