Lightnews — Scholar-powered news

Matthew Leavitt

@leavittron.bsky.social

Also we're a small startup and we want to remain nimble. Roles here are fairly fluid, though there are specific strength areas that we are trying to hire for, hence our distinct job postings.

November 25, 2024 at 7:40 PM

Matthew Leavitt

@leavittron.bsky.social

I know you're not trying to call me out, but happy to give my thoughts: Many research orgs have a "research vs. eng" divide that comes along w/ baggage: hierarchies, expectation of duties, etc. We don't want that here. Nobody is too good to touch code or insufficiently credentialed to do science

November 25, 2024 at 7:40 PM

Matthew Leavitt

@leavittron.bsky.social

Huge shoutout to @agcrnz.bsky.social @alvin-d.bsky.social @pratyushmaini.bsky.social and Mo Razzak for leading this work. You did an amazing job! Stay tuned for more announcements from us. We’ll have a booth at NeurIPS, come say hi!

November 25, 2024 at 5:49 PM

Matthew Leavitt

@leavittron.bsky.social

If you’re interested in pushing the bounds of what’s possible with data curation, we’re also looking for talented Members of Technical Staff who have experience doing data research, translating science into products, and building scalable data products
jobs.ashbyhq.com/DatologyAI

DatologyAI Jobs

jobs.ashbyhq.com

November 25, 2024 at 5:49 PM

Matthew Leavitt

@leavittron.bsky.social

We’re starting to work with early customers: if you’re an enterprise AI company interested in training multimodal and/or text models faster, better, or smaller, get in touch!

November 25, 2024 at 5:49 PM

Matthew Leavitt

@leavittron.bsky.social

If you'd prefer a quick overview, we have one of those, too!
www.datologyai.com/post/train-l...

Train LLMs Faster, Better, and Smaller with DatologyAI’s Data Curation

DatologyAI's curated data delivers substantial improvements in LLM quality, training speed, and inference efficiency over existing datasets.

www.datologyai.com

November 25, 2024 at 5:49 PM

Matthew Leavitt

@leavittron.bsky.social

If you want more details, here’s the full technical deep-dive!
www.datologyai.com/post/technic...

Technical Deep-Dive: Curating Our Way to a State-of-the-Art Text Dataset

Our data curation pipeline to obtain substantial improvements in LLM quality, training speed, and inference efficiency.

www.datologyai.com

November 25, 2024 at 5:49 PM

Matthew Leavitt

@leavittron.bsky.social

Overall I’m thrilled with these results. And I’m so very proud of our team for the amazing work that got us here. But the results aren’t the goal. The results are the first proof that it’s possible to build a product for foundation-scale data curation.

November 25, 2024 at 5:49 PM

Matthew Leavitt

@leavittron.bsky.social

We can also use our data curation to train better, smaller models to save on inference: a 1.3B model trained on 180B tokens of our data has better 5-shot performance than every 2.7B model we trained on public data sets, on token-matched (NOT FLOPs-matched) basis. FLOPs-matched is even better

November 25, 2024 at 5:49 PM

Matthew Leavitt

@leavittron.bsky.social

Our curated data also allows us to train faster! We save 86.9% on compute (7.7x speedup) training a 2.7B model on our data to reach the same avg 5-shot accuracy as training on RPJv1 for 180B tokens, and save 70.1% on compute (3.4x speedup) to reach the same accuracy as DCLM

November 25, 2024 at 5:49 PM

Matthew Leavitt

@leavittron.bsky.social

This is noteworthy because FW-Edu and DCLM have pool sizes that are 10x (DCLM) and 11.5x (FW-Edu) the curated dataset size. Our 180B token dataset is curated from a pool size of 540B tokens, which is only 3x. So we probably have a lot of room for improvement with larger datasets!

November 25, 2024 at 5:49 PM

Matthew Leavitt

@leavittron.bsky.social

Interestingly, we also find that starting with a larger dataset to curate yields a much better final dataset.

November 25, 2024 at 5:49 PM

Matthew Leavitt

@leavittron.bsky.social

Our improved model quality is general—it doesn’t come from outsize gains on a small number of tasks. We tie or surpass even the strongest baseline, DCLM, in two thirds or more of the evaluations, and are at par or outperforming other baselines on nearly all evals.

November 25, 2024 at 5:49 PM

Matthew Leavitt

@leavittron.bsky.social

With our curated data we were able to train better models: 8.4 percentage-point (pp) mean 5-shot improvement over RPJv1, +6.1pp vs FineWeb-Edu (FW-Edu), and +4.4pp vs DCLM. This is no small feat: FineWeb, FineWeb-Edu, and DCLM are VERY high-quality, meticulously-curated datasets

November 25, 2024 at 5:49 PM

Matthew Leavitt

@leavittron.bsky.social

Then we trained standard (MPT-style) transformers up to 2.7B parameters for token budgets up to 180B on our curated RPJv1 and other public pretraining corpora, and evaluated the models on a suite of 15 standard language model evals

November 25, 2024 at 5:49 PM

Matthew Leavitt

@leavittron.bsky.social

Why did we choose to curate RPJv1? Because it’s well-established, contains diverse content across a number of domains, and already has a moderate degree of curation applied to it

November 25, 2024 at 5:49 PM

Matthew Leavitt

@leavittron.bsky.social

Our data curation pipeline is a scalable, productionized system that integrates a suite of bleeding-edge algorithms to curate data in the quantity necessary for foundation model pretraining. And with it, we developed a single recipe that we used to to curate RPJv1

November 25, 2024 at 5:49 PM

Matthew Leavitt

@leavittron.bsky.social

tl;dr: We transformed RedPajama-v1 (RPJv1) into a dataset that outperforms FineWeb-Edu and DCLM, two of the strongest publicly-available text pretraining datasets. Let me walk you through how we did it

November 25, 2024 at 5:49 PM

Matthew Leavitt

@leavittron.bsky.social

Some of you may have seen our recent announcement of our state-of-the-art data curation pipeline and the fantastic results we got applying it to multimodal data for training CLIP models. Well it works pretty well for text, too!
bsky.app/profile/leav...

Matthew Leavitt @leavittron.bsky.social · Nov 14

🧵We’ve spent the last few months at @datologyai.bsky.social
building a state-of-the-art data curation pipeline and I’m SO excited to share our first results: we curated image-text pretraining data and massively improved CLIP model quality, training speed, and inference efficiency 🔥🔥🔥

November 25, 2024 at 5:49 PM

Matthew Leavitt

@leavittron.bsky.social

HUGE shoutout to Haoli Yin, Amro Abbas, and (Evil) Josh Wills for leading this work. You did an amazing job! Oh, and stay tuned for more announcements from us. Our curation pipeline works for text, too 😉

November 14, 2024 at 5:16 PM

Matthew Leavitt

@leavittron.bsky.social

If you’re interested in pushing the bounds of what’s possible with data curation, we’re also looking for talented Members of Technical Staff who have experience doing data research, translating science into products, and building scalable data products jobs.ashbyhq.com/DatologyAI

DatologyAI Jobs

jobs.ashbyhq.com

November 14, 2024 at 5:16 PM

Matthew Leavitt

@leavittron.bsky.social

We’re starting to work with early customers: if you’re an enterprise AI company interested in training multimodal and/or text models faster, better, or smaller, get in touch! forms.wix.com/f/7257903640...