TryShadowing
Shadow YouTube. Speak English.
Home
Browse
Dictation
NEW
My Library
en
0
days
Sign in
Inference, Diffusion, World Mode… — Y Combinator shadowing | TryShadowing
TryShadowing
Shadow YouTube. Speak English.
Home
Browse
Dictation
NEW
My Library
en
0
days
Sign in
Home
Browse
Y Combinator
Inference, Diffusion, World Models, and More | YC Paper Club
Inference, Diffusion, World Models, and More | YC Paper Club
Y Combinator
·
1:07:18 · May 28, 2026
Start Shadowing
0:00
0:00
Record
×1
1x
VI
EN
JA
KO
ZH
FR
PT
TH
IT
DE
IPA
Pronunciation scoring isn't supported on this browser — you can still record and listen back.
All
right.
Translating…
Turn on Record to capture your voice and get scored
Smart
Karaoke
Line
1
/1501
0:07
All right.
0:08
Hello everyone.
0:11
How you guys doing?
0:12
Welcome to the first ever YC paper club.
0:16
This is like a very exciting thing.
0:21
Absolutely thrilled with the response.
0:23
We had over a thousand folks that applied to come in.
0:26
It was a very hard selection.
0:27
If you guys have friends that didn't make the cut, I'm very sorry.
0:30
We're we kind of we need to keep it to about a hundred.
0:33
Um and so we selected a very very cool group.
0:37
Um the mission is to create this kind of community of great founders
0:44
and great researchers and try to pull them together.
0:47
I guess just for you guys to get a sense for how cool the
0:50
people in this room are.
0:52
Um, raise your hand if you have at least five citations, 10 citations,
1:02
a 100 citations, a thousand citations.
1:08
Wow, this is insane.
1:09
Okay, 10,000 citations.
1:12
Oh my god.
1:12
Okay.
1:13
All right.
1:14
This is awesome.
1:15
I I would go up to 300,000,
1:16
but I think it's like Chris Manning and that's about it.
1:18
Um, so, uh, raise your hand if you've raised at least a million dollars.
1:24
Raise your hand if you've re raised at least $5 million.
1:28
At least $10 million, at least $50 million.
1:35
We still got one.
1:35
We still got two over here.
1:37
All right.
1:38
Okay.
1:39
Awesome.
1:39
The hidden mission that I'll also kind of add on this is we had
1:43
uh Har and I had um this uh awesome uh breakfast in uh Woodside
1:48
and this place is
1:50
so so unique and special
1:52
and we kind of just don't use it enough at YC.
1:54
So the hidden mission is to make Pioneer great again.
1:57
And so I went through winter 16 here.
1:59
Um it was an unbelievable time.
2:02
I think 140 companies went through
2:04
that batch. 10 of 15 of them are unicorns.
2:07
It's an insane number. um WPY, uh Astronis, um Deep Graham,
2:12
all these companies were in the batch
2:14
and during that time uh Sam was still running the show
2:17
and basically sitting right there would be me,
2:21
Undercarpathy, Vaj Deremba and Greg Brockman
2:24
because they were starting this thing called OpenAI
2:26
and it was like the very early stages
2:28
and there was like not
2:29
that many AI companies.
2:31
So they would ask me
2:32
and Steve from Debb like what are you guys what are you working on?
2:35
What are the problems you're working on?
2:36
and they're looking for problems because they didn't even know what to research.
2:38
And so it was such a such a special time.
2:40
This place is so special uh to to me in particular uh to Har
2:45
as well.
2:45
And we just it's it we don't really use it enough.
2:48
So I wanted um to kind of make this community down here.
2:51
And I also think
2:52
that 100% of the AI talent
2:55
or AI people in the Bay Area,
2:58
probably about half of them are in the city maybe is a good number.
3:01
There's anthropic, uh there's open AI, there's cursor,
3:04
there's all this stuff in the city.
3:05
Then there's a lot
3:05
that are down here
3:06
that are not making the trek up to the city to join YC.
3:09
And so he's like, "Yes, emphatically, yes."
3:12
Um, and so you have Google DeepMind right on the corner.
3:14
You have um Tesla, you have XAI, you have Thinking Machines,
3:17
you have all these other people in Palo Alto,
3:19
you have a lot of startups.
3:20
And so uh I wanted to kind of like solve six birds with one
3:24
stone and kind of pull together this community down here
3:27
as well.
3:27
And Harj uh uh is super excited about it as well.
3:30
And so thank you very much Har for letting us do this.
3:32
We got uh five great papers here coming up.
3:35
The first one is Tanishk Speculative Speculative Decoding.
3:39
You want to come up?
3:42
All right.
3:45
Do you want me to pull it on?
3:46
Yeah, I got you.
3:51
Cool.
3:51
I know it uh looks like maybe I was sloppy
3:53
and I added an extra word in the title,
3:54
but uh it is intentional um and it'll make sense in uh good time.
3:59
Um my name is Tanishk.
4:00
I'm a grad student at Stanford.
4:01
Um, this is a project I worked on with Triau and Aar May.
4:05
I'm going to be evangelizing inference for people today.
4:09
Hopefully, you'll be inference enjoyers by the end.
4:12
So, I'm not sure how much I have to motivate inference.
4:16
I worked on training before inference.
4:18
And I sort of the sort of mental model I had in mind for
4:20
how inference works was you know you do this beautiful craftsmanship during the training
4:25
process and you get these like you know very intricate weights
4:28
and then you kind of just hand it off
4:30
and use them to generate tokens.
4:31
In my mind it's sort of like you have the weights just multiply the
4:35
matrices it's why do you need a team for it?
4:38
Um I was very confused
4:40
but there is in fact a lot of subtlety involved.
4:43
Um it's a lot of fun the algorithms and systems behind inference at scale.
4:47
I'm not sure I need to spend too long talking about why inference is
4:51
important.
4:52
Um there is one point I want to make
4:54
that I don't hear people talk about enough.
4:56
So things you may have heard are that inference costs are high.
5:01
They dominate training costs
5:02
when you're serving a model for billions of users
5:05
or you know 10 claud code power users.
5:09
That's trillions of tokens.
5:11
Um, not only are inference costs dominating training costs, but even within training,
5:17
RL is starting to exceed the compute requirements of pre-training.
5:21
And what is RL but a wrapper on inference, right?
5:25
So, these are two things you've probably heard before.
5:28
The third is one I fear isn't really talked about,
5:31
but it's the reason that I started working on inference,
5:34
and I use the phrase working on inference lightly.
5:37
This was the only inference project I've ever done.
5:39
Um, but the the reason I got interested in making inference fast was not
5:43
because of cost or for convenience.
5:45
It was entirely because of capability.
5:48
So the claim I'm going to make
5:49
and maybe this is the one thing to take away from the message I'm
5:52
trying to send in this talk is
5:54
that inference today is seen
5:56
as a sort of like cost
5:58
or convenience lever.
6:00
But uh in one two
6:01
or 3 years inference is going to be seen
6:03
as a capability.
6:05
And what I mean by that is that if you have a method,
6:08
an algorithm, a system where its performance scales with the amount of thinking it
6:13
does,
6:14
then fundamentally the speed at which you can do inference,
6:17
the tokens per second is exactly the peak intelligence that you can deliver.
6:23
So inference should be thought of
6:24
as not so much
6:25
as a a cost
6:26
or or convenience factor,
6:27
but as a capability.
6:29
Um, and that's why I got interested in it.
6:30
I I wanted to work towards the future where we have an entire data
6:34
data center of 20,000 B200s just working on the reman hypothesis.
6:38
Um okay, yes, that's the future that uh I had in mind.
6:43
Perhaps this meme is a little outdated because it has an A100 on it,
6:46
but uh yeah.
6:48
Okay.
6:49
So to motivate things, here is an example of fast inference.
6:54
So I'm going to give you a little demo of uh three algorithms side
6:57
by side.
6:57
We're going to sample, you know,
6:59
a code prompt from VLM with just normal auto reggressive decoding.
7:03
We're going to use their speculative decoding.
7:05
And then I'm going to put next to it the sort of janky handrolled
7:09
inference engine I wrote over a summer for this project.
7:11
Um, whose main strength is just
7:13
that it implements a new algorithm
7:15
and so you can see them side by side.
7:17
SSDs on the right
7:18
and you can see it is quite a bit faster than what you can
7:21
get if you try to use an open source engine.
7:23
Um, and it's not the systems, it's it's the algorithm.
7:26
Um so yeah that's what we want to work towards understanding both how speculative
7:30
decoding works as well
7:31
as the algorithm on the right.
7:34
Okay.
7:35
Um I'll start by introducing what speculative decoding is how it works
7:39
and then we'll move into what speculative speculative decoding is.
7:43
I hope that if you have like a reasonably strong understanding of how speculative
7:47
decoding works the the problem
7:49
that SSD is trying to solve will feel very motivated
7:51
and and the algorithm should just become clear in good time.
7:56
Okay, so this is the schematic I'm going to use to explain how vanilla
7:59
speculative decoding works.
8:01
Um, it has a small model, the tiny llama up top,
8:04
as well as a big model, the big llama.
8:06
And our goal is simply to sample fast from the big llama.
8:11
We want tokens generated from the big model.
8:12
And we're going to use a small model
8:13
as a sort of proxy
8:14
or an instrument to be able to sample quickly from the big model.
8:17
Okay.
8:18
So, what the draft is going to be responsible for is basically generating a
8:22
bunch of tokens one by one.
8:24
One by one is important.
8:25
It's auto reggressive.
8:26
So you need to do three forward passes on the draft
8:28
or you know however many some constant number.
8:30
Um and these are going to be guesses for what the draft believes
8:34
that the big model is going to output next.
8:37
It wants to sort of predict ahead of time.
8:38
The job that the big model has,
8:41
I'm going to call it the target model, is verifying these guesses.
8:44
What does verification mean?
8:46
Verification means doing one forward pass over these generated tokens to see how likely
8:52
it is that the big model would have generated them.
8:55
The sort of key asymmetry here,
8:57
the reason that speculation works is
8:59
that it is easier to verify than to generate.
9:04
This is a feature of the transformer architecture where you can get the probabilities
9:07
for many tokens in a sequence in parallel in one forward pass.
9:10
Um but you can't generate them in parallel. auto reggressive decoding
9:14
as uh one at a time.
9:16
Um so we're leaving the auto reggressive decoding
9:18
which is slow uh to a very quick
9:21
and small model and
9:22
then we're doing just one forward pass on these tokens.
9:25
And the way you verify tokens is basically by having the big model look
9:29
at the probabilities of each of the generated tokens
9:31
and see how plausible it is
9:33
that it would have generated those tokens.
9:36
And sort of the intuition here is
9:37
that we will accept precisely those tokens
9:40
that the big model could plausibly have generated.
9:43
Its probabilities were reasonably high.
9:45
There subtleties in exactly what the algorithm is um
9:47
that I'm going to gloss over,
9:48
but that's the way to think about it.
9:49
Um and then we're going to find a point perhaps where we don't think
9:52
it's plausible the big model would have generated those tokens
9:54
and we're going to reject those tokens.
9:56
So in the little schematic on the right uh there the draft samples three
10:01
and the big model verifies them
10:02
and concludes that only the first token was something it would plausibly have generated.
10:06
It will reject the second token onwards
10:08
and importantly this is a sort of critical
10:11
but subtle detail of vanilla specular decoding
10:14
because you have the probabilities at each of the sequence positions.
10:17
You can sample an extra token at the point at
10:20
which you rejected a token for free
10:23
as in without doing any more forward passes.
10:25
And so that yellow token is what I'm going to call a bonus token
10:28
that you sample for free.
10:29
This is going to be important in SSD.
10:31
Um, so yeah, that's uh that's an important conceptual point.
10:37
And this sort of sets the stage for how SSD works.
10:42
Okay, we have our schematic.
10:45
And the way we've set up speculative decoding is
10:48
that it's a way to exchange flops for latency.
10:50
So speculation in general is not actually something that uh only LLMs do.
10:55
It's like a a deep idea in computer science.
10:57
It's used in CPUs
10:58
as well where the general philosophy is
11:00
that you premputee something ahead of time.
11:02
Some of what you premputee may be useless
11:05
because it may be an incorrect prediction of the future,
11:07
but if you're right,
11:08
you get to fast forward in time um
11:10
and you get lower latency
11:11
as a result.
11:12
So the the sort of like moral philosophy of speculative decoding is
11:14
that it's currency exchange.
11:16
The difficulty with normal speculative decoding is that you can't push this arbitrarily far.
11:22
You cannot keep sampling more
11:23
and more tokens on the draft
11:24
and keep getting speed ups
11:26
because at some point you're going to get to a point where you're spending
11:28
a lot of time drafting
11:29
and you're not accepting all
11:30
that many tokens.
11:31
And in particular, like a big bottleneck in vanilla speculative decoding is the sequential
11:35
dependence between the small llama
11:36
and the big llama.
11:37
Um the drafting in round t has to take place before the verification of
11:42
those tokens. um and the drafting in round t+1 can't take place before you
11:47
know the outcome of verification of the previous round
11:49
because you need that
11:50
as a prefix to draft on top of.
11:52
So there's a logical dependency here.
11:55
The goal of SSD is very simple.
11:58
There's a lot of gnarly
11:59
and subtle details but the highle idea is incredibly simple.
12:02
It is simply to parallelize this sequential operation.
12:06
We want drafting and verification to be happening at the same time.
12:12
Normally in speculation they happen on the same hardware
12:15
and that's fine because there's only one of them happening at a time.
12:18
In our setup they're going to be happening at the same time.
12:20
So we're not going to be collocating them.
12:22
And the main question basically becomes how do you parallelize this inherently sequential algorithm
12:28
that has a logical dependency.
12:30
Um and the way we're going to do
12:31
that is we are going to have the draft model send back its draft
12:35
tokens in a certain round.
12:37
So we've sent back a bunch of blue tokens.
12:39
That's now the job of the verifier to do a forward passover and verify.
12:44
And this is going to take a
12:45
while because a verifier is a big model.
12:47
What we on the draft are going to do is basically start anticipating the
12:51
most likely verification outcomes immediately.
12:55
As soon as we send back like a certain round of speculation
12:58
and once we we have in mind some of the most likely verification outcomes,
13:02
we are going to start drafting the next round on top of those immediately
13:06
while verification is taking place.
13:08
If we're right, the next time the verifier asks for a draft,
13:12
we'll have it ready immediately.
13:14
We're entirely hiding the latency of drafting.
13:16
If we're wrong, well, we'll have to figure out a backup strategy.
13:18
And there's uh there's there's there's some subtleties on what you do
13:21
and how you do it there.
13:23
Um so yeah, the way that speculative decoding looks like this.
13:26
And perhaps unsurprisingly, the analog for SSD is this diagram on the right.
13:32
We're now drafting and verification happen in parallel. um the the principal difficulty
13:38
or algorithmic design space in SSD is how do you predict verification outcomes ahead
13:43
of time.
13:44
I thought verification is where you are leveraging the intelligence of the big model
13:48
that should by construction be difficult to predict.
13:50
Um and the intuition for why it's plausible at all is
13:53
that you can make many guesses on the draft for what a verification outcome
13:57
is.
13:57
And a verification outcome here is just you know a plausible number of accepted
14:01
tokens and then a bonus token on top of
14:04
that.
14:05
Now this is hard to predict
14:06
because a bonus token comes from a vocabulary
14:08
which has size you know tens to hundreds of thousands.
14:10
Um so it's a large space to cover um
14:13
but it turns out you can do it well um reasonably well.
14:16
You can get it right about 80 to 90% of the time
14:18
which is more than enough to get big speed ups.
14:20
And the way we do that,
14:21
the short of it is basically we use information on the draft to predict
14:25
what the verification outcome is likely to be.
14:27
When we generated the blue tokens on the draft,
14:29
we had other tokens that we chose not to sample.
14:31
Those other tokens are plausible verification bonus token candidates.
14:35
And so you basically use information from the token distributions of the draft model
14:40
to predict what likely outcomes on the target are.
14:43
And then once you have all of these predictions,
14:45
you can decode them in parallel
14:46
as just different sequences
14:47
that you're decoding on top of a shared prefix.
14:50
And voila, it uh it's it gives you speedups
14:54
because you get to hide the latency of drafting altogether.
14:57
Um there's also a an additional bonus
15:00
that since verification actually kind of takes a
15:02
while,
15:02
you get more time to draft uh in the first place.
15:05
So you can draft more tokens
15:06
which increases the expected tokens per round
15:08
and sort of gives you further speed ups.
15:11
There's a bunch of stuff
15:12
that we work through in the paper that's uh that's sort of reckoning with
15:15
the the implementation details of this.
15:18
One of it is how you handle cache misses.
15:20
One plausible thing you could do perhaps naively is to just fall back to
15:23
ordinary speculation just in time.
15:25
Turns out that actually this is not always optimal.
15:27
Um there's trade-offs.
15:29
You know, as batch size increases,
15:30
you're going to fail to predict some of the sequences verification outcomes.
15:34
Um and so you need different ways to predict and handle cache misses.
15:38
Should you be allocating your compute on the draft equally amongst plausible prefix length?
15:45
Uh the short answer is no.
15:46
You can be clever about it.
15:48
And all of this trickery just helps you increase your cash hit rate,
15:52
so to speak, the amount of time you're able to correctly predict verification outcomes.
15:56
And there's there's some trade-offs between cash hit rate
15:59
and the actual quality of the drafting you're doing.
16:02
Um and this is totally non-obvious.
16:04
Um, and and and we we go into why
16:06
that exists and how you can navigate it in the paper.
16:09
Um, I'm happy to talk about it in in in Q&A as well.
16:11
Um, okay.
16:13
So, what do you get for the the price of this uh mind-numbing complexity
16:21
and uh pain wrangling an inference engine?
16:23
Well, you get the privilege of watching a number go up,
16:27
which I guess is the north star of all AI research.
16:30
And so here we have uh a bunch of inference algorithms and inference engines.
16:35
The blue ones are sort of uh my inference engine
16:38
and uh the light blue is just the baseline implementation of speculative decoding.
16:42
The red is SG lang
16:44
which is you know of all the inference engines we tried the fastest with
16:48
speculative decoding and the dark blue is is SSD.
16:51
Um and normally speculative decoding um is a is a win for latency
16:55
but it's sort of unclear whether it's useful for throughput. um for us it
16:59
turn in in in this setting it's actually a win for both um
17:02
and so you get numbers going up
17:03
and you also get the ability next time you are at a San Francisco
17:07
house party um to see other people dancing
17:10
and knowing in the corner
17:11
that uh you know what it takes to sample at 300 tokens per second
17:16
uh for llama 370B on 4H100s.
17:18
So this is uh sensitive information um but yeah that's that's about it.
17:23
YOU.
17:32
All right, that was awesome.
17:34
Okay, so for this next paper, this is um my first experience being scooped.
17:43
The only issue is
17:44
that he didn't talk to me
17:45
and he did it six months before me.
17:47
Um but uh Isaac can vouch for me on this
17:51
and maybe Robert as well.
17:53
I basically fell in love with the diffusion policy paper.
17:56
I was like this is definitely like you know a full uh predicting like
18:01
th horizon steps for your robotic control.
18:05
Um we have these amazing video models.
18:07
Why don't we just use the video model to like run this like at
18:10
test time to like play out the movie
18:13
and where do I end up?
18:14
And then you have your classic push t.
18:16
And then I started like looking around uh
18:19
and then DM mind of course already did it.
18:21
So so I wasted like a month and it was not happy.
18:25
But anyway, thank you very much.
18:26
Please welcome Stannis.
18:33
>> Hi everyone.
18:34
I'm Stannis.
18:35
I'm a star research scientist at Google DeepMind.
18:38
Uh currently I'm co-leading a new project on word modeling for robotics. uh where
18:42
we try to build general purpose policies on top of video
18:45
and word models.
18:46
But uh this is an early work that I did about two years ago.
18:51
Uh so this is before I switched to working on hardcore robotics
18:55
and uh going into hardware really scaling up the data
18:58
but uh you can probably see a lot of very similar ideas early version
19:03
of ideas demonstrated on toy problems.
19:06
Okay.
19:07
So uh first to give some background what is the model predictive control.
19:12
So model predictive control also called the receding horizon control uses a dynamics model
19:17
or some people also call it a word model
19:19
and uh action selector mechanism uh
19:22
which is a planner to construct agents
19:24
that can solve a wide variety of tasks by means of maximizing a no
19:29
objective.
19:30
So the main advantages of model predictive control is uh it can adapt to
19:35
normal reward functions at test time.
19:37
So uh the dynamics model are also easier to learn
19:40
and generates better than just policies
19:43
and the action proposal dynamics model factorization also allows easy adaptation to normal dynamics.
19:50
So we're going to uh demonstrate some of these in later experiments
19:54
but basically here we are showing the overall idea
19:56
which is extremely simple.
19:58
We have a action proposal which proposes a sequence of actions.
20:02
We have a dynamics model
20:03
which can evolve these actions
20:05
and give you the future states.
20:07
And uh finally we have some objective functions that we are trying to optimize.
20:11
We basically use a planner to optimize
20:13
that and uh pick the actions
20:15
and execute it in the environment.
20:17
So what is diffusion model operative control?
20:20
So the motivation mainly is uh uh there are a couple of problems we
20:25
need to address in order to make MPC effective in practice.
20:28
One the dynamics model need to be accurate to avoid the problem of compounding
20:32
errors and uh two the planning algorithm also needs to be powerful enough to
20:37
select a good sequence of actions.
20:39
So with DMPC what we did is to use diffusion models to learn both
20:44
multi-step action proposals and multi-step uh dynamics models.
20:49
So the advantages are mainly to reduce compounding errors
20:53
and we also found
20:54
that uh it can simplify the planning algorithm.
20:57
Essentially we can just use a very simple uh sampling based planner
21:00
and we can already outperform a lot of the previous uh approaches.
21:05
So uh before we dive into the details also want to give a hierarchical
21:08
view of some related works we organized.
21:11
So there are a lot of related works in the literature
21:14
and uh we organize it uh uh in this way where we basically look
21:18
at how different approaches um
21:20
so basically all approaches essentially try to build a joint uh distribution of the
21:26
states and the actions
21:27
but they do it in different ways
21:29
and also use the different components in different ways.
21:32
So for example, you can build it in a factorized way where you have
21:36
row a which is your policy predicting the actions
21:39
and then collision on the action predict the state
21:42
which is a dynamics model
21:43
and uh for this you have the dynam paradigm where you basically learn a
21:47
model and use the model to also generate data in the imagination
21:52
and the learn policy.
21:53
But uh you can also do MPC uh where you uh essentially use a
21:58
planner to select the actions
22:00
and uh we also have uh some uh uh there are also approaches where
22:04
you build a joint model of the state
22:06
and actions and you're essentially also doing MPC
22:09
and there are also model free approaches where you directly learn a policy. uh
22:13
I won't dive into the full details
22:14
but uh uh there are basically different trade-offs in terms of runtime plan uh
22:19
whether we can do runtime planning
22:21
and uh adapting to normal rewards
22:23
and adapting to normal dynamics leveraging non-expert data
22:27
and also the uh general speed at runtime
22:31
and there is also the distinction between whether you're doing singlestep modeling
22:35
or multi-step modeling.
22:38
Okay.
22:38
So coming to diffusion model,
22:40
diffusion model has enjoyed a lot of successes uh in uh generating AI especially
22:45
for generating images and videos.
22:48
But uh in recent years they also found a lot of successes in robotics.
22:52
So currently uh so here I'm also showing a slide where uh this is
22:56
a kind of the exploration space for uh diffusion based uh I would calling
23:01
diffusion based agents.
23:02
So we of course start with the diffusion policy where we condition all the
23:07
observation and generate future actions.
23:09
But then we also have this work called the diffuser
23:12
which uh is uh you can think of it
23:16
as a way to joint jointly model uh observations
23:19
and states but in toy space.
23:22
There are of course these ideas are explored in tons of different papers
23:26
but this is just a very simple
23:28
and uh uh conceptual way to describe it.
23:31
And uh then there's also decision diffuser where we collision on the observations we
23:36
directly generate future uh we condition on the history directly generate future observations
23:41
and then try a separate inverse dynamics model to derive the actions
23:45
and uh finally we have the diffusion model predictive control where we first have
23:51
an action proposal to propose future actions
23:53
and use a dynamics model to evolve it
23:56
and uh then use planner to select the actions.
24:00
There are different uh trade-offs among these.
24:02
So for example, diffusion policy is sort of on complex uh complex control like
24:08
day-to-day we still rely on it a lot.
24:10
But this requires expert demonstrations.
24:13
So essentially you can't move out of the behavior cloning paradigm.
24:17
Uh for diffuser it's a jointly modeling state and action.
24:21
So it has implicit word modeling and also model based planning.
24:25
And this is actually something
24:27
that we are trying to explore at scale similar ideas.
24:30
But uh and then there's also uh decision diffuser where you do observation only
24:36
learning.
24:36
The main benefit of this is it allows you to leverage uh uh video
24:42
only data to learn from video only data
24:44
because for robotics uh the data is a many bottleneck.
24:48
And then finally there's a division MPC
24:50
which allows us to do runtime adaptation to normal rewards
24:54
and normal dynamics.
24:56
So what does the algorithm look like?
24:58
It actually is extremely simple.
25:00
We have uh often data set and uh we have uh some hyperparameters.
25:06
Essentially we are learning a couple of u uh learning a couple of models
25:11
all from the offline data sets.
25:13
We're learning a policy which u uh given the current observation predicts the actions.
25:18
We're learning a dynamics model
25:19
which uh given the uh given the actions uh evolves the observations to predict
25:26
the future states.
25:27
And uh uh basically after learning all this at uh um at uh inference
25:33
time when we actually deploy it
25:35
as a policy we uh sampled action proposal
25:38
and score it uh rank it
25:40
and uh pick the best.
25:42
But uh the main difference uh compared to previous approaches is uh we adopted
25:48
a multi-step action proposal
25:50
which uh is uh essentially very similar to a diffusion policy
25:54
but if you train on more diverse data it can give you uh more
25:57
coverage in terms of the action space
26:00
and uh we are also using a multi-step um uh dynamics model
26:06
which uh allows you to uh evolve for a long time horizon without a
26:11
lot of compounding error.
26:12
And uh this allows us uh to
26:15
and also uh there's a fact
26:18
that we leverage diffusion model
26:20
which is a really powerful way to model data especially multimodel data
26:25
and uh uh what we observed empirically is the uh stronger modeling uh capabilities
26:32
also allows us uh to uh simplify the planning algorithm
26:36
so that we can just use such a simple uh planner to do to
26:41
solve the task. tasks.
26:43
Yeah.
26:43
Um also contrasting with a few of the representative uh uh path works uh
26:48
including uh model based offline control offline planning
26:52
and this diffuser work
26:54
which I mentioned it learns a joint model
26:57
and uses a classifier free guidance for planning.
27:02
Okay.
27:02
Uh so yeah next to dive into some uh results uh there are lots
27:09
of numbers but the short answer is uh we obtain very competitive results in
27:14
fixed reward single task setups.
27:16
This is just to demonstrate
27:18
that uh uh the approach uh
27:20
when you deploy it in uh single reward uh fixed reward single task setup
27:25
it can perform competitively to the current state-of-the-art uh previous state-of-the-art approaches.
27:32
But uh I think uh there are a couple of uh more interesting uh
27:37
properties of DMPC.
27:39
One is it can adapt to no rewards at runtime.
27:42
Here we are showing some uh examples where uh essentially we train the model
27:48
to uh these are very simple modulo tasks
27:51
but we train the model to just uh local motion tasks run forward
27:55
and jump etc.
27:57
But uh at inference time we can just by changing the reward function to
28:02
uh make it uh exhibit uh novel behaviors like uh jumping etc.
28:08
So uh here's another example where we show
28:11
that uh uh DMPC can adapt to novel dynamics
28:15
while uh this kind of uh joint modeling approaches struggle.
28:19
This is really the benefit of the factorization of the action proposal
28:23
and the dynamics model.
28:25
So the here the idea is uh we can keep the action proposal the
28:29
same but uh we uh we have uh scenarios where the dynamics of the
28:35
environment changed.
28:36
So for example the walker has a broken left ankle
28:39
and as a result
28:40
when it starts to execute actions the consequence of the actions change.
28:45
So in such cases
28:46
because of the factorized representation in DMPC we can uh simply just adapt the
28:52
dynamics model on some play data collected in the new environment
28:57
and uh we observe
28:59
that we can recover a lot of the performance
29:02
because of the changing dynamics.
29:04
Finally, we dug into the various components of uh the DMPC design
29:10
and we demonstrated that uh the different components in DMPC basically contributed to improved
29:16
performance.
29:17
Uh this uh these include uh the diffusion active proposals, action proposals,
29:23
improve performance and simplify the planning.
29:26
We do multi-step diffusion action proposals
29:29
and the the fact
29:30
that we do multi-step also uh contributes to improved performance
29:34
and finally multi-step dynamics modeling also uh contributes to improved performance.
29:41
Uh that's it.
29:50
All right.
29:51
And that was the last Google Deep Mind paper that they're going to publish.
29:54
So, good luck out there.
29:56
Um, this next one is one of my lab mates
29:59
that I work with a lot
30:00
that is the most world model pled person
30:05
that I know.
30:08
And so, I can't imagine, you know,
30:10
anyone else presenting this paper other than Yan Lun himself.
30:14
Um, Isaac Ward.
30:18
There you go.
30:18
Thanks a lot.
30:22
>> All right, guys.
30:23
Is Is that a good distance?
30:24
You all can hear me at the back.
30:25
Cool.
30:26
Cool.
30:27
Yeah, I'm enjoying a uh a cool little period in life where I started
30:30
working on world models a couple years ago,
30:32
kind of before they got really hot
30:33
and now they're enjoying a moment in the sun
30:35
and suddenly everyone wants to talk to me
30:37
which is nice.
30:37
I'm presenting lay world model
30:39
which is a call out of course out of Yan Lacun's group.
30:41
Uh QR code here if you want to follow along with the project page,
30:44
but I'll explain through it and yeah,
30:46
really excited to talk to you about this one.
30:47
Uh hidden in this presentation is really like a billion-dollar question
30:50
and it's not hyperbole. uh Yan Lakun's raise of $1.03 billion dollars back in
30:55
March basically just to train world models is sort of what this presentation is
30:58
about.
30:58
I want to get at some of the questions
31:00
that they're going to be testing.
31:01
First five slides here just going to do some basics on world models.
31:04
I think we've all heard the term
31:05
but I want to just make sure we're all on the same page
31:07
and then we'll jump into uh what this paper is really uh offering
31:11
and what it means for world models at large.
31:13
But first of all, world models, what are they?
31:15
Why do we care about them?
31:16
So really it's about learning the dynamics of the world,
31:18
which is to say we're trying to come up with some model Typically,
31:22
we're using like a big neural network to predict how a system will change
31:24
over time based on its inputs.
31:26
So, you have your current state or scenario using S for notation here.
31:30
You're playing some action,
31:31
maybe that's like a movement or a command for a robot, um,
31:34
or a language command for a robot,
31:35
and then you're trying to predict like what its outcome is going to be,
31:38
like what scenario will it end up in once it's executed that action.
31:41
So, you're really trying to model the system
31:42
or the environment that the robot is in,
31:44
modeling the world.
31:45
It's a world model.
31:46
Uh, these kinds of models are really cool.
31:48
They enable a few really interesting capabilities.
31:50
One of them is generating imagined outcomes.
31:52
We've probably all seen like the sort of weird kind of um hallucinity uh
31:57
imagination sequences coming out of world models over the last couple years.
32:00
We'll talk more about those and why they're useful.
32:02
Uh this allows us to get to model based control.
32:05
I'm glad Stannis kind of explained that in the last talk for me,
32:07
so I'll skip over it.
32:08
Um and the last piece is really cool.
32:10
Surprise quantification.
32:11
Uh I'll get to that later.
32:12
Um but a really powerful capability of world models.
32:15
I wanted to communicate to you all
32:16
that this is not a new idea at all.
32:18
It's really just kind of new advertising or packaging on an old idea.
32:21
So I started going back through Google Scholar
32:23
and this is a paper
32:24
that I think is older than the average age of this room.
32:26
Um from Europe's 1990 and of course Richard S.
32:29
Sutton who we know from reinforcement learning basically describes exactly a modern world model
32:34
a black box that takes
32:35
as input its situation
32:36
and its action that it's going to execute
32:38
and outputs a prediction of its immediate next situation.
32:40
So really really old idea and uh that's the flyer from Europe's 1990.
32:45
Great.
32:45
Right.
32:46
So, getting a little bit more explicit um
32:47
and changing the notation from state to observation just
32:49
because in real world systems,
32:50
we typically don't have access to the exact true state.
32:52
We typically have some observation from sensors.
32:55
This is just an example
32:55
that I pulled up from some world models
32:57
that we're training on a quadrotor.
32:59
So, as an example,
33:00
the observation that the quadrotor gets might be its current kinematic state, position, velocity,
33:04
this kind of thing.
33:05
In addition to the images that it's taken from a forward- facing camera,
33:07
the action might be a control input, in this case a yaw,
33:10
and move back to the left.
33:11
And then we want to make a prediction
33:13
that says well if you do
33:13
that action you're going to end up slightly back in the room
33:16
and looking to the left.
33:17
And we actually want to generate what the sensor um would result uh in
33:21
in this case.
33:21
So highly uh dimensional observations images uh
33:24
and also LAR and things like
33:25
that are completely on the table in world models.
33:28
Uh they're really challenging because action sequences can be quite long.
33:31
Um and the really big thing is
33:32
that the minimum in the optimization landscape for these kinds of models may not
33:36
correspond to the desired behavior.
33:37
And more on that later.
33:38
Um, but hopefully you'll agree
33:40
that if you have trained a system that's capable of doing this thing,
33:42
it must have an internal model of the world.
33:44
And imbuing agents with an internal model of the world, um,
33:47
is potentially a very useful capability.
33:49
And that really is the big question.
33:51
Are we going to have model free or model based policies?
33:54
Are our agents going to have an internal model of the world
33:56
or are they not?
33:57
And this is sort of being fought out right now both in the research
33:59
community and in like the startup community.
34:02
So on the left, model free.
34:03
The idea is you're taking some observations,
34:05
you're feeding this into some kind of big neural network potentially with a bunch
34:09
of interesting learning tricks there,
34:10
but you're getting some optimal action out.
34:12
So, it's just mapping between observation and some optimal action.
34:15
But at no point is there an explicit representation of what the future might
34:18
look like if you execute
34:19
that action.
34:20
These kinds of models are pretty good.
34:22
There is growing evidence to show
34:23
that internal to these neural networks are highly obuscated
34:27
and challenging to interpret world models uh sort of in the in the weights.
34:31
uh I'll talk about a paper very briefly that's um speaks to
34:35
that and maybe someone can present on it in a future week.
34:37
And then over on the um other side, model based approaches, right?
34:40
So now we're saying we're going to train this world model up explicitly
34:42
and actually use that in our policy to be able to explicitly predict the
34:46
outcome of potential actions.
34:48
So yeah, totally like two different species of policies.
34:51
The model free stuff,
34:52
some of the weaknesses is they show a little bit of brittleleness to out
34:55
of distribution.
34:56
Um, model based ones are great
34:57
because you can kind of quantify modeling error
34:59
and this is really important
35:00
when you're deploying things in the real world.
35:02
Uh, we'll talk a little bit about this.
35:03
I have a little asterisk here, some biological precedent which we'll speak to more.
35:07
Um, and you have to have this additional mechanism of course
35:09
which is a downside where you actually need to propose action candidates to evaluate
35:12
with the world model um,
35:14
which Stannis spoke to in the previous talk.
35:16
This is a great paper.
35:17
But I just wanted to chuck this in there uh
35:19
which talks about how even model free base policies do have world models in
35:23
them and a really really cool paper
35:24
that hopefully can be presented in a future week.
35:27
Uh just to make it concrete before we jump into the paper I wanted
35:30
to just bring a little toy here just to show you what this looks
35:32
like.
35:33
So of course went to push t like all good researchers do
35:35
and in push t we basically just have an image of a little blue
35:38
ball agent and you're trying to push the blue tea into the green slot.
35:41
uh the state is comprised the observation is comprised of
35:44
that image plus the 2D position of the endeector
35:46
and the 2D action of where you're going to move the endector.
35:48
So you can make a little architecture that looks like this.
35:50
I just whipped this up.
35:51
Couple hundred thousand parameters and um oh let's play this.
35:56
So if that's the actual roll out,
35:58
this is what the model thinks the action sequence is going to do.
36:02
So you can see it's a little bit wobbly because it's a tiny model,
36:04
but we can certainly train up models of these kinds of toy environments
36:07
and indeed more complex ones.
36:09
So what are the challenges associated with training this kind of model?
36:11
Well, one is you're trying to learn the representation of the world.
36:14
So how you're going to compactly represent those highly dimensional images
36:18
or LAR inputs or highly dimensional sensor inputs at the same time
36:22
as you're trying to learn how actions change
36:24
that representation.
36:25
So you're co-learning representation and dynamics.
36:28
And there are many solutions in the optimization landscape
36:31
that will essentially just cause you to do nothing.
36:34
So for example a a local min minima in the optimization landscape is to
36:38
say well every state is just the same it's a trivial collapse basically um
36:42
and there are many techniques in the literature to say how can you avoid
36:45
these so there are solutions of a variety different kinds
36:49
that basically say there a way to avoid the collapse associated with training world
36:52
models and that's really where the world model comes in.
36:54
It says, well, instead of having to use some manner of trick
36:57
or like special method
36:58
or a bunch of like hyperparameter tuning schedule,
37:01
we're instead going to really drastically simplify this
37:03
and go for a more elegant method.
37:05
So, if you know a little bit about world models,
37:07
there's some popular ones in the top right here.
37:08
This is a figure straight out of the paper.
37:10
So, PLDM is planning in with latent dynamic models, dino, dino, um,
37:14
distillation with no labels, world model, dreamer out of deep mind,
37:17
and then temporal difference MPC as the final one.
37:20
So, in some way, shape or form,
37:22
I'll explain this. they use some kind of trick
37:24
or um like challenging to configure design to get away with uh this collapse
37:29
to avoid this collapse
37:30
and the world models coming in
37:31
and saying basically we can do this with sort of one hyperparameter
37:34
and one loss term
37:35
which I'll talk about there's really no time to go through all the different
37:38
tricks that different world model approaches use
37:41
because it really is the wild west out there right now
37:43
so many different methods
37:44
but they basically fall into one of these three categories
37:46
so one is you could do some explicit heristic
37:49
that stops collapse by like enforcing some special um healthiness in like the latent
37:54
space of your embeddings.
37:55
Um the language trick is maybe a bit unfair here,
37:58
but it's what's used in the paper.
37:59
Uh you could use some foundational methods.
38:01
So you could take some like existing autoenccoder
38:03
or diffusion model or video model
38:05
and use that as a basis for your world model
38:07
and add an action conditioning element in there.
38:10
Um or you could use some privilege data
38:12
that may not be usually available to the model outside of train time uh
38:15
to be able to avoid collapse.
38:17
and lay well model even
38:18
though it says that it's doing something very different I really think uh it's
38:21
just offering a new kind of trick uh
38:23
which I'll talk about here
38:24
so jer is joint embedding predictive architecture it's sort of yan lakun's main work
38:28
and lay world model is a kind of jepper model uh basically the way
38:31
it works is you're going to take an autoenccoder um
38:34
or I should say an image encoder uh encode this observation in this case
38:37
it's of a robot doing a push cube task that's going to turn
38:41
that image into a latent vector in the latent space of this encoder uh
38:45
you're going to train an action condition forecasting module this predictor to be able
38:48
to predict what is the next latent embedding going to look like
38:51
when I execute this action.
38:52
So not what the next image is going to look like
38:54
but what's the next latent going to look like
38:56
and you can use the decoder attached to
38:58
that encoder to decode
38:59
that back out into a useful image.
39:01
But for the most part all the interesting work is going to be done
39:03
in the latent space.
39:04
And basically what they say is over a batch all of those latent embeddings
39:08
uh should be in a healthy distribution
39:10
which they describe as a gausian distributed uh distribution in in the latent space
39:15
and thus enters the sigg regularizer
39:18
which is the sort of new term they add.
39:19
So sigg for sketching
39:21
as in uh doing one-dimensional passes over a high dimensional data.
39:25
Um I for isotropic
39:26
so this should look the same
39:27
when you slice it in any direction
39:28
and g for gaus
39:29
and distributed cigar.
39:30
So basically you're taking all of these embeddings of your different predictions doing a
39:35
one-dimensional slice over each direction like in
39:38
that highdimensional space and
39:39
then you want each of the curves across those slices to be gausian distributed
39:43
and if that's true
39:44
then your um distribution in the latent space must be very healthy.
39:48
Uh so the idea is you can quite cheaply evaluate how gausian distributed your
39:52
embeddings are and thus how healthy your world model is
39:54
and how non-olapsing it is.
39:56
So essentially I just say instead of training up on the normal predict the
39:59
next uh latent you add on this additional sigg term.
40:02
So I'd argue that basically this paper is just um providing a very elegant
40:06
kind of regularization.
40:07
And to finish off I'll just talk about three capabilities
40:09
that you get from this.
40:09
So one is the openloop prediction quality.
40:12
This is what world models do.
40:14
So you feed in like the context this push t at the top
40:16
and you can see the top row is the real example.
40:18
The bottom is the imagined and they look about the same.
40:20
This is good.
40:21
It means your world model is really good at predicting what your next action
40:24
is going to do.
40:24
They do that on push t
40:25
and then on a slightly um like a 3D analog task like a push
40:29
cube.
40:30
This is all great.
40:30
I love seeing these um these plots.
40:32
Um but really what matters is how does this actually affect the policy like
40:36
for the actual task completion.
40:38
How is this useful?
40:39
Um and that sort of brings us into how you can use these models
40:42
for model predictive control.
40:43
Basically you take your initial observation and a goal observation.
40:47
I put an asterisk there
40:48
because how often do you have a goal observation in a robotics task?
40:50
Like you don't always know exactly the situation
40:53
that you want to end up in.
40:54
But in this case, that's how they frame it.
40:55
So they say, you know, the world looks like this right now.
40:57
I want the world to look like this.
40:59
You encode both of those.
41:00
And then you're basically doing a search over the actions
41:03
that will get you in the latent space from this starting point to this
41:06
ending point.
41:07
And there are well- definfined optimization methods to um to achieve that.
41:10
It works pretty well.
41:11
I'll make it um make it simple.
41:13
The world model is better than the competition on these like small 2D tasks.
41:17
As soon as you go to 3D, Dino World model wins.
41:19
It does have a big foundational backbone trained on that kind of image data.
41:22
So you'd expect it to um to win.
41:25
Um they run on a really simple environment called two room
41:28
and kind of say you know we don't do
41:30
so well on this
41:30
but that's because we're promoting like really high dimensional healthy embeddings
41:34
and it's a very low dimensional problem.
41:36
I'm not sure if I'd truly go for that.
41:38
Um but a good takeway is
41:39
that it's about 50 times faster than any of the competition across the board
41:42
because it's doing all this work in the latent space
41:44
and it doesn't have to have any like additional tricks relating to more forward
41:47
passes or like having two copies of the model in memory.
41:50
And uh you can actually boot this thing up on like a single card,
41:52
less than 24 gigabytes of VRAM and it's only 15 million parameters.
41:56
So that is pretty nice.
41:57
Final piece, this is what I think is a really cool capability of world
42:01
models.
42:01
Um you can quantify the model error.
42:03
So basically they just come up with some trajectories
42:05
that kind of screw with the world model.
42:07
So the top one is going from left to right.
42:09
That's time.
42:09
Uh so that's just like a nominal example.
42:11
Everything's normal.
42:12
Then they take the same example, but they change the color of the tea.
42:15
And then they take the same example,
42:16
but they just teleport the tea into a different location.
42:19
And this is really cool
42:20
because you can actually see the moment they apply those perturbations,
42:23
you get a spike in the model error
42:24
and this is detectable
42:25
which is to say world model enabled agents can quantify how poor their predictions
42:29
are.
42:29
They have good estimates of their uncertainty.
42:32
This is really powerful.
42:33
Model freebased approaches don't natively give you this stuff.
42:37
This is my last slide.
42:38
Um a few discussion points and broader themes maybe we can chat about here.
42:41
Obviously, you know, are we going to go with model based?
42:43
Are we going to go with model free?
42:44
Um what's going to be the best way to enable intelligent agents to do
42:47
interesting things in the world?
42:48
regularization and representation learning.
42:51
Um, in this paper they are co-learning the representation of the world
42:55
that the agent has
42:55
and the dynamics of the world.
42:57
Should this be separated?
42:58
Can we take some bio inspiration?
43:00
Should we use pre-existing um like foundation models and stuff like that?
43:03
And then finally, how can we fight uh representational collapse elegantly?
43:07
I think this work does a really great job of that,
43:09
but the question is still out on what the best way to do it
43:11
is.
43:12
So um that's my talk.
43:13
Thanks very much for your attention.
43:21
All right.
43:24
Okay.
43:24
So, for the next two, um, we're kind of focusing on, um,
43:31
less world model stuff and more heady,
43:34
high level stuff that I think is pretty interesting.
43:37
Um, this is a a paper that's going to be presented by Ashe,
43:41
one of the YC uh, startups here named QABs. and your co-founder president.
43:47
You're president of QABs.
43:48
Is that right?
43:49
>> Okay.
43:50
Welcome Ashe.
43:54
>> Hey everybody.
43:55
Today I'm going to be talking through Andrew Gordon Wilson's paper uh deep learning
43:59
is not so mysterious
44:00
or different.
44:01
Uh we actually work with Andrew on the generalization problem at Q Labs.
44:05
So I'm really excited for more people to know about his work.
44:07
The current state of machine learning is
44:09
that we know that scaling
44:10
that scaling models leads to better generalization.
44:14
But we don't have a mechanistic understanding of why that is the case.
44:18
Um yeah, if we can understand general generalization,
44:21
then we might be able to optimize for it as well.
44:24
So the payoff to understanding it is actually really really large.
44:28
Um when you talk to people in the field,
44:30
they often explain that generalization is a mystery
44:32
and they point to examples like overparameterization,
44:36
benign overfitting and and double descent
44:38
as reasons why we might not be able to understand generalization at all.
44:43
So Andrew's work here basically dispels those mysteries by using classical theories of generalization
44:49
uh which which have to date not really been used to explain things like
44:53
like overparameterization thus far.
44:55
So the first classical theory that we'll go through is uh pack bay.
44:59
So pack bay basically bounds the test loss which is the generalization.
45:04
This is the quantity
45:04
that we care about with a training loss
45:06
and a compression term.
45:08
Um the thing is in the past
45:10
when people overparameterize models this compression term tends to dominate
45:14
and so in practice these bounds become loose
45:17
and vacuous meaning that we can't use them for anything at all.
45:21
This was basically due to a mislication of the bound.
45:24
You can compute the the compression term in an alternative way
45:27
as we'll get into sort of later in the talk here.
45:30
So let's go through the first mystery
45:32
that uh Andrew goes through in his paper.
45:34
Um the the mystery that he talks about is overparameterization.
45:38
And this is basically the idea
45:40
that as you scale up the the model parameter size from the bias various
45:44
variance trade-off,
45:45
you would expect that you might overfit.
45:48
But in practice, we see the opposite.
45:50
The scaling laws tell us that we actually get better generalization.
45:53
Um the the the scaling
45:56
and the better generalization from overparameterization is is is due to like the the
46:00
the massive gains in model capability over the last couple of years.
46:03
But we still don't really understand why it impro why it improves generalization.
46:09
So the packbased framework gives us a pretty useful way to think about the
46:13
success of over par parameterization.
46:15
The first is with empirical risk.
46:17
Empirical risk is basically training loss.
46:19
When you increase the number of parameters you can fit your data better.
46:22
Um so the empirical risk the left uh the first term goes down.
46:27
And Andrew's work also finds that when we increase the model,
46:32
when we increase the number of parameters, um we also find more compressible solutions.
46:37
So this is work by Lotfi at all at all
46:39
and they develop methods to basically compress the uh yeah they compress the the
46:44
training set you and
46:46
and and the model
46:47
and they basically find a negative correlation between the the bits required to encode
46:51
the training set and the number of parameters.
46:54
Um and so we find
46:55
that as we increase the model size we can find more efficient encodings of
47:00
the training set.
47:01
So the the second term in this bound also gets lower.
47:07
Another perspective on this model compressibility point is a perspective of flatness.
47:11
As you increase the number of parameters,
47:13
it turns out that the number of the volume of flat minima in parameter
47:18
space exponentially increases.
47:19
This is the green region
47:21
and uh and comparatively the the volume of sharp minima increases much less
47:26
and uh this is interesting
47:28
and this is useful the compressibility view
47:30
because flat minima are known to be more compressible than sharp minima
47:34
and so overparameterization fits within existing theories
47:38
and through Andrew's work we actually see useful bounds on generalization even for models
47:43
at at like a billion parameter scale
47:45
and so we go to the next so-called mystery of deep learning
47:48
which is called uh benign overfitting
47:50
which Andrew also dispels in
47:52
or at least partially explains in his paper.
47:55
So the idea of benign overfitting is
47:57
that deep neural networks are able to fit totally random noise
48:00
but at the same time they are able to to to generalize well
48:03
when you have structured data.
48:05
The mystery is how can you have an inductive bias
48:08
that allows you to generalize well
48:10
if you can also fit totally random data.
48:12
I think a regularized polomial model um in Andrew's paper gives us pretty good
48:17
intuition for how this might be the case.
48:19
Here you can see that on random data,
48:21
so section C of the figure
48:23
that we have enough parameters to fit the data
48:25
and so we we can we can fit the totally random data.
48:29
But on structured data,
48:30
the the regularization pushes us to use the lower order terms.
48:33
And so we are able to both get the flexibility
48:36
but also have inductive bias
48:38
that allows us to generalize.
48:40
And generally this is this is the view to take um for for neural
48:44
networks like there are expressive models with a soft inductive bias.
48:49
Um we can go through this concept um just using this figure right here.
48:53
So uh on the left hand side we have an example of of what's
48:57
like a flexible hypothesis space.
48:59
And a flexible hypothesis space would allow you to fit the data
49:02
that you have.
49:02
But the problem is
49:03
that you would almost certainly overfit
49:06
if you if you um
49:08
if you do not have a bias towards one solution over the other.
49:11
But on the other hand, if you have an inductive bias,
49:13
you would solve this overfitting problem,
49:15
but instead you wouldn't you wouldn't be able to model all of the details
49:19
of reality.
49:20
Um and so the middle ground is to have a very expressive hypothesis space,
49:24
but also have a bias towards solutions that might generalize.
49:28
For example, in the pack bay framework,
49:30
we might want to bias towards more compressible models if we can.
49:34
And so we see
49:35
that uh deep learning so-called mysteries are actually consistent
49:38
and partially explained by existing theories such
49:41
as soft inductive biases
49:42
and pack bays.
49:45
And sort of the thing I want to leave you with is
49:46
that um if if we can find the right inductive biases building on these
49:51
theories,
49:52
we might be able to optimize for them as well.
49:54
And by the no free lunch theorem,
49:56
the only way that we get improvements in learning efficiency is through inductive biases.
50:00
So I I think
50:01
that this is that working on this problem is is a really good bet
50:04
to make.
50:04
Given the massive sample efficiency gap between AI and humans,
50:08
we might actually see massive gains in capability.
50:10
If we work on this problem um and so yeah,
50:13
that's where I want to leave you with short presentation.
50:19
Okay.
50:20
Um so for this last paper
50:22
then after this we have some boba for everyone.
50:25
So sit tight 15 minutes.
50:28
Um this is an idea that you know I've been obsessed with.
50:33
Back to the sample efficiency thing.
50:34
I think that like the two major problems we have left really to solve
50:36
in in AI is intelligence per watt um
50:39
and intelligence per sample.
50:40
And if you compare that to to where we're at today compared to humans,
50:44
um I would say
50:45
that we're still or an order
50:47
or two magnitude off on intelligence per watt.
50:50
Uh and we're me like orders of magnitude off on intelligence per sample.
50:54
I don't know what percent of the internet that you guys have read,
50:57
but I have not read the entire internet.
50:59
In Chris Ray's lab in particular,
51:00
we've been obsessed with this idea
51:02
that um if I have uh under the the a fixed size amount of
51:07
data and I have infinite compute,
51:09
just go nuts, how much generalization can I actually achieve?
51:12
And so this is exactly uh the paper that starts to answer that question.
51:16
And I'm really excited to uh introduce uh Con Woo.
51:24
>> Uh hi, I'm Ku.
51:26
Um this is a paper
51:28
that I co-led with my amazing collaborator Suhas
51:31
as well as Percy
51:32
and Potsu.
51:35
So part of the motivation for this paper is just the fact
51:38
that over the past uh six
51:40
or seven years pre-training has continued to improve model capabilities in pretty surprising ways.
51:46
So in 2020 with GPT3 we had sort of the emergence of incontext learning.
51:52
In 2022 with Anthropics RHF, we had sort of the advent of alignment.
51:58
And maybe most notably in 2024 with both 01 from OpenAI
52:03
and then later Deepseek R1,
52:05
we had the emergence of reasoning.
52:07
And in fact, even still today,
52:09
we see that with these newer
52:10
and bigger pre-training runs like Mythos
52:13
and 5.5,
52:14
the models just continue to keep better.
52:16
And so because pre-training is very expensive,
52:19
a lot of the focus on the research side of things has been on
52:22
how do we improve compute efficiency.
52:25
And in general, people have found that to improve compute efficiency,
52:28
you need to scale both the number of parameters in your model
52:31
and the number of data points
52:32
that you train your model on.
52:34
And so these were quantified with the so-called chinchilla scaling laws.
52:38
The problem with compute efficiency is
52:39
that we're soon going to be constrained by data.
52:42
And so if you look at these sort of public projections of the rate
52:45
of growth of internet data,
52:47
they suggest that the amount of sort of human generated text on the internet
52:51
grows by roughly 3% per year.
52:54
And the amount of compute
52:55
that we're spending on pre-training is growing by roughly four
52:58
or 5x per year.
53:00
And so what this suggests is that as time passes on,
53:05
the amount of compute
53:06
that we're willing to spend per data point is going to continue to increase
53:09
by roughly 4x year-over-year.
53:11
And so this sort of motivates the core question in this paper
53:14
which is how should you approach pre-training
53:17
when you're constrained by data
53:19
but totally unconstrained by compute.
53:21
And it's worth maybe spending a few seconds to think for yourself
53:25
if you haven't already seen this paper like what would you do in this
53:28
situation.
53:29
This is a very different algorithmic regime from sort of the computer efficient pre-training
53:33
world that we've sort of lived in for sort of most of uh uh
53:37
modern time.
53:38
And it's also worth noting
53:40
that this question is not
53:41
that different from how machine learning worked before the modern alm.
53:46
So for things like classical statistics where maybe you really care about your rates
53:50
with respect to the number of points of data you have
53:52
and you don't care about compute
53:54
or even older benchmarks like emnest
53:56
and pen treebank where you're sort of implicitly data constrained
53:59
because the benchmarks don't have
54:00
that many data points.
54:03
And so sort of the core contribution
54:05
that I'll explain in this paper is
54:07
that we bring the modern toolkit of scaling laws to to sort of answer
54:12
this problem.
54:13
And so what we'll show is
54:14
that we'll propose a few different scaling recipes
54:17
and we'll sort of chase scaling recipes
54:20
that monotonically decrease your iid validation laws.
54:24
So sort of in distribution generalization
54:26
and we'll show that these scaling laws have a really clean functional form
54:29
and they follow a super clean power law.
54:31
And when you're able to fit these power laws,
54:33
what you can do is you can estimate the best possible loss of your
54:37
recipe by looking at the asmtote of the power law.
54:40
And this is in some sense a quantification of your best possible performance under
54:44
infinite compute.
54:46
And our goal in this paper is sort of to think more carefully about
54:49
what types of algorithms allow you to lower your compute asmtote.
54:54
Uh and we're sort of going to chase these types of infinite compute wins.
54:57
And so to start,
54:58
I'm going to introduce this canonical setting that we referenced in this paper,
55:01
which is that we're going to simulate a data constrained world by just constraining
55:05
the number of pre-training tokens we have to be a very small amount.
55:08
So we're going to assume access to only 200 million tokens from DCLM,
55:11
which is general web data.
55:13
And what we're going to do is we're going to pre-train large
55:16
and larger models,
55:17
which is the x-axis, using different kinds of pre-training recipes.
55:21
And the y-axis here is going to be again our ID validation loss on
55:25
DS DCLM.
55:26
And our goal is going to be to find recipes
55:29
that allow us to spend more compute
55:30
and train larger models
55:32
while monotonically decreasing our loss.
55:34
So to start, we can consider sort of the obvious approach
55:36
that you might take
55:37
when you're in this setting,
55:38
which is first to epoch your data.
55:40
So to train on the same data points over
55:42
and over again until you start overfitting
55:45
as well as scaling up your model.
55:46
So making your model larger and larger.
55:48
And what we can do is we can do both of these at the
55:50
same time.
55:51
And we can do sort of an exhausted grid search over these parameters until
55:55
we start over until we start overfitting
55:56
and then we do early stopping.
55:58
And this is sort of the red line
56:00
which is what we call the standard recipe.
56:02
And what you'll see with the standard recipe is
56:04
that even if you are willing to spend more compute,
56:07
as you train more and more overparameterized models,
56:10
you start to overfit more quickly
56:12
and your loss starts to increase after a certain point.
56:16
And so if you see this line,
56:18
sort of the natural instinct you should have is how do we fix this?
56:21
And one possible approach is to do really aggressive regularization.
56:24
And so sort of the first baseline in this paper is going to be
56:28
doing really aggressive regularization by cranking up your weight decay.
56:32
And so what we do is we show
56:33
that if you optimally tune your weight decay for each total parameter count.
56:38
So we're going to optimally tune learning rate, weight decay,
56:40
and epoch count for each one of these purple points.
56:43
You can show that your loss follows a really clean power law
56:47
as you increase the number of parameters in your model.
56:50
And this is really aggressive regularization.
56:52
So for context, we use weight decays
56:55
that are something like 30 times larger than the weight decays
56:58
that people do for compute optimal pre-training.
57:00
And so on the legend here,
57:02
you can see the the sort of the form of this power law.
57:05
And it has a few nice properties.
57:07
One is that the exponent on the model parameters n is one.
57:11
And this is actually predicted by sort of the data constraint theory.
57:15
The second nice property
57:16
that it has is
57:17
that the scaling law has an asmtote
57:20
which is 3.43 in this case.
57:22
And this characterizes the performance of the best possible regularized model in this setting
57:27
if you had like infinite compute.
57:30
So you'll notice that the baseline approaches because they overfit more quickly.
57:33
They don't even have a measurable asmtote.
57:35
And so once we start going down the rabbit hole of regularization
57:38
and these other types of classical machine learning techniques,
57:41
there's a whole basket of techniques to to get into.
57:44
And so perhaps maybe the most famous one is to do ensembling.
57:48
And so what we show in this paper is
57:50
that you can bring back ensembling in the modern world of pre-training language models
57:55
and they turn out to be incredibly data efficient.
57:58
So what these light blue points correspond to is they correspond to 300 million
58:03
parameter models that were ensembling with more
58:06
and more members.
58:07
So the fifth point will correspond to 1.5 total billion total parameters
58:12
which is five five ensemble of 300 million parameter models.
58:16
We show that you can also fit really clean scaling laws to ensembles.
58:20
So you also get a power law
58:21
that has exponent one
58:22
and the number of ensemble members
58:24
and it also has an asmtote.
58:26
But most importantly the asmtote of ensembling is much lower than the asmtote of
58:31
the regularized recipe.
58:32
So it's giving you a true data efficiency win
58:35
if you had an infinite amount of compute.
58:37
There's also this interesting property
58:39
which is that ensemblings
58:40
if you do a compute matched comparison
58:42
so the same number of parameters are actually better than the regularized recipe.
58:46
So if your goal is just to train the best 1.5 billion parameter model
58:51
it's better to train an ensemble of a bunch of small models
58:53
when you're data constrained than to train one really large model.
58:57
The last thing we show in this plot is
58:59
that you can actually compose the benefits of regularization
59:03
and ensembling.
59:04
So one way to think about this is
59:06
that regularization gives you this ability to continue to make the models larger
59:10
and larger while ensembling introduces this new axis for scaling compute
59:16
which is by training more
59:17
and more models.
59:18
And so what this gold line
59:20
which we call the joint scaling recipe is we quantify this hypothetical performance
59:25
if we were able to train an ensemble an infinitely large ensemble of infinitely
59:30
large models.
59:31
And so the way in
59:32
which we actually quantify this performance is we fit two scaling laws.
59:37
So we'll take a double limit.
59:39
What we'll first do is we'll train ensembles of 150 million parameter models,
59:44
300 million parameter models and so on and so forth.
59:47
And then we'll look at the asmmptotes of the ensembles.
59:50
And then we'll take a second we'll fit a second scaling law to the
59:52
asmmptotes of these ensembles.
59:54
And this is essentially taking the first limit is taking the limit over K.
59:58
And the second limit is taking the limit over n.
60:00
And what we find is
60:02
that if you're willing to sort of go through the effort of training infinitely
60:06
large models and infinitely many ensembles,
60:09
uh you get a huge loss improvement.
60:10
And so all of these experiments are sort of in this toy data constrained
60:14
setup of 200 million tokens.
60:15
And obviously this is very different from sort of the standard regime of pre-training.
60:19
So what we also do in this paper is we spend some effort on
60:22
trying to confirm that our recipes scale.
60:24
So the first way in
60:25
which we do this is
60:26
that we build data scaling laws.
60:28
So what data scaling laws are is
60:30
that we repeat the exact same set of experiments from the previous slide at
60:33
four different pre-training token counts up to 1.7 billion uh tokens.
60:38
And so for each slice on the x-axis at each seat token count,
60:42
we're going to quantify the best possible performance of each recipe
60:46
if we had an infinite amount of compute.
60:48
So for the red points, they overfit more quickly.
60:51
So these will be actual models.
60:52
While for the purple and the gold points,
60:54
these will correspond to sort of a single limit or a double limit.
60:57
What these data scaling laws let us do is they let us quantify the
61:01
data efficiency numbers of our approaches.
61:04
So one way in
61:04
which we do this is
61:05
if we have some new recipe
61:07
that we believe should improve upon the standard recipe
61:09
that we're using right now,
61:11
you can take the loss of your new recipe
61:14
and you can project it onto the data scaling law.
61:16
So the red line of a standard recipe
61:19
and this projection lets you measure essentially the effective number of extra tokens
61:23
that your algorith algorithmic improvement is buying you.
61:26
So in this case what we see is
61:28
that this joint scaling recipe gives you roughly a 5x data efficiency win over
61:33
uh the the standard recipe.
61:35
It's also worth noting
61:36
that uh these data efficiency wins are something
61:39
that we can realize with sort of finite models not just double limits.
61:43
So for example if you're willing to train a five ensemble of 1 billion
61:46
parameter models this will give you roughly a 3.7x data efficiency win.
61:50
The other interesting aspect about these data scaling laws is
61:53
if you look at the functional form in the legend,
61:56
you'll see that they all have really similar exponents
61:58
and they all have very similar asmtotes.
62:00
And so the reason why this matters is this suggests
62:03
that even if you repeated these experiments at a much much larger token scale,
62:08
if you believe that these data scaling law laws extrapolate,
62:11
this data efficiency win is going to be constant over the actual number of
62:15
token counts that you have.
62:16
So they suggest that this double joint scaling well recipe has a 5x data
62:21
efficiency win even if you are willing to send the seed token count to
62:25
like 10 trillion tokens
62:26
or whatever people are doing pre-training at these days.
62:29
So now I'll go over some methods to sort of make this data efficiency
62:32
win perhaps slightly more practical.
62:34
And so even though these recipes require a lot of training compute we also
62:38
show that you can reduce the amount of inference compute you need by using
62:41
distillation.
62:43
So the plot on the right here,
62:44
the purple line corresponds to the same regularized recipe.
62:48
The light blue points correspond to the same ensemble skilling.
62:51
So we first show
62:52
that what you can do is you can take an eight ensemble
62:54
which is roughly 2.4 billion total parameters
62:57
and you can distill it into a single dense 300 million parameter model
63:01
which is the pink star in the bottom.
63:03
And you can do this while retaining roughly 83% of the loss improvement.
63:08
So this shows you
63:09
that data efficiency is not something
63:11
that you need a large amount of inference compute for.
63:15
If you're willing to amort amortize the test time compute during training time,
63:20
you can get an extremely data efficient model that's still very very small.
63:24
The other surprising result we show in this section is
63:27
that you can do self-distillation to even improve your loss.
63:30
So with self-distillation,
63:32
what we're doing is we're starting with the 300 million parameter model at the
63:35
start of the light blue curve
63:37
and then we're distilling this model into a fresh 300 million parameter model
63:42
which is the green star.
63:43
And what we find is very surprisingly even doing self distillation gives you huge
63:47
loss improvement.
63:48
It even beats the asmtote of the regularized recipe.
63:51
This is actually pretty counterintuitive
63:53
and we have a longer sort of uh description of this result in the
63:57
paper but it turns out to have pretty surprising connections to uh ensembling
64:02
and there's actually a view uh from prior work on viewing self-distillation
64:06
as implicitly training a two ensemble.
64:09
We also show that even
64:10
though we're only chasing IID VAT loss in all of our experiments,
64:14
pretty much all of the trends in this paper directly work on downstream benchmarks.
64:19
And this is like a fully held out sort of test set where we
64:23
only looked at the benchmarks at the very end of the paper
64:25
because the advisers told us to.
64:27
Um, and you can see that everything tracks the standard recipe overfits.
64:32
Still model scaling gives you improvements.
64:35
Ensembling is even better.
64:36
and you can still retain a lot of the benefits through distillation.
64:39
And finally, we also show
64:41
that you can do this for other settings beyond pre-training.
64:43
So things like continued pre-training.
64:45
So we consider a setup where you're trying to CPT a 3B model
64:49
and we assume access to sort of this restricted set of 4 billion math
64:54
related tokens where the whole corpus of data is actually 73 billion tokens.
64:59
And what we show is
65:00
that if you're willing to do these data efficiency tricks like aggressive epoing
65:04
and things like ensembling,
65:06
you can match the performance of training on the full 73 billion tokens even
65:10
using only 4 billion tokens
65:12
which is roughly a 17x data efficiency win.
65:15
So to sort of wrap up this talk,
65:17
maybe the main point I want to make is
65:19
that when you're constrained by data
65:21
and you're unconstrained by compute
65:22
and this sort of new algorithmic regime,
65:25
the types of algorithmic choices you make matter a lot
65:28
and we should be willing to sort of rethink every aspect of a stack.
65:31
In this paper, we mostly do this by revisiting a lot of these classical
65:35
ideas from uh machine learning
65:37
and deep learning.
65:38
Things like regularization, ensembling, distillation have existed for for many many years.
65:44
And we also introduced this evaluative tool of asmmptotes.
65:48
And maybe the hope is
65:49
that if you're willing to chase algorithms
65:51
that have lower compute asmmptotes,
65:53
uh these will give you like better ideas for data efficiency.
65:56
But like ultimately what we really want to do is we want these asmtotes
65:59
to help us develop new
66:01
and better ideas under infinite compute
66:03
that that don't already exist.
66:05
And so if you're interested in the details,
66:07
that's a QR code for the paper.
66:09
And we've also done some follow-up work on looking at how synthetic data interacts
66:12
with data efficiency.
66:13
So feel free to check that out as well if you're interested.
66:16
Thanks.
66:22
>> All right.
66:23
Thank you guys so much for coming.
66:25
This is like a dream come true.
66:26
I'm in one of my favorite places
66:28
that um was most important places of my life
66:31
and now I get to talk about AI here.
66:34
So super super fun.
66:35
I think there's a lot of potential for this club.
66:37
I think I don't have nearly, you know,
66:40
1% of all the ideas
66:42
that we probably have to make this club really great um in all of
66:46
your heads.
66:47
And so we want to make sure all of you guys get in on
66:50
the Slack.
66:50
So I'll make sure that you know,
66:52
please send me a note if you're not already on there.
66:54
And then we can kind of make this thing whatever we want.
66:56
So it's kind of fun and I intend to.
66:58
So like please come with ideas.
67:01
We want to make this super fun.
67:02
Um obviously, you know, there's some round rules, be respectful,
67:05
all that kind of stuff.
67:06
Um, and definitely be involved.
67:07
And that's kind of the the the biggest thing
67:09
that we really only really ask.
67:11
That's all I got.
67:11
That's a wrap.
67:12
Go get some boba tea.
67:13
Thank you.
Like
Share
Y Combinator
View all →
C1
Business
Karaoke
22:36
OpenClaw Creator: Why 80% Of Apps Will Disappear
Y Combinator
5
C1
Business
Karaoke
7:51
The New Way To Build A Startup
Y Combinator
C1
Business
Karaoke
13:07
Inside The Startup Reinventing The $6 Trillion Chemical Manufacturing Industry
Y Combinator
C1
Business
Karaoke
40:57
Demis Hassabis: Agents, AGI & The Next Big Scientific Breakthrough
Y Combinator
C1
Business
Karaoke
10:27
How To Build A Company With AI From The Ground Up
Y Combinator
C1
Business
Karaoke
53:20
The Future Of Brain-Computer Interfaces
Y Combinator
C1
Business
Karaoke
50:10
Inside Claude Code With Its Creator Boris Cherny
Y Combinator
C1
Business
Karaoke
21:49
How to Make Claude Code Your AI Engineering Team
Y Combinator
1
Suggested videos
C1
Business
Karaoke
5:49
You Need to Be Bored. Here's Why.
Harvard Business Review
7
C1
Business
Karaoke
4:28
Lean Into Imposter Syndrome, Don't Give In to It
Harvard Business Review
C1
Business
6:30
Identity Crisis: Why Defining Yourself by Your Career Is a Problem
Harvard Business Review
C1
Business
Karaoke
22:36
OpenClaw Creator: Why 80% Of Apps Will Disappear
Y Combinator
5
C1
Business
Karaoke
6:36
Fighting Workaholism: You Are Not a Success Machine
Harvard Business Review
C1
Business
Karaoke
2:47
Can Work Make You Happy? Should It?
Harvard Business Review
C1
Business
Karaoke
7:51
The New Way To Build A Startup
Y Combinator
C1
Business
Karaoke
13:07
Inside The Startup Reinventing The $6 Trillion Chemical Manufacturing Industry
Y Combinator