TryShadowing
Shadow YouTube. Nói tiếng Anh.
Trang chủ
Khám phá
Chép chính tả
NEW
Thư viện
vi
0
ngày
Đăng nhập
The Strange Math That Predic… — Veritasium luyện shadowing | TryShadowing
TryShadowing
Shadow YouTube. Nói tiếng Anh.
Trang chủ
Khám phá
Chép chính tả
NEW
Thư viện
vi
0
ngày
Đăng nhập
Trang chủ
Khám phá
Veritasium
The Strange Math That Predicts (Almost) Anything
The Strange Math That Predicts (Almost) Anything
Veritasium
·
32:32 · 25 thg 7, 2025
Bắt đầu học
0:00
0:00
Ghi âm
×1
1x
VI
EN
JA
KO
ZH
FR
PT
TH
IT
DE
IPA
Chấm điểm phát âm chưa hỗ trợ trên trình duyệt này — bạn vẫn ghi âm & nghe lại được.
-
How
many
times
do
you
need
to
shuffle
a
deck
of
cards
to
Đang dịch…
Bật Ghi âm để được thu giọng và chấm điểm
Thông minh
Karaoke
Câu gốc
Câu
1
/597
0:00
- How many times do you need to shuffle a deck of cards to
0:03
make them truly random?
0:05
How much uranium does it take to build a nuclear bomb?
0:10
(explosion booming) How can you predict the next word in a sentence?
0:12
And how does Google know which page you're actually searching for?
0:16
Well, the reason we know the answer to all of these questions is
0:19
because of a strange math feud in Russia
0:22
that took place over 100 years ago.
0:26
In 1905, socialist groups all across Russia rose up against the Tsar,
0:32
the ruler of the empire.
0:33
They demanded a complete political reform, or failing that,
0:36
that he stepped down from power entirely. - This divided the nation into two.
0:41
So on one side you got the Tsarists, right?
0:44
They wanted to defend the status quo and keep the Tsar in power.
0:48
But then on the other side,
0:49
you had the socialists who wanted this complete political reform.
0:53
And this division was
0:54
so bad that it crept into every part of society to the point where
0:58
even mathematicians started picking sides. - On the side of the Tsar was Pavel
1:02
Nekrasov,
1:03
unofficially called the Tsar of Probability.
1:06
Nekrasov was a deeply religious and powerful man,
1:10
and he used his status to argue
1:12
that math could be used to explain free will
1:14
and the will of God. - His intellectual nemesis on the socialist side was
1:19
Andrey Markov,
1:21
also known as Andrey The Furious.
1:24
Markov was an atheist
1:25
and he had no patience for people who were being unrigorous,
1:29
something he considered Nekrasov to be, because in his eyes,
1:32
math had nothing to do with free will or religion.
1:35
So he publicly criticized Nekrasov's work, listing it among "the abuses of mathematics."
1:41
Their feud centered on the main idea people had used to do probability for
1:45
the last 200 years.
1:47
And we can illustrate this with a simple coin flip.
1:50
When I flip the coin 10 times,
1:51
I get six times heads and four times tails,
1:54
which is obviously not the 50/50 you'd expect.
1:57
But if I keep flipping the coin,
1:59
then at first the ratio jumps all over the place.
2:02
But after a large number of flips,
2:04
we see that it slowly settles down and approaches 50/50.
2:08
And in this case, after 100 flips,
2:10
we end up on 51 heads and 49 tails,
2:14
which is almost exactly what you would expect.
2:17
This behavior that the average outcome gets closer
2:20
and closer to the expected value
2:22
as you run more
2:23
and more independent trials is known
2:25
as the law of large numbers.
2:27
It was first proven by Jacob Bernoulli in 1713,
2:30
and it was the key concept at the heart of probability theory right up
2:34
until Markov and Nekrasov.
2:36
But Bernoulli only proved
2:38
that it worked for independent events like a fair coin flip,
2:42
or when you ask people to guess how much they think an item is
2:45
worth,
2:45
where one event doesn't influence the others.
2:49
But now imagine that instead of asking each person to submit their guess individually,
2:54
you ask people to shout out their answer in public.
2:57
Well, in this case, the first person might think it's an extraordinarily valuable item,
3:02
and say it's worth around $2,000,
3:05
but now all the other people in the room are influenced by this value,
3:09
and so, their guesses have become dependent.
3:12
And now the average doesn't converge to the true value,
3:16
but instead it clusters around a higher amount. - And so, for 200 years,
3:21
probability had relied on this key assumption,
3:24
that you need independence to observe the law of large numbers.
3:28
And this was the idea that sparked Nekrasov and Markov's feud.
3:32
See, Nekrasov agreed with Bernoulli
3:34
that you need independence to get the law of large numbers.
3:37
But he took it one step further.
3:39
He said, if you see the law of large numbers,
3:42
you can infer that the underlying events must be independent. - Take this table
3:48
of Belgian marriages from 1841 to 1845.
3:52
Now you see that every year the average is about 29,000.
3:55
And so, it seems like the values converge
3:58
and therefore that they follow the law of large numbers.
4:01
And when Nekrasov looked at other social statistics like crime rates and birth rates,
4:05
he noticed a similar pattern.
4:08
But now think about where all this data is coming from.
4:10
It's coming from decisions to get married, decisions to commit crimes,
4:14
and decisions to have babies, at least for the most part.
4:18
So Nekrasov reasoned that because these statistics followed the law of large numbers,
4:22
the decisions causing them must be independent.
4:24
In other words, he argued that they must be acts of free will.
4:28
So to him, free will wasn't just something philosophical,
4:32
it was something you could measure.
4:34
It was scientific. - But to Markov, Nekrasov was delusional.
4:40
He thought it was absurd to link mathematical independence to free will.
4:45
So Markov set out to prove
4:47
that dependent events could also follow the law of large numbers,
4:51
and that you can still do probability with dependent events. - To do this,
4:56
he needed something where one event clearly depended on what came before,
5:01
and he got the idea that this is what happens in text.
5:04
Whether your next letter is a consonant
5:06
or a vowel depends heavily on what the current letter is.
5:10
So to test this,
5:11
Markov turned to a poem at the heart of Russian literature,
5:15
"Eugene Onegin" by Alexander Pushkin. - He took the first 20,000 letters of the
5:21
poem,
5:21
stripped out all punctuation and spaces,
5:24
and pushed them together into one long string of characters.
5:27
He counted the letters and found that 43% were vowels and 57% were consonants.
5:33
Then Markov broke the string into overlapping pairs, that gave him four possible combinations,
5:40
vowel-vowel, consonant-consonant, vowel-consonant, or consonant-vowel.
5:44
Now, if the letters were independent,
5:46
the probability of a vowel-vowel pair would just be the probability of a vowel
5:50
twice,
5:51
which is about 0.18 or an 18% chance.
5:56
But when Markov actually counted,
5:58
he found vowel-vowel pairs only show up 6% of the time,
6:03
way less than if they were independent.
6:06
And when he checked the other pairs,
6:07
he found that all actual values differed greatly from what the independent case would
6:13
predict.
6:13
So Markov had shown that the letters were dependent.
6:18
And to beat Nekrasov,
6:19
all he needed to do now was show
6:21
that these letters still followed the law of large numbers.
6:24
So he created a prediction machine of sorts.
6:28
He started by drawing two circles,
6:30
one for a vowel and one for a consonant.
6:32
These were his states.
6:34
Now, say you're at a vowel,
6:36
then the next letter could either be a vowel or a consonant.
6:39
So he drew two arrows to represent these transitions.
6:43
But what are these transition probabilities?
6:46
Well, Markov knew that if you pick a random starting point,
6:49
there is a 43% chance that it'll be a vowel.
6:52
He also knew that vowel-vowel pairs occur about 6% of the time.
6:57
So to find the probability of going from a vowel to another vowel,
7:00
he divided 0.06 by 0.43 to find a transition probability of about 13%.
7:07
And since there is a 100% chance that another letter comes next,
7:11
all the arrows going from the same state need to add up to one.
7:15
So the chance of going to a consonant is one minus 0.13,
7:19
or 87%.
7:21
He repeated this process for the consonants to complete his predictive machine.
7:26
So let's see how it works.
7:28
We'll start at a vowel.
7:31
Next, we generate a random number between zero and one.
7:34
If it's below 0.13, we get another vowel, if it's above,
7:38
we get a consonant.
7:39
We got 0.78, so we get a consonant, then we generate another number,
7:43
and check if it's above or below 0.67, 0.21,
7:47
so we get a vowel.
7:50
Now, we can keep doing this
7:51
and keep track of the ratio of vowels to consonants.
7:54
At first, the ratio jumps all over the place, but after a while,
7:58
it converges to a steady value, 43% vowels and 57% consonants,
8:04
the exact split Markov had counted by hand.
8:09
So Markov had built a dependent system, a literal chain of events,
8:13
and he showed that it still followed the law of large numbers,
8:17
which meant that observing convergence in social statistics didn't prove
8:20
that the underlying decisions were independent.
8:23
In other words, those statistics don't prove free will at all.
8:27
Markov had shattered Nekrasov's argument, and he knew it.
8:31
So he ended his paper with one final dig at his rival.
8:35
"Thus, free will is not necessary to do probability."
8:39
In fact, independence isn't even necessary to do probability.
8:43
With this Markov chain, as it came to be known,
8:45
he found a way to do probability with dependent events.
8:49
This should have been a huge breakthrough, because in the real world,
8:53
almost everything is dependent on something else.
8:56
I mean, the weather tomorrow depends on the conditions today.
9:00
How a disease spreads depends on who's infected right now,
9:03
and the behavior of particles depends on the behavior of particles around them.
9:08
Many of these processes could be modeled using Markov chains.
9:13
Do people think it was like a mic drop moment and like, "Oh,
9:16
Nekrasov's out, like, Markov's the man"?
9:19
Or people didn't really notice, or it was obscure,
9:22
or? - I feel like people didn't really notice, like,
9:25
it wasn't a really big thing.
9:28
And Markov himself seemingly didn't care much about how it might be applied to
9:32
practical events.
9:34
He wrote, "I'm concerned only with questions of pure analysis.
9:38
I refer to the question of the applicability with indifference."
9:43
Little did he know
9:44
that this new form of probability theory would soon play a major role in
9:49
one of the most important developments of the 20th century.
9:54
On the morning of the 16th of July, 1945,
9:58
the United States detonated The Gadget, the world's first nuclear bomb.
10:04
The six kilogram plutonium bomb created an explosion
10:08
that was equivalent to nearly 25,000 tons of TNT.
10:12
This was the culmination of the top secret Manhattan Project,
10:16
a three-year long effort by some of the smartest people alive,
10:20
including people like J.
10:22
Robert Oppenheimer, John von Neumann,
10:24
and a little known mathematician named Stanislaw Ulam. - Even after the war ended,
10:31
Ulam continued trying to figure out how neutrons behave inside a nuclear bomb.
10:35
Now, a nuclear bomb works something like this.
10:38
Say you have a core of uranium-235,
10:41
then when a neutron hits a U-235 nucleus,
10:44
the nucleus splits releasing energy and, crucially, two or three more neutrons.
10:50
If, on average, those new neutrons go on to hit
10:52
and split more than one other U-235 nucleus,
10:56
you get a runaway chain reaction, so you have a nuclear bomb.
11:00
But uranium-235, the fissile fuel needed for the bombs was really hard to get.
11:05
So one of the key questions was just how much of it do you
11:08
need to build a bomb?
11:10
And this is why Ulam wanted to understand how the neutrons behave. -
11:15
But then in January of 1946,
11:18
everything came to a halt.
11:20
Ulam was struck by a sudden and severe case of encephalitis,
11:24
an inflammation of the brain, that nearly killed him.
11:28
His recovery was long and slow,
11:30
with Ulam spending most of his time in beds.
11:34
To pass the time, he played a simple card game, Solitaire.
11:38
But as he played countless games, winning some, losing others,
11:42
one question kept nagging at him,
11:45
what are the chances that a randomly-shuffled game of Solitaire could be won?
11:50
It was a deceivingly difficult problem to solve.
11:53
Ulam played with all 52 cards where each arrangement created a unique game,
11:58
so the total number of possible games was 52 factorial,
12:02
or about eight times 10 to 67.
12:06
So solving this analytically was hopeless.
12:10
But then Ulam had a flash of insight,
12:12
what if I just play hundreds of games
12:14
and count how many could be won?
12:16
That would give him some sort of statistical approximation of the answer.
12:21
Back at Los Alamos, the remaining scientists grappled with much harder problems than Solitaire,
12:26
like figuring out how neutrons behave inside a nuclear core.
12:32
In a nuclear core,
12:32
there are trillions and trillions of neutrons all interacting with their surroundings.
12:36
So the number of possible outcomes is immense,
12:39
and computing it directly seemed impossible. - But when Ulam returned to work,
12:44
he had a sudden revelation.
12:46
What if we could simulate these systems by generating lots of random outcomes like
12:50
I did with Solitaire?
12:52
He shared this idea with von Neumann, who immediately recognized its power,
12:57
but also spotted a key problem. - See, in Solitaire, each game is independent.
13:03
How the cards are dealt in one game have no effect on the next,
13:07
but neutrons aren't like that.
13:09
A neutron's behavior depends on where it is and what it has done before.
13:14
So you couldn't just sample random outcomes like in Solitaire.
13:18
Instead, you needed to model a whole chain of events where each step influenced
13:23
the next.
13:24
What von Neumann realized is that you needed a Markov chain.
13:28
So they made one
13:30
and a much simplified version of it works something like this.
13:34
Now, the starting state is just a neutron traveling through the core,
13:37
and from there, three things can happen.
13:39
It can scatter off an atom and keep traveling,
13:42
so that gives you an arrow going back to itself.
13:45
It can leave the system or get absorbed by a non-fissile material,
13:49
in which case it no longer takes part in the chain reaction,
13:52
and so it ends its Markov chain,
13:55
or it can strike another uranium-235 atom,
13:58
triggering a fission event
14:00
and releasing two or three more neutrons
14:02
that then start their own chains.
14:05
But in this chain, the transition probabilities aren't fixed,
14:08
they depend on things like the neutron's position, velocity and energy,
14:12
as well as the overall configuration and mass of uranium.
14:16
So a fast-moving neutron might have a 30% chance to scatter,
14:20
a 50% chance to be absorbed or leave,
14:22
and a 20% chance to cause fission.
14:25
But a slower-moving neutron would have different probabilities.
14:29
Next, they ran this chain on the world's first electronic computer, the ENIAC.
14:34
The computer started by randomly generating a neutron starting conditions
14:38
and stepped through the chain to keep track of how many neutrons were produced
14:41
on average per run,
14:43
known as the multiplication factor k.
14:46
So if, on average, one neutron produces another two neutrons,
14:50
then k is equal to two.
14:52
And if on average every two neutrons produce three neutrons,
14:55
then k is equal to three over two, and so on.
14:59
Then, after stepping through the full chain for a specified number of steps,
15:03
we collect the average k-value and record that number in a histogram.
15:07
This process was then repeated hundreds of times, and the results tallied up,
15:11
giving you a statistical distribution of the outcome.
15:15
If you find that in most cases, k is less than one,
15:18
the reaction dies down.
15:19
If it's equal to one, there's a self-sustaining chain reaction,
15:23
but it doesn't grow.
15:24
And if k is larger than one,
15:26
the reaction grows exponentially and you've got a bomb. - With it,
15:31
von Neumann and Ulam had a statistical way to figure out how many neutrons
15:35
were produced without having to do any exact calculations.
15:39
In other words, they could approximate differential equations
15:42
that were too hard to solve analytically.
15:45
All that was needed was a name for the new method.
15:48
Now, Ulam's uncle was a gambler,
15:50
and the random sampling
15:52
and high stakes reminded Ulam of the Monte Carlo Casino in Monaco,
15:56
and the name stuck.
15:58
The Monte Carlo method was born.
16:01
The method was so successful that it didn't stay secret for long.
16:05
By the end of 1948, scientists at another lab, Argonne, in Chicago,
16:10
used it to study nuclear reactor designs, and from there, the idea spread quickly.
16:16
Ulam later remarked, "It is still an unending source of surprise for me to
16:21
see how a few scribbles on a blackboard could change the course of human
16:25
affairs."
16:27
And it wouldn't be the last time Markov chain based method changed the course
16:31
of human affairs. (upbeat music) - In 1993,
16:37
the internet was open to the public, and soon it exploded.
16:41
By the mid-1990s, thousands of new pages appeared every day,
16:44
and that number was only growing.
16:48
This created a new kind of problem.
16:50
I mean, how do you find anything in this ever-expending sea of information?
16:55
In 1994, two Stanford PhD students, Jerry Yang and David Filo,
17:00
founded the search engine Yahoo, to address this issue, but they needed money.
17:06
So a year later, they arranged to meet with Japanese billionaire, Masayoshi Son,
17:11
also known as the Bill Gates of Japan. (gong clangs) - They were looking
17:15
to raise $5 million for their next startup,
17:19
but Son has other plans.
17:22
He offers to invest a full $100 million instead.
17:26
That's 20 times more than what the founders asked for.
17:29
So Jerry Yang declines saying, "We don't need that much," but Son disagrees, "Jerry,
17:36
everyone needs $100 million."
17:40
(Son laughs) Before the founders get a chance to respond,
17:42
Son jumps in again and asks, "Who are your biggest competitors?"
17:46
"Excite and lycos," the pair respond.
17:49
Son orders his associate to write those names down.
17:51
And then he says, "If you don't let me invest in Yahoo,
17:54
I will invest in one of them and I'll kill you."
17:59
See, Son had realized something.
18:01
None of the leading search engines at the time had any superior technology.
18:05
They didn't have a technological advantage over the others.
18:09
They all just ranked pages by how often a search term appears on a
18:13
given page.
18:14
So the battle for the number one search engine would be decided by who
18:17
could attract the most users,
18:19
who could spend the most on marketing. - Lycos, go get it. - Get Lycos,
18:24
or get lost. - This is revolution. (upbeat funky music) ♪ Yahoo ♪ -
18:33
And marketing required a lot of money,
18:35
money that Son had, so he could decide who won the war.
18:40
Yahoo's founders realized they were left with no real choice
18:43
but to accept Son's investment. -
18:46
So here we are,
18:46
right in the middle of Yahoo. - And within four years,
18:49
Yahoo became the most popular site on the planet. - In the time it
18:53
takes to say this sentence,
18:55
Yahoo will answer 79,000 information requests worldwide,
19:00
the two men are now worth $120 million each.
19:07
♪ Yahoo ♪ - But Yahoo had a critical weakness.
19:11
See, Yahoo's keyword search was easy to trick.
19:14
To get your page ranked highly, you could just repeat keywords hundreds of times,
19:18
hidden with white text on a white background. - One thing they didn't have
19:23
in those early days was a notion of quality of the result.
19:28
So they had a notion of relevance saying,
19:31
does this document talk about the thing that you're interested in?
19:35
But there wasn't really a notion of
19:37
which ones are better. - What they really needed was a way to rank
19:41
pages by both relevance
19:42
and quality.
19:44
But how do you measure the quality of a webpage?
19:46
Well, to understand that,
19:48
we need to borrow an idea from libraries. -
19:50
So I'm old enough
19:51
that library books used to have a paper card in it
19:55
that was a stamp of all the due dates of
19:57
when it was due back.
19:58
You took a book and if it had a lot of those, you said,
20:00
"Oh, this is probably a good book."
20:01
And if it didn't have any, you said, "Well,
20:04
maybe this isn't the best book." - Stamps acted like endorsements.
20:07
The more stamps, the better the book must be.
20:10
And the same idea can be applied to the web.
20:12
Over at Stanford, two PhD students, Sergey Brin and Larry Page,
20:17
were working on this exact problem.
20:19
Brin and Page realized
20:20
that each link to a page can be thought of
20:23
as an endorsement.
20:24
And the more links a page sends out, the less valuable each vote becomes.
20:29
So what they realized is
20:31
that we can model the web
20:32
as a Markov chain. - To see how this works,
20:36
imagine a toy internet with just four webpages.
20:39
Call them Amy, Ben, Chris, and Dan.
20:42
These are our states.
20:44
Typically, one webpage links to others, allowing you to move between them.
20:48
These are our transitions.
20:50
In this setup, Amy only links to Ben,
20:52
so there's a 100% chance of going from Amy to Ben.
20:56
Ben links to Amy, Chris, and Dan,
20:59
so there's a 33% chance of going to any of those pages,
21:02
and we can fill out the other transition probabilities in the same way.
21:07
So now we can run this Markov chain and see what happens.
21:10
Imagine you're a surfer on this web.
21:13
You start on a random page, say, Amy,
21:15
and you keep running the machine
21:17
and keep track of the percentage of time you spend on each page.
21:21
Over time, the ratio settles
21:23
and the scores give us some measure of the relative importance of these pages.
21:28
You spend the most time on Ben, so Ben is ranked first,
21:30
followed by Amy, then Dan, and lastly Chris.
21:35
It might seem like there's an easy way to beat the system,
21:37
just make 100 pages all linking to your website.
21:40
Now you get 100 full votes and you'll always rank on top,
21:44
but that is not the case.
21:47
While during their first few steps, they might make your page seem important,
21:50
none of the other websites link to them.
21:53
So over many steps, their contributions don't matter.
21:57
You might have many links, but they're not quality links,
22:00
so they don't affect the algorithm. - But there is still one problem, though,
22:05
not all pages are connected.
22:07
In networks like this one, a random server can get stuck in a loop,
22:11
never reaching the rest of the web.
22:13
So to fix this, we can set a rule that 85% of the time,
22:17
our random server just follows a link like normal.
22:20
But then for about 15% of the time,
22:22
they just jump to a page at random.
22:25
This damping factor makes sure
22:27
that we explore all possible parts of the web without ever getting stuck.
22:32
By using Markov chains, Page and Brin had built a better search engine,
22:36
and they called it PageRank. - Because it's talking about how pages react,
22:42
webpages react with each other and also 'cause the founder's name is Larry Page,
22:46
so he snuck that in. - With PageRank, Google got much better search results,
22:51
often getting you to the site you were looking for in one go.
22:54
Although, to some, this sounded like a terrible idea. - Others said, "Oh,
22:59
well you're telling me you get a search
23:00
that will get the right result on the first answer?
23:04
Well, I don't want that because if it takes them three or four chances,
23:08
searches to get the right answer,
23:10
then I have three or four chances to show ads,
23:13
and if you get 'em the answer right away, I'm just gonna lose them.
23:16
So, you know, I don't see why better search is better." -
23:20
But Page and Brin disagreed.
23:22
They were convinced that if their product was far superior,
23:24
then people would flock to it. - I would say it actually is a
23:28
democracy that works.
23:30
If all pages were equal, anybody can manufacture as many pages as they want.
23:35
I can set up a billion pages in my server tomorrow.
23:38
We shouldn't treat them all as equal.
23:40
Just looking at the data out of curiosity,
23:42
we found that we had technology to do a better job of search,
23:45
and we realized how impactful having great search can be. - And so, in 1998,
23:51
they launched their new search engine to take on Yahoo.
23:54
Initially, they called it BackRub, after the backlinks it analyzed,
23:58
but then they realized that maybe that's not the most attractive name.
24:02
Now, their ambitions were big to essentially index all the pages on the internet,
24:06
and they needed a name equally as big.
24:09
So they thought of the largest number they could think of,
24:12
10 to the power of 100, a googol.
24:15
But then when trying to register their domain, they accidentally misspelled it.
24:19
And so, Google was born. (dramatic music) Over the next four years,
24:28
Google overthrew Yahoo to become the most used search engine. - Everyone who knows
24:32
the internet almost certainly knows Google. - Googling is like oxygen to teenagers. -
24:37
And today,
24:38
Alphabet, which is Google's parent company,
24:40
is worth around $2 trillion. -
24:43
When Google makes even the slightest change in its algorithms,
24:46
it can have huge effects. - Google. - Google. - Google. - Google. - They're on fire.
24:52
And the reason why they're on fire is
24:54
because they're focused and they're more focused than Yahoo who does search,
24:57
they're more focused than Microsoft who does search with Bing.
24:59
Yahoo has lots of traffic, they always have, they have some really great properties,
25:03
but I don't think Yahoo is the go-to place,
25:05
you know. - And at the heart of this trillion dollar algorithm is a
25:09
Markov chain,
25:11
which only looks at the current state to predict what's going to happen next.
25:16
But in the 1940s, Claude Shannon, the father of information theory,
25:20
started asking a different question.
25:23
He went back to Markov's original idea of predicting text,
25:26
but instead of just using vowels and consonants, he focused on individual letters.
25:31
And he wondered, what
25:32
if instead of looking at only the last letter
25:35
as a predictor,
25:36
I look at the last two?
25:37
Well, with that, he got text that looked like this.
25:41
Now, it doesn't make much sense, but there are some recognizable words like "whey",
25:45
"of", and "the".
25:47
But Shannon was convinced he could do better.
25:49
So next, instead of looking at letters, he wondered,
25:52
what if I use entire words as predictors?
25:55
That gave him sentences like this,
25:58
"The head and in frontal attack on an English writer
26:00
that the character of this point is therefore another method for the letters
26:04
that the time of who ever told the problem for an unexpected."
26:09
Now, clearly, this doesn't make any sense,
26:12
but Shannon did notice
26:13
that sequences of four words
26:14
or so generally did make sense.
26:17
For instance, "attack on an English writer" kind of makes sense.
26:21
So Shannon learned that you can make better
26:23
and better predictions about what the next word is going to be by taking
26:26
into account more and more of the previous words.
26:30
It's kind of like what Gmail does
26:32
when it predicts what you're going to type next.
26:35
And this is no coincidence,
26:37
the algorithms that make these predictions are based on Markov chains. - They're not
26:41
necessarily using letters,
26:43
you know, -Yeah they use what they call tokens, some of which are letters,
26:48
some of which are words, marks of punctuation, whatever.
26:52
So it's a bigger set than just the alphabet.
26:56
The game is simply, we have this string of tokens that, you know,
27:01
might be 30 long,
27:04
and we're asking what are the odds
27:06
that the next token is this
27:08
or this or this? -
27:09
But today's large language models don't treat all those tokens equally,
27:13
because unlike simple Markov chains, they also use something called attention,
27:17
which tells the model what to pay attention to.
27:20
So in the phrase,
27:20
"the structure of the cell," the model can use previous context like blood
27:25
and mitochondria to know the cell most likely refers to biology rather than a
27:29
prison cell.
27:30
And it uses that to tune its prediction.
27:33
But as large language models become more widespread,
27:36
one concern is that the text they produce ends up on the internet
27:39
and that becomes training data for future models. -
27:43
When you start doing
27:44
that,
27:46
the game is very soon over.
27:48
You come, in this case, to us, a very dull, stable state,
27:52
it just says the same thing over and over and over again forever.
27:55
The language models are vulnerable to this process. -
27:59
And any system like this where we have a feedback loop,
28:02
will become hard to model using Markov chains.
28:05
Take global warming, for instance,
28:07
as we increase the amount of carbon dioxide in the air,
28:09
the average temperature of the Earth increases.
28:12
But as the temperature increases, the atmosphere can hold more water vapor,
28:15
which is an incredibly powerful greenhouse gas.
28:18
And with more water vapor,
28:19
the temperature increases further allowing for even more water vapor.
28:23
So you get this positive feedback loop,
28:25
which makes it hard to predict what's going to happen next.
28:28
So there are some systems where Markov chains don't work,
28:31
but for many other dependent systems,
28:33
they offer a way of doing probability. -
28:36
But what's fascinating is
28:37
that all these systems have extremely long histories.
28:41
I mean, you could trace back all the letters in a text,
28:43
trace back all the interactions of what a neutron did,
28:46
or trace back the weather for weeks.
28:49
But the beautiful thing Markov
28:50
and others found is
28:51
that for many of these systems you can ignore almost all of
28:55
that.
28:55
You can just look at the current state and forget about the rest,
29:00
that makes these systems memoryless.
29:02
And it's this memoryless property
29:05
that makes Markov chains
29:06
so powerful because it's what allows you to take these extremely complex systems
29:11
and simplify them a lot to still make meaningful predictions. -
29:16
As one paper put it,
29:17
"Problem-solving is often a matter of cooking up an appropriate Markov chain." - It's
29:22
kind of ridiculous to me
29:23
that this basic fact of mathematics would come out of a fight like
29:28
that,
29:28
which, you know, really had nothing to do with it.
29:31
But all the evidence suggests
29:33
that it really was this determination to show up Nekrasov
29:38
that led Markov to do the work. -
29:41
But there's one question we still haven't answered.
29:44
When playing Solitaire, how did Ulam know his cards were perfectly shuffled?
29:49
I mean, how many shuffles does it take to get a completely random arrangement
29:54
of cards? - If you have a deck of cards,
29:57
you need to shuffle it, right?
30:00
How often, if you're shuffling, like, you know, you split it in half,
30:02
and then you do the (cards riffling).
30:05
How often do you have to shuffle it to make it completely random? -
30:10
Two. - Two?
30:11
I'm going with 26. - Yeah, four times. - Four times? - I don't know,
30:14
52 times? - Okay.
30:16
Okay.
30:16
It's not a bad guess. - Seven? - It is seven. - Really? - Yeah.
30:22
So you can think of card shuffling
30:23
as a Markov chain where each deck arrangement is a state,
30:26
and then each shuffle is a step.
30:28
And so for a deck of 52 cards,
30:30
if you riffle shuffle it seven times,
30:32
then every arrangement of the deck is about equally likely, so it's basically random.
30:38
But I can't shuffle like that.
30:40
So for me, what I do is I do it like this.
30:43
How many times do you think you have to shuffle like this to get
30:46
it random?
30:48
(beeping) - What do you think?
30:49
And perhaps more importantly, how would you go about working it out?
30:54
Well, that's where today's sponsor Brilliant comes in.
30:56
Brilliant is a learning app
30:57
that gets you hands-on with problems just like this.
31:01
Whether it's math, physics, programming, or even AI,
31:04
Brilliant's interactive lessons and challenges let you play your way to a sharper mind.
31:09
You can discover how large language models actually work from basic Markov chains to
31:14
complex neural networks,
31:16
or dig into the math behind this shuffling question.
31:19
It's a fun way to build knowledge
31:20
and skills that help you solve all kinds of problems,
31:24
which brings us back to our shuffle.
31:26
So Casper, what actually is the answer? - It's actually over 2,000. - What?
31:30
- Over- - Crazy,
31:33
right? - Yeah. - So the next time someone offers to shuffle before a game,
31:36
make sure they're doing it right, seven riffles or it doesn't count.
31:40
But the interesting part isn't just knowing that,
31:42
it's understanding why and seeing how a simple question can lead you to some
31:46
surprisingly complex mathematics.
31:49
And that's what Brilliant is all about.
31:52
So to try everything Brilliant has to offer for free for a full 30
31:56
days,
31:56
visit brilliant.org/veritasium,
31:58
click that link in the description or scan this handy QR code.
32:02
And if you sign up, you'll also get 20% off their annual premium subscription.
32:06
So I wanna thank Brilliant for sponsoring this video
32:09
and I wanna thank you for watching. (upbeat music) - Easy. (light music)
Thích
Chia sẻ
Veritasium
Xem tất cả →
C1
Khoa học
Karaoke
54:08
How One Company Secretly Poisoned The Planet
Veritasium
2
C1
Khoa học
Karaoke
55:00
The World's Most Important Machine
Veritasium
C1
Khoa học
Karaoke
33:39
How a Student's Question Saved This NYC Skyscraper
Veritasium
C1
Khoa học
Karaoke
46:58
Exposing Why Farmers Can't Legally Replant Their Own Seeds
Veritasium
C1
Khoa học
Karaoke
33:00
Something Strange Happens When You Trust Quantum Mechanics
Veritasium
C1
Khoa học
Karaoke
22:17
Why Do Escalator Steps Have Teeth?
Veritasium
C1
Khoa học
Karaoke
44:15
There Is Something Faster Than Light
Veritasium
C1
Khoa học
Karaoke
53:00
The Internet Was Weeks Away From Disaster and No One Knew
Veritasium
Video gợi ý
B1
Khoa học
Karaoke
26:56
Egg Drop From Space
Mark Rober
96
B1
Khoa học
18:05
Beating 5 Scam Arcade Games with Science
Mark Rober
4
B1
Khoa học
Karaoke
19:00
Mark Rober vs Dude Perfect- Ultimate Robot Battle
Mark Rober
4
B1
Khoa học
Karaoke
17:13
Octopus vs Underwater Maze
Mark Rober
4
B1
Khoa học
24:20
This Ball is Impossible to Hit
Mark Rober
2
B1
Khoa học
20:37
Acid vs Lava- Testing Liquids That Melt Everything
Mark Rober
B2
Khoa học
28:49
The Power of Suggestion
Vsauce
2
B1
Khoa học
19:16
My Rock, Paper, Scissors Robot Never Loses (+9 Other Inventions)
Mark Rober