TryShadowing
Shadow YouTube. Nói tiếng Anh.
Trang chủ
Khám phá
Chép chính tả
NEW
Thư viện
vi
0
ngày
Đăng nhập
Google just casually disrupted… — Fireship luyện shadowing | TryShadowing
TryShadowing
Shadow YouTube. Nói tiếng Anh.
Trang chủ
Khám phá
Chép chính tả
NEW
Thư viện
vi
0
ngày
Đăng nhập
Trang chủ
Khám phá
Fireship
Google just casually disrupted the open-source AI narrative…
Google just casually disrupted the open-source AI narrative…
Fireship
·
5:15 · 8 thg 4, 2026
Bắt đầu học
0:00
0:00
Ghi âm
×1
1x
VI
EN
JA
KO
ZH
FR
PT
TH
IT
DE
IPA
Chấm điểm phát âm chưa hỗ trợ trên trình duyệt này — bạn vẫn ghi âm & nghe lại được.
Last
week,
Google
did
something
Đang dịch…
Bật Ghi âm để được thu giọng và chấm điểm
Thông minh
Karaoke
Câu
1
/116
0:00
Last week, Google did something
0:01
that no other Fang company has had the balls to do.
0:04
They released a large language model
0:05
that qualifies as truly free
0:07
and open source under the Apache 2.0 license.
0:11
That means free as in total freedom, not openish, not research only,
0:15
not please don't make money or we'll sue you.
0:17
That model is Gemma 4, and my initial thought was, "Oh great,
0:21
another half-baked open model that's technically free
0:24
as long as you also own a small data center to run it."
0:26
But the craziest thing about Gemma 4 is that it's small.
0:29
Like suspiciously small.
0:31
The big model is small enough to run on a consumer GPU,
0:34
and the edge model is small enough to run on your phone
0:36
or Raspberry Pi while hitting intelligence levels
0:39
that are on par with other open models
0:41
that would normally require data center caliber GPUs just to run.
0:45
That shouldn't be possible.
0:46
And in today's video,
0:47
we'll find out how it works
0:49
and look at some other crazy compression techniques developed by Google.
0:52
It is April 8th, 2026, and you're watching The Code Report.
0:55
To be fair, several other companies in the game
0:58
and family have released open weight models.
1:00
Like Meta's Llama models are quasi-free and open,
1:03
but under a special license
1:04
that gives Meta leverage to any developer
1:06
that actually starts printing cash with them.
1:08
Then we have OpenAI's GPT-OSS models, which are also Apache 2.0 licensed,
1:14
but they're bigger and dumber than Gemma.
1:16
Outside of that, we basically rely on Mistral and the Chinese models like Qwen,
1:20
GLM, Kimmi, and DeepSeek.
1:22
Gemma 4 hits different though because it's made in America, Apache 2.0 licensed,
1:26
intelligent, and most importantly, tiny.
1:29
For comparison, the 31 billion parameter version of Gemma 4 is scoring in the
1:33
same ballpark as models like Kimmi K 2.5 thinking.
1:37
But here's the absurd part.
1:38
I can run Gemma 4 locally with a 20 GB download getting roughly 10
1:43
tokens per second on a single RTX 4090.
1:46
But if I wanted to run Kimmi K 2.5,
1:48
I'd be looking at a 600-plus GB download,
1:51
at least 256 GB of RAM, aggressive quantization,
1:55
and multiple H100s just to get it off the ground.
1:58
Is Kimmi still a better model than Gemma?
2:00
But there's no way in hell I'm going to run it locally.
2:02
So the obvious question is how did Google achieve this unbelievable shrinkage?
2:06
Well, the answer is they didn't just shrink the model.
2:09
They attacked the real bottleneck in AI, memory.
2:12
To run a massive large language model locally, you don't need a better CPU.
2:15
You need more memory bandwidth.
2:17
Every time a model generates a token,
2:19
it has to read through a massive amount of model weights in VRAM,
2:22
which is the video random access memory on your GPU.
2:25
It doesn't really matter how big the model is.
2:27
It's more about how expensive it is to read it.
2:30
And this is where things get interesting because alongside Gemma 4,
2:33
Google quietly dropped a research note on something called Turboquant,
2:37
which sounds like a marketing buzzword, but it's actually kind of insane.
2:40
It's a new approach to quantization,
2:42
which is the process of compressing model weights so they take up less space.
2:45
Normally, through this process, you get a simple tradeoff,
2:48
a smaller model but worse performance.
2:50
But Turboquant improves this tradeoff with two steps.
2:53
First, it compresses data that's normally in an XYZ Cartesian coordinate system into polar
2:59
coordinates that include a radius
3:00
and angle.
3:01
Because these angles follow a predictable pattern,
3:04
the model can skip the typical normalization steps and store information more efficiently,
3:08
thus reducing memory overhead.
3:10
Then it uses this mathematical technique called the Johnson-Lindenstrauss transform to shrink high-dimensional data
3:17
by compressing it down to single sign bits,
3:19
positive one, negative one, while preserving the distances between these data points.
3:23
Frankly, I'm too stupid to understand how the math actually works,
3:27
but Turboquant is actually not the secret behind Gemma 4's small models.
3:31
You'll notice that some of the Gemma 4 models have an E in the
3:33
model name like E2B
3:35
and E4B.
3:36
And what that stands for is effective parameters
3:39
because these models incorporate something called per-layer embeddings,
3:43
which is like giving every layer in the neural network its own mini cheat
3:46
sheet for each token.
3:47
In a normal transformer, each token gets one embedding at the start,
3:51
and the model has to carry that information through every layer.
3:54
And most of that information isn't needed.
3:56
Per-layer embeddings changes that by giving each layer its own small custom version of
4:00
the token.
4:01
So information can be introduced exactly when it's useful instead of all at once.
4:05
There's an incredible visual guide by Martin Groothuis
4:08
that I'll link in the description
4:09
if you want to dive into more detail.
4:11
The end result is a small, smart, and efficient model.
4:14
I'm running it here with Ollama on my RTX 4090,
4:17
and my initial impression is that it's a solid all-around model,
4:20
and it would also be a great model for fine-tuning with your own data
4:23
using tools like Unsloth.
4:24
But if you're a programmer,
4:26
it's still not good enough to replace any high-end coding tools like Code Rabbit,
4:30
the sponsor of today's video.
4:31
They just launched a CLI update
4:33
that lets it review all the code your agent writes,
4:36
then tells it exactly how to fix any bugs it finds.
4:38
You can enable this with a new {dash} {dash} agent flag,
4:41
which turns Code Rabbit into a tool your agent can call directly.
4:45
From there, it'll give your agent structured JSON with all of the issues plus
4:49
instructions on how to fix them
4:51
so your agent can go back
4:52
and clean everything up before it opens up a pull request.
4:55
They also simplified the setup process
4:57
and removed their rate limits
4:59
so you can get started with a single terminal command
5:01
and run as many reviews
5:02
as your agents need.
5:03
Try it out for free today using the Code Rabbit off login command
5:07
and use it free forever on any open source project.
5:10
This has been The Code Report.
5:11
Thanks for watching, and I will see you in the next one.
Thích
Chia sẻ
Fireship
Xem tất cả →
C1
Công nghệ
Karaoke
7:22
Tragic mistake... Anthropic leaks Claude’s source code
Fireship
1
C1
Công nghệ
Karaoke
5:18
The wild rise of OpenClaw
Fireship
C1
Công nghệ
Karaoke
4:50
Google just changed the future of UI/UX design...
Fireship
C1
Công nghệ
Karaoke
9:10
The unhinged world of tech in 2026...
Fireship
C1
Công nghệ
Karaoke
10:03
10 open source tools that feel illegal...
Fireship
C1
Công nghệ
Karaoke
5:36
He just crawled through hell to fix the browser…
Fireship
C1
Công nghệ
Karaoke
5:36
Claude Mythos is too dangerous for public consumption...
Fireship
C1
Công nghệ
Karaoke
5:00
Anthropic just released the real Claude Bot...
Fireship
Video gợi ý
B2
Công nghệ
Karaoke
17:24
Driving Xiaomi's Electric Car: Are we Cooked?
Marques Brownlee
B2
Công nghệ
Karaoke
8:47
Macbook Neo Impressions: Reincarnated!
Marques Brownlee
B2
Công nghệ
Karaoke
11:51
Xiaomi 17 Pro Max: An iPhone... But Better!
Marques Brownlee
B2
Công nghệ
Karaoke
16:11
The Problem with this Humanoid Robot
Marques Brownlee
B2
Công nghệ
Karaoke
10:47
So This is Peak Foldable
Marques Brownlee
B2
Công nghệ
Karaoke
12:52
OnePlus 15 Review: This is Not Normal!
Marques Brownlee
B2
Công nghệ
Karaoke
32:38
Smartphone Awards 2025!
Marques Brownlee
B2
Công nghệ
Karaoke
12:53
Macbook Neo Review: Better than you Think!
Marques Brownlee