WEBVTT

00:01.340 --> 00:04.940
In this video,
I am going to talk about four technical

00:05.000 --> 00:08.880
why I believe super intelligence may be
possible by 2030.

00:08.890 --> 00:11.940
The first reason is super intelligence via
speed.

00:11.960 --> 00:15.440
Then we have model scaling,
then we have RL scaling on top of

00:15.500 --> 00:19.410
transformers. Lastly, we have RL scaling,
uh, outside of

00:19.500 --> 00:23.420
transformers. The

00:23.460 --> 00:27.420
third reason you might be convinced
is more of a speculative argument.

00:27.440 --> 00:31.320
There isn't a lot of actual data backing
this, but this is something I

00:31.360 --> 00:35.220
believe. Uh, so today's, uh, AI models

00:35.540 --> 00:38.340
can do tasks that can be done quickly.

00:38.380 --> 00:41.160
Like, if you can do a task that will take,
you know, one minute or 10 minutes, you

00:41.200 --> 00:45.040
know the AI can do it for you.
If you want to do a task

00:45.060 --> 00:47.960
know, one hour of your time
and lots of failed attempts,

00:48.020 --> 00:51.980
probably not do it for you. Uh, however,
the tasks

00:52.000 --> 00:54.810
which it can do,
those it can do way faster than you.

00:54.840 --> 00:58.070
Like, if it will take, you know,
10 minutes, the AI will do it in,

00:58.140 --> 01:02.120
seconds. Uh, AIs operate, at least as

01:02.160 --> 01:05.760
of today, between 10 to, you know,
100 times the speed at which you are

01:05.880 --> 01:07.500
operating.

01:09.020 --> 01:12.960
Also, the time horizon of the tasks

01:13.040 --> 01:16.880
is increasing. So, you know,
AI six months ago could do tasks

01:16.940 --> 01:17.320
that were

01:18.240 --> 01:21.460
maybe one minute long
and the AI today can do tasks that are,

01:21.500 --> 01:24.080
long. There is a trend that
is increasing in that direction.

01:24.100 --> 01:28.020
The technical reason behind this is, uh,
increasing

01:28.060 --> 01:31.550
compute being given to reinforcement
learning on top of transformers.

01:32.060 --> 01:35.480
And that's something I'll get to later,
but this is specific technical reason why

01:35.520 --> 01:39.380
this is happening. Uh,
my view of intelligence, at

01:39.420 --> 01:39.660
least

01:40.480 --> 01:44.440
to some extent, is
that you get better at a thing by

01:44.480 --> 01:47.480
doing it a lot of times. You know,
you get better at giving speeches by

01:47.500 --> 01:50.300
lot of speeches.
You get better at writing code by writing

01:50.360 --> 01:53.100
code. You get better

01:54.360 --> 01:57.340
at, you know,
producing movies by producing lots of

01:57.660 --> 02:00.560
And if you can run faster,
you can run more

02:00.640 --> 02:04.440
iterations. Uh, all scientific discovery,
in my

02:04.480 --> 02:08.180
view, works like this. Uh, every lab

02:08.220 --> 02:11.880
experiment you run
is yet another iteration

02:11.940 --> 02:15.800
system. You know,
it could be microscopes

02:15.840 --> 02:19.440
biology, it could be, uh, you know,

02:19.460 --> 02:22.740
telescopes that help you study stars.

02:22.760 --> 02:26.720
It could be any field. And
if you can run more iterations,

02:26.760 --> 02:30.400
you can make progress faster. Uh,
there are some people who

02:30.440 --> 02:34.380
think that the amount of AI-driven
scientific progress

02:34.420 --> 02:36.720
we can get is limited because

02:37.580 --> 02:41.480
the physical world is slow. You know,
labs are slow, labs are expensive, running

02:41.540 --> 02:45.440
experiments takes time. And this is true.

02:45.480 --> 02:49.100
However, A, there
are domains where the labs are cheap.

02:49.220 --> 02:49.410
Uh,

02:52.340 --> 02:54.660
learning how to get good speeches
is very cheap.

02:54.720 --> 02:57.420
Learning how to get good at math
and software is cheap.

02:57.520 --> 03:00.200
Uh, like, it does not cost a lot of money.

03:00.210 --> 03:00.820
And I think

03:01.800 --> 03:05.120
that is partly why we have gotten a lot of
good capabilities

03:05.460 --> 03:08.290
from reinforcement learning in math
and software.

03:08.300 --> 03:11.820
Like, AI is extremely good at math
and software in particular as of

03:11.940 --> 03:15.880
2025. I think this also carries over to

03:15.960 --> 03:19.320
other fields of scientific discovery.

03:19.360 --> 03:22.980
Maybe the speed up is not that much,
but there will be a

03:23.060 --> 03:26.920
speed up. Uh, you can still get

03:26.960 --> 03:30.609
super intelligence this way. If, you know,
humanity would

03:30.640 --> 03:33.800
otherwise have taken, you know,
100 years to, you know, invent a certain

03:33.840 --> 03:35.600
technology, if AI can

03:36.540 --> 03:39.760
do that in 10 years, that
is still super intelligence.

03:39.780 --> 03:43.640
Like, imagine if you
were sitting in 1800 and you could get

03:43.700 --> 03:47.580
entire 1800 to 1900, you know,
technological inventions to you in

03:47.620 --> 03:51.440
10 years. Or imagine you were in 1900,
you know,

03:51.460 --> 03:55.300
before World War I, you know,
people used sticks and stones to literally

03:55.340 --> 03:59.180
fight.
And by the end of the century people used,

03:59.260 --> 04:02.300
and, you know, space satellites and,

04:03.740 --> 04:06.600
uh, bomber planes and so on.

04:08.620 --> 04:12.540
If you can speed that up,
you can compress the amount of time

04:12.560 --> 04:16.100
requires to invent things,
that gets you super

04:16.140 --> 04:18.640
intelligence, and that is my third

04:18.860 --> 04:22.760
reason. The fourth

04:22.800 --> 04:26.720
reason I believe we might get super
intelligence in the next five years

04:26.960 --> 04:30.800
is, uh, model scaling. This
is a well-known trend among

04:30.820 --> 04:34.440
machine learning researchers. However,
if you are not very in this space, you may

04:34.480 --> 04:38.240
not be aware of it. Uh,
most of the AI progress in the last

04:38.280 --> 04:42.170
six years from 2019 to 2025 has been
driven by one trend and

04:42.200 --> 04:46.000
one trend only, which
is you take a transformer, you give it

04:46.060 --> 04:49.620
more data, more computation,
and it will produce better

04:49.680 --> 04:53.260
results. This is magic.
People don't know how this happens,

04:53.280 --> 04:56.380
This predictability repeatedly happens.

04:56.520 --> 04:59.500
Uh, the two ingredients are compute
and data.

04:59.520 --> 05:03.460
Uh,
data is just the entire internet's data

05:03.520 --> 05:06.940
model. Compute is just more computers
that people have

05:06.980 --> 05:10.080
purchased. Uh,
there's a special type of computer used.

05:10.100 --> 05:12.540
They're called graphics processing units.

05:12.560 --> 05:16.000
They are good at this operation of matrix
multiplication in

05:16.040 --> 05:19.220
particular. But basically,
it's just more computers.

05:19.260 --> 05:20.280
We have gone from

05:21.360 --> 05:24.420
spending, you know, like,

05:24.430 --> 05:28.040
$100,000,
which could be literally just one machine

05:28.100 --> 05:31.980
person's room, to, you know, $100 billion

05:32.100 --> 05:36.020
worth of compute to train a single AI
model that is, you know, an

05:36.100 --> 05:39.620
entire data center spanning, you know,
thousands of

05:39.700 --> 05:40.190
acres.

05:41.880 --> 05:45.240
This scale up in, you know,
six orders of magnitude

05:45.780 --> 05:49.740
has just been driven,
because every time we notice we use more

05:49.780 --> 05:52.260
more data, we get more AI performance.

05:52.960 --> 05:56.260
We can now even predict this up to three
decimal places.

05:56.280 --> 05:59.420
There is a law called the Chinchilla
Scaling Law.

05:59.520 --> 06:02.657
Uh,
I'll put that on the screen right now....

06:02.668 --> 06:06.548
that N is the number of parameters,
like how big the model

06:06.608 --> 06:09.828
is,
and D is how much compute we give to it.

06:09.888 --> 06:10.488
Assume

06:11.428 --> 06:15.407
we s-
show the model every token in the data

06:15.548 --> 06:19.188
exactly once.
This L gives you the training loss.

06:19.227 --> 06:22.428
This is a proxy for how good the model

06:22.587 --> 06:23.248
performs.

06:25.087 --> 06:28.828
There is one important metric here that's
really important, which is

06:28.907 --> 06:32.788
that the training loss correlates with
your actual performance

06:32.847 --> 06:36.668
that you care about, but it
is not the performance itself.

06:36.727 --> 06:40.628
It is possible for training loss to reduce
and for capabilities to not increase

06:40.668 --> 06:44.147
that much.
It is possible for training loss to reduce

06:44.188 --> 06:48.008
capabilities to increase a lot.
Predicting this

06:48.068 --> 06:51.888
relationship is hard between training loss
and capabilities,

06:51.948 --> 06:54.788
but there is a relationship.
More training loss does mean more

06:54.808 --> 06:58.788
capabilities.
Most AI researchers across the field have

06:58.847 --> 07:02.727
failed to predict exactly what training
loss leads to exactly

07:02.768 --> 07:06.727
what capabilities. There
are few people who have predicted it,

07:06.808 --> 07:10.568
and these people are, for example,
Ilya Sutskever at OpenAI was very

07:10.587 --> 07:14.428
bullish on this, and he predicts,
you know, we will get

07:14.448 --> 07:17.347
super intelligence in the next 5 to 10
years.

07:18.168 --> 07:21.988
People who have actually believed in this
law have generally predicted that AI will

07:22.048 --> 07:24.647
come sooner than what most other people in
the field

07:24.788 --> 07:29.668
predict.

07:31.828 --> 07:35.727
Uh,
we still have at least two orders of

07:35.768 --> 07:39.607
up left. Uh, we have spent, like I said,
you know, $10 billion has

07:39.647 --> 07:43.448
already been spent and the next set of la-
data centers we're making are being

07:43.488 --> 07:46.628
made so that we spend $100 trillion on the
training of a single

07:46.707 --> 07:50.628
model. We might scale
that up another order of magnitude,

07:50.688 --> 07:52.308
trillion dollars on a single model.

07:52.688 --> 07:56.548
It seems likely we will not go beyond a
trillion dollars on a single model

07:57.268 --> 08:00.168
unless we get, you know,
really exceptional capabilities.

08:00.227 --> 08:04.147
This alone means, you know,
two more orders of magnitude from 10

08:04.248 --> 08:08.128
trillion. Which you can compare to,
you know, the six

08:08.168 --> 08:10.448
orders of magnitude they have already gone
through.

08:10.467 --> 08:13.948
There are some people who say, you know,
the latest model, which was

08:13.987 --> 08:17.048
GPT-4.5, was not that

08:17.207 --> 08:21.048
impressive, meaning, you know,
this law might be weakening, and that

08:21.107 --> 08:25.028
might be true,
but just remember we still have two more

08:27.268 --> 08:30.908
The fifth reason you might be convinced
super intelligence might be coming in the

08:30.948 --> 08:34.308
next five years is reinforcement learning
scaling.

08:34.328 --> 08:37.188
Reinforcement learning
is an old technique in machine learning.

08:37.227 --> 08:41.087
However,
it has been applied to transformers in

08:41.127 --> 08:44.928
Literally only one year ago we figured out
how

08:44.968 --> 08:48.508
to apply reinforcement learning on a
transformer to get

08:48.568 --> 08:52.188
even more capabilities. Uh, reinforcement

08:52.248 --> 08:56.168
learning allows the model to try
something, fail, try something, fail,

08:56.208 --> 08:58.887
try again,
and until it gets the correct answer.

08:59.008 --> 09:02.887
Uh,
this is different from the base model

09:02.928 --> 09:05.288
know, cached answer which it gives out.

09:05.308 --> 09:08.807
Reinforcement learning allows us to spend
more compute on a single

09:08.948 --> 09:12.768
task. Uh,
whereas earlier it would just give the

09:12.808 --> 09:14.088
answer for the task.

09:15.068 --> 09:18.908
There are multiple, you know, tasks
which people consider

09:18.948 --> 09:22.848
difficult, which, you know,
AI would never crack, you know, as of 2023

09:22.887 --> 09:26.808
and 2024. Uh, researchers said that

09:26.948 --> 09:28.948
certain benchmarks would not be cracked.

09:28.968 --> 09:29.867
People said, you know,

09:30.848 --> 09:32.938
humanity's last exam would not be cracked.

09:32.968 --> 09:35.308
People said, you know,
ARC AGI would not be cracked.

09:35.348 --> 09:38.488
People said, you know,
mathematics olympiads would not be

09:38.568 --> 09:40.728
All three of these have been cracked.

09:40.828 --> 09:44.627
Uh,
humanity's last exam consists of lot of

09:44.668 --> 09:48.328
answers by PhDs in various fields of both
science

09:48.568 --> 09:51.268
and social science. This includes biology,
chemistry,

09:51.387 --> 09:55.348
archeology, uh, journalism, law and so on.

09:55.448 --> 09:59.028
Uh, all these PhDs agreed that, you know,
if these questions get

09:59.088 --> 10:03.028
cracked, that means the model
is as good as a PhD in those fields,

10:03.088 --> 10:05.928
seeing AI is already solved like almost
half of that data

10:05.988 --> 10:09.877
set. We have

10:10.028 --> 10:13.137
started seeing a loss to predict,
you know,

10:14.448 --> 10:18.387
how much computation leads to how much
performance.

10:18.448 --> 10:22.428
Like where we had with model scaling where
we have a very clear empirical law which

10:22.548 --> 10:26.407
gives you up to three decimal places
accuracy predictions.

10:26.528 --> 10:30.328
Uh,
we do not have as accurate predictions for

10:30.407 --> 10:33.718
scaling of reinforcement learning because
it is relatively new.

10:33.768 --> 10:36.848
We have less data points,
we have tried this for less longer.

10:36.867 --> 10:40.728
However,
people are already starting to make some

10:40.768 --> 10:42.127
are starting to see that

10:43.088 --> 10:47.068
every time we increase the compute by,
you know, 10x, we get a

10:47.188 --> 10:51.068
little bit of improvement. So
if you spend, you know, $10 per task,

10:51.147 --> 10:54.668
spend $100,
then you spend $1,000 per task,

10:54.788 --> 10:56.428
improvement, little improvement.

10:56.468 --> 11:00.147
We have gone all the way,
at least as per public data of OpenAI,

11:00.168 --> 11:03.948
the way up to $500,000 per task,
and we also know

11:04.008 --> 11:07.748
that, uh, AI has been able to get,
you know, the gold

11:07.808 --> 11:11.708
medal on the International Mathematics
Olympiad, which is like the olympiad

11:11.788 --> 11:15.617
for mathematics worldwide.
We still have many,

11:15.647 --> 11:18.718
many,
many more orders of magnitude to go from,

11:18.718 --> 11:22.348
$100,000 and in theory what we can

11:22.448 --> 11:25.588
spend is,
which is at least a billion dollars per

11:25.627 --> 11:28.998
Like if you were really solving, you know,
novel physics, novel mathematics, you

11:29.008 --> 11:32.887
know, novel chemistry. If you were,
you know, trying to solve something like

11:32.907 --> 11:36.407
cancer or trying to, you know,
solve math theorems that have, you know,

11:36.448 --> 11:40.147
unsolved for, you know, the past century,
like, why would you not spend a billion

11:40.188 --> 11:44.168
dollars on this task?
So we have many more orders

11:44.188 --> 11:48.088
of magnitude to go here.
We have a lot more compute we're going to

11:48.488 --> 11:52.428
and we don't know how far this law scales,
but we are going to find out very

11:52.508 --> 11:55.147
soon.

11:57.348 --> 12:00.688
The sixth and final reason you might be
convinced super intelligence might be

12:00.748 --> 12:04.627
coming by 2030 is
that reinforcement learning has worked

12:04.668 --> 12:07.928
before. Like I said, this
is a not new technique.

12:07.968 --> 12:11.668
We are applying it to transformers newly
because transformers can, you know,

12:11.688 --> 12:15.528
understand English
and generalize across so many,

12:15.548 --> 12:19.387
tasks.
Whereas earlier before transformers we

12:19.407 --> 12:22.948
doing specific things. Even
when we just had narrow models doing

12:22.968 --> 12:25.328
specific things,
like we had just one model and

12:25.338 --> 12:26.068
(laughs)

12:26.088 --> 12:29.028
... one task. Even back then we knew
that reinforcement learning actually works

12:29.068 --> 12:32.728
pretty well. Uh, we have, uh, an AI

12:32.768 --> 12:35.988
model that is, you know,
the world champion in chess

12:36.008 --> 12:39.897
chess.
Like no chess experts had to tell

12:39.928 --> 12:43.348
it's the best at chess. Uh, this
is obviously by

12:43.387 --> 12:46.387
DeepMind. This is, uh, AlphaZero.

12:46.448 --> 12:46.668
We

12:47.768 --> 12:51.708
also have DeepMind's model
that beat the best

12:51.788 --> 12:55.548
Go player in the world.
This happened again before transformers

12:55.588 --> 12:59.228
Like this is very old news,
but we have gotten superhuman

12:59.268 --> 13:02.968
performance in narrow domains. You know,
poker has been solved.

13:03.008 --> 13:06.268
Dota has been solved.
StarCraft has been solved.

13:06.308 --> 13:10.228
People,
AI researchers moved on from games because

13:10.248 --> 13:13.778
They found it too easy to make an AI
that will just completely obliterate all

13:13.808 --> 13:17.588
human players in a game. And that
is why we are now

13:17.608 --> 13:21.528
working with general domains.
So we know point number A

13:21.608 --> 13:25.048
on narrow domains,
reinforcement learning gets superhuman

13:25.068 --> 13:28.867
Point number B,
transformers can work on general domains.

13:28.907 --> 13:32.788
Point number C,
reinforcement learning plus transformers

13:32.808 --> 13:36.768
domains actually works. Point number D,
this thing scales up.

13:36.808 --> 13:38.627
Point number E, we are going to scale it

13:38.748 --> 13:41.147
up.
