1 00:00:01,340 --> 00:00:04,940 In this video, I am going to talk about four technical 2 00:00:05,000 --> 00:00:08,880 why I believe super intelligence may be possible by 2030. 3 00:00:08,890 --> 00:00:11,940 The first reason is super intelligence via speed. 4 00:00:11,960 --> 00:00:15,440 Then we have model scaling, then we have RL scaling on top of 5 00:00:15,500 --> 00:00:19,410 transformers. Lastly, we have RL scaling, uh, outside of 6 00:00:19,500 --> 00:00:23,420 transformers. The 7 00:00:23,460 --> 00:00:27,420 third reason you might be convinced is more of a speculative argument. 8 00:00:27,440 --> 00:00:31,320 There isn't a lot of actual data backing this, but this is something I 9 00:00:31,360 --> 00:00:35,220 believe. Uh, so today's, uh, AI models 10 00:00:35,540 --> 00:00:38,340 can do tasks that can be done quickly. 11 00:00:38,380 --> 00:00:41,160 Like, if you can do a task that will take, you know, one minute or 10 minutes, you 12 00:00:41,200 --> 00:00:45,040 know the AI can do it for you. If you want to do a task 13 00:00:45,060 --> 00:00:47,960 know, one hour of your time and lots of failed attempts, 14 00:00:48,020 --> 00:00:51,980 probably not do it for you. Uh, however, the tasks 15 00:00:52,000 --> 00:00:54,810 which it can do, those it can do way faster than you. 16 00:00:54,840 --> 00:00:58,070 Like, if it will take, you know, 10 minutes, the AI will do it in, 17 00:00:58,140 --> 00:01:02,120 seconds. Uh, AIs operate, at least as 18 00:01:02,160 --> 00:01:05,760 of today, between 10 to, you know, 100 times the speed at which you are 19 00:01:05,880 --> 00:01:07,500 operating. 20 00:01:09,020 --> 00:01:12,960 Also, the time horizon of the tasks 21 00:01:13,040 --> 00:01:16,880 is increasing. So, you know, AI six months ago could do tasks 22 00:01:16,940 --> 00:01:17,320 that were 23 00:01:18,240 --> 00:01:21,460 maybe one minute long and the AI today can do tasks that are, 24 00:01:21,500 --> 00:01:24,080 long. There is a trend that is increasing in that direction. 25 00:01:24,100 --> 00:01:28,020 The technical reason behind this is, uh, increasing 26 00:01:28,060 --> 00:01:31,550 compute being given to reinforcement learning on top of transformers. 27 00:01:32,060 --> 00:01:35,480 And that's something I'll get to later, but this is specific technical reason why 28 00:01:35,520 --> 00:01:39,380 this is happening. Uh, my view of intelligence, at 29 00:01:39,420 --> 00:01:39,660 least 30 00:01:40,480 --> 00:01:44,440 to some extent, is that you get better at a thing by 31 00:01:44,480 --> 00:01:47,480 doing it a lot of times. You know, you get better at giving speeches by 32 00:01:47,500 --> 00:01:50,300 lot of speeches. You get better at writing code by writing 33 00:01:50,360 --> 00:01:53,100 code. You get better 34 00:01:54,360 --> 00:01:57,340 at, you know, producing movies by producing lots of 35 00:01:57,660 --> 00:02:00,560 And if you can run faster, you can run more 36 00:02:00,640 --> 00:02:04,440 iterations. Uh, all scientific discovery, in my 37 00:02:04,480 --> 00:02:08,180 view, works like this. Uh, every lab 38 00:02:08,220 --> 00:02:11,880 experiment you run is yet another iteration 39 00:02:11,940 --> 00:02:15,800 system. You know, it could be microscopes 40 00:02:15,840 --> 00:02:19,440 biology, it could be, uh, you know, 41 00:02:19,460 --> 00:02:22,740 telescopes that help you study stars. 42 00:02:22,760 --> 00:02:26,720 It could be any field. And if you can run more iterations, 43 00:02:26,760 --> 00:02:30,400 you can make progress faster. Uh, there are some people who 44 00:02:30,440 --> 00:02:34,380 think that the amount of AI-driven scientific progress 45 00:02:34,420 --> 00:02:36,720 we can get is limited because 46 00:02:37,580 --> 00:02:41,480 the physical world is slow. You know, labs are slow, labs are expensive, running 47 00:02:41,540 --> 00:02:45,440 experiments takes time. And this is true. 48 00:02:45,480 --> 00:02:49,100 However, A, there are domains where the labs are cheap. 49 00:02:49,220 --> 00:02:49,410 Uh, 50 00:02:52,340 --> 00:02:54,660 learning how to get good speeches is very cheap. 51 00:02:54,720 --> 00:02:57,420 Learning how to get good at math and software is cheap. 52 00:02:57,520 --> 00:03:00,200 Uh, like, it does not cost a lot of money. 53 00:03:00,210 --> 00:03:00,820 And I think 54 00:03:01,800 --> 00:03:05,120 that is partly why we have gotten a lot of good capabilities 55 00:03:05,460 --> 00:03:08,290 from reinforcement learning in math and software. 56 00:03:08,300 --> 00:03:11,820 Like, AI is extremely good at math and software in particular as of 57 00:03:11,940 --> 00:03:15,880 2025. I think this also carries over to 58 00:03:15,960 --> 00:03:19,320 other fields of scientific discovery. 59 00:03:19,360 --> 00:03:22,980 Maybe the speed up is not that much, but there will be a 60 00:03:23,060 --> 00:03:26,920 speed up. Uh, you can still get 61 00:03:26,960 --> 00:03:30,609 super intelligence this way. If, you know, humanity would 62 00:03:30,640 --> 00:03:33,800 otherwise have taken, you know, 100 years to, you know, invent a certain 63 00:03:33,840 --> 00:03:35,600 technology, if AI can 64 00:03:36,540 --> 00:03:39,760 do that in 10 years, that is still super intelligence. 65 00:03:39,780 --> 00:03:43,640 Like, imagine if you were sitting in 1800 and you could get 66 00:03:43,700 --> 00:03:47,580 entire 1800 to 1900, you know, technological inventions to you in 67 00:03:47,620 --> 00:03:51,440 10 years. Or imagine you were in 1900, you know, 68 00:03:51,460 --> 00:03:55,300 before World War I, you know, people used sticks and stones to literally 69 00:03:55,340 --> 00:03:59,180 fight. And by the end of the century people used, 70 00:03:59,260 --> 00:04:02,300 and, you know, space satellites and, 71 00:04:03,740 --> 00:04:06,600 uh, bomber planes and so on. 72 00:04:08,620 --> 00:04:12,540 If you can speed that up, you can compress the amount of time 73 00:04:12,560 --> 00:04:16,100 requires to invent things, that gets you super 74 00:04:16,140 --> 00:04:18,640 intelligence, and that is my third 75 00:04:18,860 --> 00:04:22,760 reason. The fourth 76 00:04:22,800 --> 00:04:26,720 reason I believe we might get super intelligence in the next five years 77 00:04:26,960 --> 00:04:30,800 is, uh, model scaling. This is a well-known trend among 78 00:04:30,820 --> 00:04:34,440 machine learning researchers. However, if you are not very in this space, you may 79 00:04:34,480 --> 00:04:38,240 not be aware of it. Uh, most of the AI progress in the last 80 00:04:38,280 --> 00:04:42,170 six years from 2019 to 2025 has been driven by one trend and 81 00:04:42,200 --> 00:04:46,000 one trend only, which is you take a transformer, you give it 82 00:04:46,060 --> 00:04:49,620 more data, more computation, and it will produce better 83 00:04:49,680 --> 00:04:53,260 results. This is magic. People don't know how this happens, 84 00:04:53,280 --> 00:04:56,380 This predictability repeatedly happens. 85 00:04:56,520 --> 00:04:59,500 Uh, the two ingredients are compute and data. 86 00:04:59,520 --> 00:05:03,460 Uh, data is just the entire internet's data 87 00:05:03,520 --> 00:05:06,940 model. Compute is just more computers that people have 88 00:05:06,980 --> 00:05:10,080 purchased. Uh, there's a special type of computer used. 89 00:05:10,100 --> 00:05:12,540 They're called graphics processing units. 90 00:05:12,560 --> 00:05:16,000 They are good at this operation of matrix multiplication in 91 00:05:16,040 --> 00:05:19,220 particular. But basically, it's just more computers. 92 00:05:19,260 --> 00:05:20,280 We have gone from 93 00:05:21,360 --> 00:05:24,420 spending, you know, like, 94 00:05:24,430 --> 00:05:28,040 $100,000, which could be literally just one machine 95 00:05:28,100 --> 00:05:31,980 person's room, to, you know, $100 billion 96 00:05:32,100 --> 00:05:36,020 worth of compute to train a single AI model that is, you know, an 97 00:05:36,100 --> 00:05:39,620 entire data center spanning, you know, thousands of 98 00:05:39,700 --> 00:05:40,190 acres. 99 00:05:41,880 --> 00:05:45,240 This scale up in, you know, six orders of magnitude 100 00:05:45,780 --> 00:05:49,740 has just been driven, because every time we notice we use more 101 00:05:49,780 --> 00:05:52,260 more data, we get more AI performance. 102 00:05:52,960 --> 00:05:56,260 We can now even predict this up to three decimal places. 103 00:05:56,280 --> 00:05:59,420 There is a law called the Chinchilla Scaling Law. 104 00:05:59,520 --> 00:06:02,657 Uh, I'll put that on the screen right now.... 105 00:06:02,668 --> 00:06:06,548 that N is the number of parameters, like how big the model 106 00:06:06,608 --> 00:06:09,828 is, and D is how much compute we give to it. 107 00:06:09,888 --> 00:06:10,488 Assume 108 00:06:11,428 --> 00:06:15,407 we s- show the model every token in the data 109 00:06:15,548 --> 00:06:19,188 exactly once. This L gives you the training loss. 110 00:06:19,227 --> 00:06:22,428 This is a proxy for how good the model 111 00:06:22,587 --> 00:06:23,248 performs. 112 00:06:25,087 --> 00:06:28,828 There is one important metric here that's really important, which is 113 00:06:28,907 --> 00:06:32,788 that the training loss correlates with your actual performance 114 00:06:32,847 --> 00:06:36,668 that you care about, but it is not the performance itself. 115 00:06:36,727 --> 00:06:40,628 It is possible for training loss to reduce and for capabilities to not increase 116 00:06:40,668 --> 00:06:44,147 that much. It is possible for training loss to reduce 117 00:06:44,188 --> 00:06:48,008 capabilities to increase a lot. Predicting this 118 00:06:48,068 --> 00:06:51,888 relationship is hard between training loss and capabilities, 119 00:06:51,948 --> 00:06:54,788 but there is a relationship. More training loss does mean more 120 00:06:54,808 --> 00:06:58,788 capabilities. Most AI researchers across the field have 121 00:06:58,847 --> 00:07:02,727 failed to predict exactly what training loss leads to exactly 122 00:07:02,768 --> 00:07:06,727 what capabilities. There are few people who have predicted it, 123 00:07:06,808 --> 00:07:10,568 and these people are, for example, Ilya Sutskever at OpenAI was very 124 00:07:10,587 --> 00:07:14,428 bullish on this, and he predicts, you know, we will get 125 00:07:14,448 --> 00:07:17,347 super intelligence in the next 5 to 10 years. 126 00:07:18,168 --> 00:07:21,988 People who have actually believed in this law have generally predicted that AI will 127 00:07:22,048 --> 00:07:24,647 come sooner than what most other people in the field 128 00:07:24,788 --> 00:07:29,668 predict. 129 00:07:31,828 --> 00:07:35,727 Uh, we still have at least two orders of 130 00:07:35,768 --> 00:07:39,607 up left. Uh, we have spent, like I said, you know, $10 billion has 131 00:07:39,647 --> 00:07:43,448 already been spent and the next set of la- data centers we're making are being 132 00:07:43,488 --> 00:07:46,628 made so that we spend $100 trillion on the training of a single 133 00:07:46,707 --> 00:07:50,628 model. We might scale that up another order of magnitude, 134 00:07:50,688 --> 00:07:52,308 trillion dollars on a single model. 135 00:07:52,688 --> 00:07:56,548 It seems likely we will not go beyond a trillion dollars on a single model 136 00:07:57,268 --> 00:08:00,168 unless we get, you know, really exceptional capabilities. 137 00:08:00,227 --> 00:08:04,147 This alone means, you know, two more orders of magnitude from 10 138 00:08:04,248 --> 00:08:08,128 trillion. Which you can compare to, you know, the six 139 00:08:08,168 --> 00:08:10,448 orders of magnitude they have already gone through. 140 00:08:10,467 --> 00:08:13,948 There are some people who say, you know, the latest model, which was 141 00:08:13,987 --> 00:08:17,048 GPT-4.5, was not that 142 00:08:17,207 --> 00:08:21,048 impressive, meaning, you know, this law might be weakening, and that 143 00:08:21,107 --> 00:08:25,028 might be true, but just remember we still have two more 144 00:08:27,268 --> 00:08:30,908 The fifth reason you might be convinced super intelligence might be coming in the 145 00:08:30,948 --> 00:08:34,308 next five years is reinforcement learning scaling. 146 00:08:34,328 --> 00:08:37,188 Reinforcement learning is an old technique in machine learning. 147 00:08:37,227 --> 00:08:41,087 However, it has been applied to transformers in 148 00:08:41,127 --> 00:08:44,928 Literally only one year ago we figured out how 149 00:08:44,968 --> 00:08:48,508 to apply reinforcement learning on a transformer to get 150 00:08:48,568 --> 00:08:52,188 even more capabilities. Uh, reinforcement 151 00:08:52,248 --> 00:08:56,168 learning allows the model to try something, fail, try something, fail, 152 00:08:56,208 --> 00:08:58,887 try again, and until it gets the correct answer. 153 00:08:59,008 --> 00:09:02,887 Uh, this is different from the base model 154 00:09:02,928 --> 00:09:05,288 know, cached answer which it gives out. 155 00:09:05,308 --> 00:09:08,807 Reinforcement learning allows us to spend more compute on a single 156 00:09:08,948 --> 00:09:12,768 task. Uh, whereas earlier it would just give the 157 00:09:12,808 --> 00:09:14,088 answer for the task. 158 00:09:15,068 --> 00:09:18,908 There are multiple, you know, tasks which people consider 159 00:09:18,948 --> 00:09:22,848 difficult, which, you know, AI would never crack, you know, as of 2023 160 00:09:22,887 --> 00:09:26,808 and 2024. Uh, researchers said that 161 00:09:26,948 --> 00:09:28,948 certain benchmarks would not be cracked. 162 00:09:28,968 --> 00:09:29,867 People said, you know, 163 00:09:30,848 --> 00:09:32,938 humanity's last exam would not be cracked. 164 00:09:32,968 --> 00:09:35,308 People said, you know, ARC AGI would not be cracked. 165 00:09:35,348 --> 00:09:38,488 People said, you know, mathematics olympiads would not be 166 00:09:38,568 --> 00:09:40,728 All three of these have been cracked. 167 00:09:40,828 --> 00:09:44,627 Uh, humanity's last exam consists of lot of 168 00:09:44,668 --> 00:09:48,328 answers by PhDs in various fields of both science 169 00:09:48,568 --> 00:09:51,268 and social science. This includes biology, chemistry, 170 00:09:51,387 --> 00:09:55,348 archeology, uh, journalism, law and so on. 171 00:09:55,448 --> 00:09:59,028 Uh, all these PhDs agreed that, you know, if these questions get 172 00:09:59,088 --> 00:10:03,028 cracked, that means the model is as good as a PhD in those fields, 173 00:10:03,088 --> 00:10:05,928 seeing AI is already solved like almost half of that data 174 00:10:05,988 --> 00:10:09,877 set. We have 175 00:10:10,028 --> 00:10:13,137 started seeing a loss to predict, you know, 176 00:10:14,448 --> 00:10:18,387 how much computation leads to how much performance. 177 00:10:18,448 --> 00:10:22,428 Like where we had with model scaling where we have a very clear empirical law which 178 00:10:22,548 --> 00:10:26,407 gives you up to three decimal places accuracy predictions. 179 00:10:26,528 --> 00:10:30,328 Uh, we do not have as accurate predictions for 180 00:10:30,407 --> 00:10:33,718 scaling of reinforcement learning because it is relatively new. 181 00:10:33,768 --> 00:10:36,848 We have less data points, we have tried this for less longer. 182 00:10:36,867 --> 00:10:40,728 However, people are already starting to make some 183 00:10:40,768 --> 00:10:42,127 are starting to see that 184 00:10:43,088 --> 00:10:47,068 every time we increase the compute by, you know, 10x, we get a 185 00:10:47,188 --> 00:10:51,068 little bit of improvement. So if you spend, you know, $10 per task, 186 00:10:51,147 --> 00:10:54,668 spend $100, then you spend $1,000 per task, 187 00:10:54,788 --> 00:10:56,428 improvement, little improvement. 188 00:10:56,468 --> 00:11:00,147 We have gone all the way, at least as per public data of OpenAI, 189 00:11:00,168 --> 00:11:03,948 the way up to $500,000 per task, and we also know 190 00:11:04,008 --> 00:11:07,748 that, uh, AI has been able to get, you know, the gold 191 00:11:07,808 --> 00:11:11,708 medal on the International Mathematics Olympiad, which is like the olympiad 192 00:11:11,788 --> 00:11:15,617 for mathematics worldwide. We still have many, 193 00:11:15,647 --> 00:11:18,718 many, many more orders of magnitude to go from, 194 00:11:18,718 --> 00:11:22,348 $100,000 and in theory what we can 195 00:11:22,448 --> 00:11:25,588 spend is, which is at least a billion dollars per 196 00:11:25,627 --> 00:11:28,998 Like if you were really solving, you know, novel physics, novel mathematics, you 197 00:11:29,008 --> 00:11:32,887 know, novel chemistry. If you were, you know, trying to solve something like 198 00:11:32,907 --> 00:11:36,407 cancer or trying to, you know, solve math theorems that have, you know, 199 00:11:36,448 --> 00:11:40,147 unsolved for, you know, the past century, like, why would you not spend a billion 200 00:11:40,188 --> 00:11:44,168 dollars on this task? So we have many more orders 201 00:11:44,188 --> 00:11:48,088 of magnitude to go here. We have a lot more compute we're going to 202 00:11:48,488 --> 00:11:52,428 and we don't know how far this law scales, but we are going to find out very 203 00:11:52,508 --> 00:11:55,147 soon. 204 00:11:57,348 --> 00:12:00,688 The sixth and final reason you might be convinced super intelligence might be 205 00:12:00,748 --> 00:12:04,627 coming by 2030 is that reinforcement learning has worked 206 00:12:04,668 --> 00:12:07,928 before. Like I said, this is a not new technique. 207 00:12:07,968 --> 00:12:11,668 We are applying it to transformers newly because transformers can, you know, 208 00:12:11,688 --> 00:12:15,528 understand English and generalize across so many, 209 00:12:15,548 --> 00:12:19,387 tasks. Whereas earlier before transformers we 210 00:12:19,407 --> 00:12:22,948 doing specific things. Even when we just had narrow models doing 211 00:12:22,968 --> 00:12:25,328 specific things, like we had just one model and 212 00:12:25,338 --> 00:12:26,068 (laughs) 213 00:12:26,088 --> 00:12:29,028 ... one task. Even back then we knew that reinforcement learning actually works 214 00:12:29,068 --> 00:12:32,728 pretty well. Uh, we have, uh, an AI 215 00:12:32,768 --> 00:12:35,988 model that is, you know, the world champion in chess 216 00:12:36,008 --> 00:12:39,897 chess. Like no chess experts had to tell 217 00:12:39,928 --> 00:12:43,348 it's the best at chess. Uh, this is obviously by 218 00:12:43,387 --> 00:12:46,387 DeepMind. This is, uh, AlphaZero. 219 00:12:46,448 --> 00:12:46,668 We 220 00:12:47,768 --> 00:12:51,708 also have DeepMind's model that beat the best 221 00:12:51,788 --> 00:12:55,548 Go player in the world. This happened again before transformers 222 00:12:55,588 --> 00:12:59,228 Like this is very old news, but we have gotten superhuman 223 00:12:59,268 --> 00:13:02,968 performance in narrow domains. You know, poker has been solved. 224 00:13:03,008 --> 00:13:06,268 Dota has been solved. StarCraft has been solved. 225 00:13:06,308 --> 00:13:10,228 People, AI researchers moved on from games because 226 00:13:10,248 --> 00:13:13,778 They found it too easy to make an AI that will just completely obliterate all 227 00:13:13,808 --> 00:13:17,588 human players in a game. And that is why we are now 228 00:13:17,608 --> 00:13:21,528 working with general domains. So we know point number A 229 00:13:21,608 --> 00:13:25,048 on narrow domains, reinforcement learning gets superhuman 230 00:13:25,068 --> 00:13:28,867 Point number B, transformers can work on general domains. 231 00:13:28,907 --> 00:13:32,788 Point number C, reinforcement learning plus transformers 232 00:13:32,808 --> 00:13:36,768 domains actually works. Point number D, this thing scales up. 233 00:13:36,808 --> 00:13:38,627 Point number E, we are going to scale it 234 00:13:38,748 --> 00:13:41,147 up.