@dwarkesh_sp - The @JeffDean & @NoamShazeer episode. We talk about 25
Dwarkesh Patel✓@dwarkesh_sp
2025-02-12The @JeffDean & @NoamShazeer episode.
We talk about 25 years at Google, from PageRank to MapReduce to the Transformer to MoEs to AlphaChip – and soon to ASI.
My favorite part was Jeff's vision for AGI as one giant MoE that is grown in bits and pieces over time like a forest, rather than trained all at once.
Specialization, distillation, inference time scaling all emerge organically rather than by design.
Noam bites every bullet: 100x world GDP soon; let’s get a million automated researchers running in the Google datacenter; living to see the year 3000.
Links below. Enjoy!
Timestamps
0:00:00 - Intro
0:03:29 - Joining Google in 1999
0:06:20 - Future of Moore's Law
0:11:04 - Future TPUs
0:13:56 - Jeff’s undergrad thesis: parallel backprop
0:15:54 - LLMs in 2007
0:25:09 - “Holy shit” moments
0:27:28 - AI fulfills Google’s original mission
0:32:00 - Doing Search in-context
0:36:12 - The internal coding model
0:37:29 - What will 2027 models do?
0:43:20 - A new architecture every day?
0:49:10 - Automated chips and intelligence explosion
0:53:07 - Future of inference scaling
1:02:38 - Already doing multi-datacenter runs
1:08:15 - Debugging at scale
1:12:41 - Fast takeoff and superalignment
1:20:51 - A million evil Jeff Deans
1:24:22 - Fun times at Google
1:27:51 - World compute demand in 2030
1:34:37 - Getting back to modularity
1:44:48 - Keeping a giga-MoE in-memory
1:49:35 - All of Google in one model
1:57:59 - What’s missing from distillation
2:03:10 - Open research, pros and cons
2:09:58 - Going the distance