机器翻译,已尽力保留原意与数字
内容摘要
Tesla 关于完全自动驾驶的投资者活动:定制 FSD 芯片、神经网络方法和现场自动驾驶演示,埃隆·马斯克出席。
Tesla's investor event on full self-driving: the custom FSD chip, the neural-network approach and a live autonomous demo, with Elon Musk.
中文实录Transcript
715 个段落
说话人
它。
说话人
萨。
说话人
萨姆。
说话人
它。
说话人
范围萨姆它。
说话人
萨。
皮特·班农
萨。
发言人
比赛。
发言人
萨。
发言人
萨。
发言人
萨。
发言人
萨姆。
发言人
萨。
发言人
萨姆。
分析师
萨。
发言人
它。
发言人
萨姆。
发言人
萨。
发言人
萨姆。
发言人
它。
发言人
萨。
发言人
萨姆萨。
发言人
萨姆。
发言人
萨。
发言人
萨。
发言人
萨。
发言人
萨。
发言人
它。
发言人
萨姆。
发言人
萨。
发言人
萨。
发言人
萨姆。
发言人
萨。
发言人
萨。
发言人
萨。
发言人
它。
发言人
萨姆。
分析师
萨。
发言人
萨姆。
发言人
萨。
发言人
它。
发言人
萨。
投资者关系代表
大家好。很抱歉我迟到了。欢迎参加我们首次举办的自动驾驶分析师日。我真的希望今后能更定期地举办这样的活动,让大家及时了解我们在自动驾驶方面所做的开发。大约3个月前,我们正与埃隆和其他相当多的高管一起为第四季度财报电话会议做准备。我告诉团队的一件事是,根据我经常与投资者进行的所有交流,我所看到的公司内部情况与外界认知之间最大的差距,是我们在自动驾驶方面的能力。
投资者关系代表
这也算合乎情理,因为过去几年里,我们一直都在谈论Model 3的产能爬坡。而且,你们知道,很多争论都围绕着Model 3展开,但实际上,许多事情一直在幕后发生。我们一直在开发新的完全自动驾驶芯片。我们对用于视觉识别等功能的神经网络进行了全面彻底的改造。因此,现在我们终于开始生产完全自动驾驶计算机了,我们认为揭开面纱、邀请大家进来,并谈谈过去2年里我们一直在做的一切,是个好主意。
投资者关系代表
所以大约3年前,我们想要使用,我们想要找到用于完全自动驾驶的最佳芯片。我们发现,没有任何芯片是从头开始专为神经网络设计的。因此,我们邀请了我的同事、硅工程副总裁皮特·班农,为我们设计这样的芯片。他在制造和设计芯片方面拥有大约35年的经验。其中大约12年是在一家名为PA Semi的公司度过的,这家公司后来被苹果公司收购。
投资者关系代表
因此,他参与过数十种不同架构和设计的工作,而且我想,在加入Tesla之前,他是Apple iPhone 5的首席设计师。埃隆·马斯克将与他一起登台。谢谢。
埃隆·马斯克
实际上,我本来要介绍皮特,但马丁·斯坦特。所以,他简直是我所认识的全世界最优秀的芯片和系统架构师。你和你的团队能来到Tesla是我们的荣幸,接下来交给你。就告诉他们你和你的团队完成的那些不可思议的工作吧。
皮特·班农
谢谢,埃隆。很高兴今天上午来到这里,而且能够向大家介绍过去3年里我和同事们在Tesla这里所做的全部工作,确实是一种真正的享受。
皮特·班农
我想,我们会稍微讲讲整件事是如何开始的,然后我会向大家介绍完全自动驾驶计算机,并稍微讲讲它是如何工作的。我们会深入到芯片本身,介绍其中的一些细节。我会说明我们设计的定制神经网络加速器是如何工作的,然后会向大家展示一些结果,希望到那时大家都还醒着。
皮特·班农
我是2016年2月受聘的。我问埃隆,他是否愿意说出进行完全定制系统设计所需的全部资金。他说,嗯,我们会赢吗?我说,嗯,会,当然会。所以他说,我加入。于是我们就这样开始了。我们招聘了一群人,并开始思考用于完全自动驾驶的定制芯片会是什么样子。我们花了18个月进行设计,并于2017年8月发布了设计以供制造。
皮特·班农
我们在12月拿到了它,当时封装已经通电,而且第一次尝试时它实际上就运行得非常、非常好。我们做了几处修改,并在2018年4月发布了B0修订版。2018年7月,芯片通过了认证,我们开始全面生产达到量产质量的部件。2018年12月,我们已经让自动驾驶栈在新硬件上运行,并且能够开始对员工的汽车进行改装,在现实世界中测试硬件和软件。
皮特·班农
就在3月,我们开始在Model S和Model X中交付这款新计算机。而就在4月早些时候,我们开始在Model 3中投入生产。因此,从招聘最初几名员工到在我们的全部3款汽车中全面量产,整个项目只用了略多于3年的时间,而且可能是我参与过的最快的系统开发项目。这确实充分说明了高度垂直整合所带来的优势,它让你能够开展并行工程并加快部署。
皮特·班农
在目标方面,我们完完全全只专注于Tesla的要求,这让事情容易了许多。如果你只有一个、也仅有一个客户,就不必担心其他任何事情。其中一个目标是将功耗控制在100瓦以下,以便我们能将这台新机器改装装入现有汽车。
皮特·班农
我们还希望降低部件成本,以便实现完全冗余来保障安全。当时,我们凭感觉粗略估计,驾驶汽车至少需要50万亿次运算。每秒的神经网络性能。所以我们希望至少达到这么多,而且实际上是越多越好。批量大小是指你同时处理多少个项目。例如,谷歌的TP批量大小为256,你必须一直等到有256个待处理项目之后才能开始。
皮特·班农
我们不想那么做,因此将机器设计成批量大小为1。这样一旦图像出现,我们就会立即处理。为了最大限度降低延迟,从而最大限度提高安全性,我们需要一块GPU来运行一些后处理。当时我们做了相当多这样的处理。但我们推测,随着神经网络变得越来越好,GPU上的后处理量会逐渐减少。而事实确实如此。
皮特·班农
因此,正如你们将看到的,我们冒险在设计中采用了一块性能相当适中的GPU,事实证明这个赌注押对了。安全防护极其重要。如果没有一辆具备安全防护的汽车,就不可能有一辆安全的汽车。因此,我们非常重视安全防护,当然还有芯片实际设计方面的安全性。正如埃隆之前提到的,2016年实际上还不存在从零开始设计的神经网络加速器。外面的所有人都在给自己的CPU、GPU或DSP添加指令,使其更适合推理,但没有人真正以原生方式来做这件事。
皮特·班农
所以我们着手自己来做。至于芯片上的其他组件,我们购买了用于CPU和GPU的行业标准IP。这让我们能够最大限度缩短设计时间,同时也降低项目风险。
皮特·班农
我刚来时,另一件有些出乎意料的事情,是我们能够利用Tesla现有的团队。Tesla拥有出色的电源设计团队、信号完整性分析、封装设计、系统软件、固件电路板设计,以及一个非常优秀的系统验证项目;我们能够利用这些资源来加快这个项目。这就是它的样子。
皮特·班农
右边那里,你可以看到接收车内8个摄像头传入视频的所有连接器。你可以看到电路板中央的2台自动驾驶计算机,然后左边是电源和一些控制连接。所以,当一个解决方案被精简到最基本的要素时,我真的很喜欢。你有视频、计算和电力,它直截了当、简单明了。这是计算机装入的原始 Hardware 2.5 外壳,我们过去2年一直在交付它。
皮特·班农
这是 FSD 计算机的新设计。它基本相同,而这当然是由为车辆提供改装方案的约束所决定的。我想指出,这实际上是一台相当小的计算机。它安装在手套箱后面,也就是车内手套箱与防火墙之间。它不会占掉你一半的后备箱。
皮特·班农
正如我先前所说,电路板上有2台完全独立的计算机。你可以看到它们在那里分别以蓝色和绿色突出显示。在大型 SoC 的两侧,你可以看到我们用于存储的 DRAM 芯片。然后在左下方,你可以看到构成文件系统的闪存芯片。所以这是2台独立的计算机,它们各自启动并运行自己的操作系统。
埃隆·马斯克
是的。如果我可以补充一点,这里的基本原则是,这其中任何部件都可能发生故障,而车辆仍会继续行驶。所以摄像头可能发生故障,电源电路可能发生故障,Tesla 全自动驾驶计算机芯片中的一个可能发生故障,车辆仍会继续行驶。这台计算机发生故障的概率远低于某个人失去意识的概率。那是关键指标,至少在数量级上是如此。
分析师
对。
皮特·班农
所以,为了让机器继续运行,我们额外做的事情之一,是在车内配备冗余电源。所以,一台,一台机器使用一个电源运行,另一台使用另一个电源。摄像头也是如此。所以一半摄像头由蓝色电源供电,另一半由绿色电源供电,而2颗芯片都会接收所有视频并独立处理。因此,就驾驶车辆而言,基本顺序是从周围世界收集大量信息。
皮特·班农
我们不仅有摄像头,还有雷达、GPS 地图、IMU、车辆周围的超声波传感器。我们有车轮脉冲、转向角。我们知道车辆应有怎样的加速和减速。所有这些都会整合到一起形成一个计划。一旦有了计划,2台机器就会交换各自独立生成的计划版本,以确保两者相同。假设我们达成一致,随后就采取行动并驾驶车辆。
皮特·班农
现在,一旦你通过某项新控制让车辆行驶起来,你就要对它进行验证。所以我们会验证,我们所传输的内容就是我们打算传输给车内其他执行器的内容。然后你可以使用传感器套件来确保它确实发生。所以,如果你要求车辆加速、制动、向右或向左转向,你可以查看加速度计,确保你实际上确实在这样做。因此,在这里,我们的数据采集能力和数据监控能力都具有极其大量的冗余和重叠。
皮特·班农
接下来稍微谈谈全自动驾驶芯片。它采用37.5毫米、带有1600个焊球的 BGA 封装。其中大多数用于电源和接地,但也有很多用于信号。如果取下盖子,它看起来是这样的。你可以看到封装基板,也可以看到位于中央的裸片。如果取下裸片并把它翻过来,它看起来是这样的。裸片顶部散布着13,000个 C4 凸点,而在其下方有12层金属层,它们遮住了设计的所有细节。
皮特·班农
所以,如果把它剥离掉,它看起来是这样的。
皮特·班农
这是一种14纳米 FinFET 解决方案 DMOS 工艺。它的尺寸是260毫米,属于中等尺寸的裸片。作为对比,典型的手机芯片约为100平方毫米,所以我们比它大得多。但高端 GPU 则更接近600至800平方毫米。所以我们算是处在中间。我会称其为最佳区间。这是一个适合制造的尺寸。上面有2.5亿个逻辑门,总计60亿个晶体管,即使,即使我一直在做这项工作,这对我来说也令人难以置信。
皮特·班农
这款芯片按照 AEC Q100 标准制造和测试,这是一项标准的汽车行业准则。接下来,我想沿着芯片逐一讲解它的所有不同部分。我大致会按照一个从摄像头传入的像素访问所有不同部分的顺序来讲。所以在左上方,你可以看到,看到摄像头串行接口。我们每秒可以接收25亿个像素,这足以覆盖我们所知的所有传感器,而且绰绰有余。
皮特·班农
我们有一个片上网络,用于分发来自内存系统的数据。因此,像素会通过该网络传输到芯片左右边缘的内存控制器。我们采用行业标准的 LPDDR4 内存,运行速率为每秒 4266 吉比特,这让我们的峰值带宽达到每秒 68 吉字节,这是相当充裕的带宽。但同样,这并不是什么离谱的水平。所以出于成本原因,我们算是在努力维持在一个舒适的甜点区间内。
皮特·班农
图像信号处理器拥有一条 24 位内部流水线,使我们能够充分利用车身周围配备的 HDR 传感器。它会进行高级色调映射,有助于呈现细节和阴影。然后它还有高级降噪功能,这只会改善、改善我们在神经网络中使用的图像的整体质量。神经网络加速器本身。芯片上有其中的 2 个。
皮特·班农
它们各自配有 32 兆字节的 SRAM,用于保存临时结果,并尽量减少必须在芯片内外传输的数据量,这有助于降低功耗。每个阵列都有一个 96×96 的乘加阵列,支持原位累加,使我们每个周期能够执行近 10,000 次乘加运算。它有专用的 RELU 硬件、专用的池化硬件,而其中每一个都能提供 306。抱歉,每一个都能提供每秒 36 万亿次运算,并以 2 吉赫兹运行。
皮特·班农
裸片上的这2个单元合在一起,每秒可完成72万亿次运算。所以我们大幅超越了50万亿次运算的目标。
皮特·班农
还有一个视频编码器。我们对视频进行编码,并将其用于车内的各种地方,包括倒车摄像头显示。还可以选择启用面向用户的行车记录仪功能,以及将片段记录数据上传到云端的功能,斯图尔特和安德烈稍后会进一步谈到这些。芯片上还有一个 GPU。它的性能一般。它同时支持 32 位和 16 位浮点运算。然后,我们还有 12 个 A72 64 位 C CPU,用于通用处理。
皮特·班农
它们以 2.2 吉赫兹运行。这相当于当前解决方案所能提供性能的约 2 1/2 倍。
皮特·班农
有一个安全系统,其中包含 2 个以锁步方式运行的 CPU。这个系统是判断实际驱动车内执行器是否安全的最终裁决者。所以,2 套方案会在这里汇合,我们会决定继续前进是否安全。最后还有一个安全系统。基本上,这个安全系统的职责是确保该芯片只运行经过 Tesla 加密签名的软件。
皮特·班农
如果软件没有经过 Tesla 签名,那么芯片就不会运行。
皮特·班农
现在,我已经告诉了大家许多不同的性能数据,我想或许把它们稍微放到具体背景中会有所帮助。所以在这次演讲中,我会谈到一个来自我们窄视角摄像头的神经网络。它需要 350 亿次运算,也就是 35 吉次运算;如果我们使用全部 12 个 CPU 来处理这个网络,就能达到每秒 1.5 帧,这慢得不得了,远远不足以驾驶汽车。
皮特·班农
如果我们使用每秒 600 吉次浮点运算的 GPU,运行同一个网络,就能达到每秒 17 帧,但这仍然不足以驾驶汽车。在使用 8 个摄像头的情况下,通道芯片上的神经网络加速器可以达到每秒 2100 帧。从我们一路展示的性能扩展可以看出,CPU 和 GPU 中的计算能力,与神经网络加速器所能提供的计算能力相比,基本上微不足道。
皮特·班农
确实有天壤之别。
皮特·班农
那么,接下来谈谈神经网络加速器,我们先停一下喝点水。
皮特·班农
左边有一幅神经网络的示意图,只是为了让大家了解正在发生什么。数据从顶部进入,经过每一个方框。数据沿箭头流向不同的方框。这些方框通常是带 relu 的卷积或反卷积。绿色方框是池化层。这里重要的一点是,一个方框产生的数据随后会被下一个方框使用,之后你就不再需要它了。
皮特·班农
你可以把它丢掉。所以,在数据流经网络的过程中,所有这些被创建和销毁的临时数据都不需要存储到芯片外的 dram 中。因此,我们把所有这些数据都保存在 sram 中。几分钟后我会解释为什么这一点极其重要。如果看一下这里的右侧,你会发现,在这个网络的 350 亿次运算中,几乎全部都是卷积,而卷积以点积为基础。
皮特·班农
其余的是反卷积,同样以点积为基础,然后是 relu 和池化,它们都是相对简单的运算。所以,如果你要设计某种硬件,显然会把目标对准执行以乘法、加法为基础的点积,并真正把它做到极致。但设想一下,你把它加速了 10,000 倍。那么 100% 突然就变成了 0.1%、0.01%,而 relu 和池化运算突然就会变得相当重要。
皮特·班农
所以我们的硬件并非如此。我们的硬件设计也包括用于处理 ReLU 和池化的专用资源。
皮特·班农
现在,这款芯片运行在一个受散热限制的环境中,所以我们必须非常谨慎地考虑如何消耗这些功率。我们希望最大限度地提高可完成的算术运算量。因此,我们选择了整数加法。它所需的能量仅为相应浮点加法的1/9;我们还选择了8位乘8位整数乘法,其功耗显著低于其他乘法运算,而且可能有足够的精度来获得良好结果。
皮特·班农
在内存方面,我们选择尽可能多地使用 SRAM。你们可以看到,访问芯片外的 DRAM 所需能耗大约是使用本地 SRAM 的100倍。因此显然,我们希望尽可能多地使用本地 SRAM。在控制方面,这些数据来自 Mark Horowitz 在 ISSCC 发表的一篇论文,他在其中对普通整数 CPU 执行一条指令需要消耗多少功率进行了一番评析。
皮特·班农
你们可以看到,加法运算只占总功率的0.15厘米百分比。其余所有功率都用于控制开销和记录管理。因此,在我们的设计中,我们力求尽可能彻底地摆脱这一切,因为我们真正关注的是算术运算。这就是我们最终完成的设计。你们可以看到,其中占主导的是32兆字节的 SRAM。左侧、右侧以及中间底部都有大型存储体。
皮特·班农
然后,所有计算都在中上部完成。每一个时钟周期,我们从 SRAM 阵列中读取256字节的激活数据,从 SRAM 阵列中读取128字节的权重数据,然后在一个96乘96的乘加阵列中将它们结合起来;该阵列在2吉赫兹下,每个时钟周期执行9000次乘加运算。总计为3.63 36.8万亿次运算。
皮特·班农
现在,完成点积后,我们会从引擎中卸载数据,将数据移出并通过专用的 ReLU 单元,也可以选择通过池化单元,最后进入写缓冲区,所有结果会在那里汇总。然后,我们每个周期将128字节写回 SRAM。整个过程一直持续不断地循环。因此,我们在执行点积的同时,也在卸载之前的结果、执行池化并写回内存。
皮特·班农
如果把这一切加起来,在2吉赫兹下,需要每秒1太字节的 SRAM 带宽来支持所有这些工作。而硬件提供了这样的带宽。因此,每个引擎每秒有1太字节的带宽。芯片上有2个。每秒2太字节。
皮特·班农
该加速器的指令集相对较小。我们有 DMA 读取操作,用于从内存调入数据。我们有 DMA 写入操作,用于将结果推回内存。我们有3条基于点积的指令、指令,即卷积、反卷积和内积。然后还有2种相对简单的操作:scale 是1个输入、1个输出的操作,outwise 是2个输入、1个输出的操作。当然,完成后就停止。
皮特·班农
我们必须为此开发一个神经网络编译器。因此,我们取得由视觉团队训练、原本会部署在旧款汽车上的神经网络,然后将其拿来编译,供新的加速器使用。
皮特·班农
编译器会进行层融合,使我们每次从 SRAM 读出数据再放回时,都能最大限度地提高计算量。它还会进行一些平滑处理,使内存系统承受的需求不会过于参差不齐。然后,我们还会进行通道填充,以减少存储体冲突。我们还会进行存储体感知的 SREM 分配。在这种情况下,在这种情况下,我们本可以在设计中加入更多硬件来处理存储体冲突。
皮特·班农
但通过将其推到软件中,我们以增加一些软件复杂性为代价,节省了硬件和功率。我们还会自动把 DMA 插入计算图,使数据恰好及时到达以供计算,而不必让机器停顿。然后在最后,我们生成所有代码,生成所有权重数据,对其进行压缩,并添加 CRC 校验和以确保可靠性。
皮特·班农
为了运行一个程序,所有神经网络描述、程序都会在启动时加载到 SRAM 中,然后始终待在那里,随时可以运行。因此,要运行一个网络,你必须设定输入缓冲区的地址,其中想必是一幅刚从摄像头传来的新图像。你设定输出缓冲区地址,设定指向网络权重的指针,然后设定、设定、开始,接着机器就会自行启动,并完全靠自己依次运行整个神经网络,通常会运行100万或200万个周期。
皮特·班农
然后,当它完成时,你会收到一个中断,并可以对结果进行后处理。接下来谈谈结果,我们的目标是保持在100瓦以下。这是汽车行驶过程中、运行完整 Autopilot 软件栈时测得的数据。我们的耗散功率是72瓦,比之前的设计功率稍高一些。但鉴于性能有了巨大提升,这仍然是一个相当不错的结果。在这72瓦中,大约15瓦消耗在运行神经网络上。
皮特·班农
在成本方面,这一解决方案的芯片成本约为我们之前所支付成本的80%。所以,改用这一解决方案能为我们省钱。在性能方面,我们采用了我一直在谈的窄视野摄像头神经网络,其中包含350亿次运算。我们让它在旧硬件上以尽可能快的速度循环运行,达到了每秒110帧。我们采用相同的数据、相同的网络,将它编译到新款FSD计算机的硬件上。
皮特·班农
使用全部4个加速器,我们每秒可以处理2,300帧。因此是21倍。
埃隆·马斯克
我认为这或许是最重要的一张幻灯片。简直是天壤之别。
皮特·班农
我从未参与过性能提升超过3倍的项目,所以这相当有趣。
皮特·班农
如果把它与比如英伟达的Drive Xavier解决方案相比,单颗芯片可提供21万亿次运算。我们配备2颗芯片的全停性能驾驶计算机可达到144万亿次运算。
皮特·班农
所以总结一下,我认为我们打造出了一项性能卓越的设计。神经网络处理能力达到144万亿次运算。它的功耗性能十分出色。我们设法将所有这些性能都塞进了既定的热设计功耗范围。它实现了完全冗余的计算解决方案。成本适中,而真正重要的是,这台FSD计算机将在不影响成本或续航里程的情况下,让Tesla车辆达到全新的安全和自动驾驶水平。
皮特·班农
我想这是我们所有人都期待的事情。
埃隆·马斯克
我想,不如我们在每个环节之后进行问答,这样如果大家对硬件有疑问,现在就可以提问。
埃隆·马斯克
我之所以请皮特对Tesla全自动驾驶计算机进行详细介绍,详细程度可能远超大多数人愿意了解的程度,是因为乍看之下,这似乎不太可能。此前从未设计过芯片的Tesla,怎么可能设计出全世界最好的芯片?但客观事实就是如此。不是,不是只领先一点点,而是遥遥领先。它现在就在车里。
埃隆·马斯克
目前生产的所有Tesla都配备了这台计算机。大约1个月前,我们在SNX上从英伟达解决方案切换了过来,大约10天前又在Model 3上完成了切换。所有正在生产的汽车都具备全自动驾驶所需的全部硬件,包括计算硬件和其他硬件。
埃隆·马斯克
我再说一遍。目前生产的所有Tesla汽车都具备全自动驾驶所需的一切。你需要做的只是改进软件,今天晚些时候,你们将驾驶搭载改进版软件开发版本的汽车,你们自己就会亲眼看到。
埃隆·马斯克
有问题要问皮特吗?
分析师
Y。
分析师
问题。我是Global Equities Research的Trip Chaudhary。方方面面都非常、非常令人印象深刻。我在想,就像我。我记了些笔记。你们使用的是激活函数ReLU,即修正线性单元。但如果考虑深度神经网络,它有多个层,而某些算法可能会针对不同的隐藏层使用不同的激活函数,比如softmax或tanh。你们的平台能否灵活整合LU之外的不同激活函数?
分析师
然后我还有一个后续问题。
皮特·班农
是的,可以。例如,我们实现了Tanh和Sigmoid。
分析师
太棒了。最后一个问题。比如在纳米制程方面,你提到了14nm。我在想,采用更小一点的制程不是更合理吗?也许是10nm,2年后,或者也许是7?
皮特·班农
在我们开始设计时,并非所有我们想购买的IP都有10纳米版本。所以我们最终以14纳米完成了设计。
埃隆·马斯克
也许值得指出的是,我们大概在1年半、2年前就完成了这项设计,并开始设计下一代产品。我们今天不谈下一代产品,但它已经完成大约一半了。
埃隆·马斯克
下一代芯片那些显而易见的东西,我们都会做。
说话人
对。
皮特·班农
你好。
分析师
你谈到软件现在是关键部分。你做得很出色。我被震撼到了。你说的内容我听懂了10%,但我相信这件事交到了可靠的人手中。
埃隆·马斯克
谢谢。
分析师
所以,感觉你们已经完成了硬件部分,而那确实非常难做到,现在你们必须完成软件部分。也许这超出了你的专业领域,但我们应该如何看待软件这一部分?
埃隆·马斯克
嗯,没有比这更好的方式来引出 Andre 和 Stuart 了。在演示的下一部分,也就是神经网络和软件之前,大家对芯片部分有什么问题吗。
分析师
那么,也许说到芯片方面,上一张幻灯片是每秒 144 万亿次运算,而 Nvidia 是不是 21?
皮特·班农
没错。
分析师
也许你能否为金融从业者解释一下,为什么这个差距如此显著?谢谢。
皮特·班农
嗯,我的意思是,性能差距是 7 倍。所以,这意味着你每秒可以处理 7 倍数量的帧。你可以运行规模大 7 倍、更为复杂的神经网络。因此,这是一笔非常大的现有资源,你可以把它用在许多有意思的事情上,让汽车变得更好。
埃隆·马斯克
我认为 Xavier 的功耗比我们的高。我觉得 Xavier 的功耗比我们的高。或者差不多。
皮特·班农
我不知道我是否相信这是
分析师
比如
埃隆·马斯克
据我所知,功率需求至少会以同等程度增加,即增加到7倍,成本也会增加到7倍。
埃隆·马斯克
很好。所以,是的,功率是个真正的问题,因为它也会缩短续航里程。所以功率带来的代价非常高,然后你还必须把这些功率产生的热量排出去,散热问题会变得非常严重,因为你必须把所有这些功率产生的热量都排出去。所以
投资者关系
非常感谢。我想我们有,你知道,一个
埃隆·马斯克
很多,相当多要提的问题。如果各位不介意今天的活动拖得有点久。只是我们之后会进行试驾演示。所以,如果你们有人需要先离开、早点进行试驾演示,完全可以这么做。但我确实想确保我们回答你们的问题。好的。
埃隆·马斯克
我是来自瑞银的普拉迪普·拉马尼,英特尔以及某种程度上的AMD已经开始转向基于小芯片的架构。我在这里没有看到基于小芯片的设计。展望未来,你们是否认为那
斯图尔特·鲍尔斯
会是你们从架构角度可能
埃隆·马斯克
感兴趣的东西?
皮特·班农
基于小芯片的架构?
分析师
是的。
皮特·班农
我们目前没有考虑任何类似的东西。我认为,主要是在你需要使用不同类型的技术时,它才有用。所以,如果你想在同一个硅衬底上集成硅锗或DRAM技术,那就会变得很有意思。但在裸片尺寸大得令人难以接受之前,我不会走那条路。
埃隆·马斯克
需要说明的是,这里的策略基本上始于3年,略多于3年前,就是设计并制造一台经过全面优化、以完全自动驾驶为目标的计算机。然后编写专门设计用于在那台计算机上运行的软件,并充分发挥那台计算机的性能。所以你拥有的是定制硬件,它只精通1项本领,即自动驾驶。
埃隆·马斯克
英伟达是一家很棒的公司,但他们有许多客户,因此当他们投入资源时,需要提供一种通用解决方案。
埃隆·马斯克
我们只关心1件事:自动驾驶。所以它被设计得能极其出色地完成这件事。软件也被设计得能在该硬件上极其出色地运行。我认为,这种软件和硬件的组合是无可匹敌的。
投资者关系
你好,这款芯片被设计用于处理视频输入。
安德烈·卡帕西
假如你使用,比如说激光雷达,
埃隆·马斯克
它也能够处理吗,还是说,它主要用于视频?我们今天要向你们说明的是,激光雷达是徒劳之举,任何依赖激光雷达的人都注定失败。
埃隆·马斯克
注定失败。昂贵、昂贵而又没有必要的传感器。这就像长了一大堆昂贵的阑尾。1个阑尾就够糟糕了。好吧,现在他们想装上一大堆。这太荒谬了。你们会看到的。
皮特·班农
这里前面有人。
发言者
你好。
分析师
你好。
皮特·班农
哦,那里有位先生。你好。
分析师
你好。
安德烈·卡帕西
那么,只问2个关于功耗的问题。是否可以给我们一个大致的经验法则,比如,你知道,每1瓦会使续航里程降低某个百分比或某个数值
斯图尔特·鲍尔斯
好让我们能够了解
安德烈·卡帕西
改善幅度有多大,一辆
皮特·班农
Model 3的,的目标能耗是每英里250瓦。
埃隆·马斯克
这会影响多少英里,取决于驾驶的性质。在城市里,它造成的影响会比在高速公路上好得多、大得多。所以,如果你在城市里驾驶1小时,而你采用的方案,假设其功率为1千瓦,那么Model 3会损失4英里的续航里程。所以,如果你的时速只有,比如说12英里,那么城市续航里程会受到20, 25%的影响。基本上,系统的功率对城市续航里程有巨大影响,而我们认为机器人出租车市场的大部分需求会在城市。
埃隆·马斯克
所以功率极其重要。
埃隆·马斯克
塔莎。
皮特·班农
抱歉,我没听见你说什么。
分析师
谢谢。
安德烈·卡帕西
下一代芯片的首要设计目标是什么?
埃隆·马斯克
我们不想过多谈论下一代芯片,但它是
皮特·班农
安全。
埃隆·马斯克
它至少会比当前系统好3倍,可以这么说。
皮特·班农
请讲。
埃隆·马斯克
大约还有2小时。
分析师
开发这款芯片,是这款芯片正在被——你们不生产这款芯片,而是外包生产。这能让整车成本降低多少?
皮特·班农
我提到的20%成本降低,指的是每辆车的零部件成本降低。那不是开发成本,只是实际的。
分析师
不,我是说。但是,比如说,如果我要大批量生产这些,自己做是否能省钱?
皮特·班农
是的,能省一点。
埃隆·马斯克
我的意思是,大多数芯片都是制造出来的。大多数人不会用自己的晶圆厂制造芯片。这很不寻常。
分析师
我想你们认为让芯片量产不会有任何供应问题。
皮特·班农
节省下来的成本足以支付开发费用。我的意思是,向埃隆提出的基本策略是,我们要制造这款芯片,它会降低成本。而埃隆说,乘以每年100万辆车。
分析师
成交。
埃隆·马斯克
没错,是的。
皮特·班农
抱歉。
埃隆·马斯克
如果确实有专门针对芯片的问题,我们可以回答。否则,在安德烈发言后以及斯图尔特发言后,还会有问答机会。所以还会有另外2次问答机会。现在只回答非常具体的芯片问题,
皮特·班农
另外,我整个下午都会在这里。
埃隆·马斯克
是的,正是如此。而且皮特最后也会在这里。所以请讲。
安德烈·卡帕西
哦,是的,谢谢。
斯图尔特·鲍尔斯
你刚才那张染料照片。
皮特·班农
有。
斯图尔特·鲍尔斯
神经处理器占据了染料的相当大一部分。我很好奇,那是你们自己的设计,还是其中有一些外部IP?
皮特·班农
是的,那是Tesla的定制设计。
说话人
好的。
斯图尔特·鲍尔斯
然后我想,后续问题可能是,随着你们调整设计,应该有相当大的机会缩小它的占用面积。
皮特·班农
它实际上相当密集。所以说到缩小它,我认为不行。下一代会大幅增强功能能力。
斯图尔特·鲍尔斯
好的,然后是最后一个问题。你能透露这款部件是在哪里制造的吗?
埃隆·马斯克
哪里什么?
皮特·班农
我们在哪里制造它?
埃隆·马斯克
哦,三星。
皮特·班农
三星,是的。得克萨斯州奥斯汀。
埃隆·马斯克
谢谢。
投资者关系
后面有一位。
皮特·班农
格兰特·田中,田中资本。只是好奇你们的芯片技术
分析师
和设计从、从知识产权角度看有多大的防御性,并希望
皮特·班农
你们不会把大量的
安德烈·卡帕西
知识产权免费提供给外界。
埃隆·马斯克
谢谢。
皮特·班农
我们已经为这项技术申请了大约12项专利。
皮特·班农
从根本上说,它是线性代数,我觉得这个不能申请专利。我不确定。
埃隆·马斯克
我认为,如果有人今天开始做,而且他们真的很优秀,他们也许能在3年内做出某种类似我们现在所拥有的东西,但2年后,我们会有好上3倍的东西。
分析师
谈到知识产权保护,你们拥有最好的知识产权,而有些人只是为了好玩就去窃取它。我在想,如果我们看看与 Aurora 的几次互动,业内公司认为他们窃取了你们的知识产权。我认为你们需要保护的关键要素,是与各种参数相关联的权重。你认为你们的芯片能做些什么来防止任何人吗?
分析师
也许把所有权重都加密,这样甚至连你们自己在芯片层面也不知道权重是什么,从而让你们的知识产权留在芯片内部,没有人知道,也没有人能直接窃取它。
埃隆·马斯克
本,我想见见能做到这一点的人,因为我会立刻雇用他们。是的,所以这会是个难题。
埃隆·马斯克
是的。你要不要,我是说,我们确实加密了。
埃隆·马斯克
这是一块很难破解的芯片,所以如果他们能破解它,那就非常厉害。如果他们随后能破解它,并且还、还弄明白软件、神经网络系统和其他一切。他们可以从头开始设计,就这样。
皮特·班农
我们的意图是防止人们窃取所有那些东西。而如果他们真的窃取了,我们希望这至少要花很长时间。
埃隆·马斯克
他们肯定要花很长时间。是的。我是说,我只是不认为如果那是我们的目标。我们会怎么做?那会非常困难。但我认为,对我们来说非常强大且可持续的一项优势是车队。没有人拥有这样的车队。那些权重在不断更新和改进。基于数十亿英里的行驶里程,Tesla 配备完全自动驾驶硬件的汽车数量,是其他所有公司总和的100倍。
埃隆·马斯克
你知道,我们,我们到本季度末将拥有500,000辆配备完整8摄像头配置、12个超声波传感器的汽车,其中一些在硬件上也仍然如此,但我们仍然具备收集数据的能力。然后从现在起1年后,我们将拥有超过100万辆配备完全自动驾驶计算机硬件、一切的汽车。
埃隆·马斯克
是的,所以我们有创始人。这就是巨大的数据优势。这有点类似于,你知道,Google 搜索引擎拥有巨大优势,因为人们使用它,而人们实际上是在用他们的查询和结果编程、实际上是对 Google 进行编程。
分析师
我能否就此进一步追问一下,并且请重新表述这个问题,因为我是技术外行,如果这样合适的话。但是你知道,当我们与 Waymo 或英伟达交流时,他们也以同样坚定的口吻谈论自己的领先地位,因为、因为他们具备模拟行驶里程的能力。你能谈谈拥有真实世界里程相较于模拟里程的优势吗?因为我认为他们的说法是,你知道,等你获得100万英里时,他们可以模拟10亿英里。
分析师
而且例如,没有任何一级方程式赛车手能在不使用模拟器驾驶的情况下,成功跑完真实世界的赛道。你能谈谈听起来你所认为的、从真实世界里程获取数据相较于模拟里程获取数据所带来的优势吗?
埃隆·马斯克
当然。模拟器,我们也有相当不错的模拟,但它就是无法捕捉现实世界中发生的那些奇怪事情的长尾。如果模拟完全捕捉了现实世界。嗯,我是说,我想那将证明我们生活在模拟之中。是的,它没有做到,我倒希望如此。
埃隆·马斯克
但模拟无法捕捉现实世界。现实世界真的很怪异,也很混乱。你需要让汽车上路。实际上,我们会在安德烈和斯图尔特的演示中讲到这一点。所以。好的,我们为什么不转到安德烈?
投资者关系
很好,谢谢。谢谢。
皮特·班农
谢谢大家。
投资者关系
非常感谢。
投资者关系
刚才最后一个问题实际上是一个很好的过渡,因为关于我们的 FSD 计算机,有一点需要记住,那就是它能运行复杂得多的神经网络,实现精确得多的图像识别。为了向大家介绍我们实际上如何获取这些图像数据以及如何分析它们,有请我们的人工智能高级总监安德烈·卡帕西,他会向大家解释这一切。安德烈拥有斯坦福大学博士学位,他在那里学习计算机科学,研究重点是视觉识别和深度学习。
埃隆·马斯克
安德烈,你直接讲吧?自己做介绍。斯坦福有很多博士。那不重要。是的,好的,我们不在乎。
投资者关系
请上来。
安德烈·卡帕西
谢谢。
埃隆·马斯克
安德烈在斯坦福开设了计算机视觉课程。那要重要得多。那才是关键。
安德烈·卡帕西
只是。
埃隆·马斯克
所以你能不能用一种不羞怯的方式谈谈你的背景?就告诉我你做过的事情,然后。
安德烈·卡帕西
当然。
埃隆·马斯克
对。
安德烈·卡帕西
所以,是的,我想我训练神经网络基本上已经有现在算来10年了。而这些神经网络直到大概5或6年前,实际上才真正应用于行业。所以我训练这些神经网络已经有一段时间了,其中包括在斯坦福、在、在OpenAI、在Google等机构,而且确实训练了大量神经网络,不只是用于图像,也用于自然语言,并设计将这2种模态结合起来的架构。
安德烈·卡帕西
在我攻读博士期间,所以。
埃隆·马斯克
还有计算机计算机科学课。
安德烈·卡帕西
哦,对。我其实还在斯坦福教授过卷积神经网络课程。所以我是那门课的主要授课教师。事实上,是我开设了这门课程,并设计了全部课程内容。所以一开始大约有150名学生,随后在接下来的2或3年里增加到了700名学生。所以这是一门非常受欢迎的课。它目前是斯坦福规模最大的课程之一。所以那也真的非常成功。
埃隆·马斯克
我的意思是,安德烈确实算得上是世界上最优秀的计算机视觉专家之一。可以说是最优秀的。
安德烈·卡帕西
好的,谢谢。
安德烈·卡帕西
对。所以。
安德烈·卡帕西
大家好。皮特已经向各位介绍了我们设计的、在车内运行神经网络的芯片。我的团队负责训练这些神经网络,其中包括从车队收集全部数据、训练神经网络,然后将其中一些部署到那块芯片上。
安德烈·卡帕西
那么,神经网络在车内究竟做什么?我们在这里看到的是来自车辆各处、汽车各处的视频流。这是向我们发送视频的8个摄像头。然后这些神经网络会查看这些视频,对其进行处理,并对看到的内容作出预测。因此,我们感兴趣的一些内容,以及各位在这里的可视化画面中看到的一些内容,包括车道线标记、其他物体、与这些物体的距离、我们所称的可行驶空间,以蓝色显示,也就是汽车获准行驶的地方,以及交通信号灯、交通标志等许多其他预测。
安德烈·卡帕西
现在,我的演讲大致分为3个阶段。首先,我会简短介绍神经网络,以及、以及它们如何工作和如何训练。我需要这样做,因为我需要在第2部分解释,为什么我们拥有车队是如此重大的一件事,为什么它如此重要,以及为什么它是能够真正训练这些神经网络并让它们在道路上有效工作的关键促成因素。
安德烈·卡帕西
在第3个阶段,我会谈到视觉和激光雷达,以及我们如何仅凭视觉估算深度。所以这些网络在车内解决的核心问题就是视觉识别。对你我而言,这些非常,这是一个非常简单的问题。你可以看这全部4张图像,并看出其中包含一把大提琴、一艘船、一只鬣蜥或一把剪刀。所以这对我们来说非常简单,毫不费力。
安德烈·卡帕西
计算机的情况并非如此。原因在于,对计算机来说,这些图像实际上只是一张巨大的像素网格。在每个像素上,都有该点的亮度值。因此,计算机并不是仅仅看到一幅图像,它真正得到的是网格中的 100 万个数字,这些数字告诉你所有位置上的亮度值。
埃隆·马斯克
可以说,是一个矩阵。它确实就是矩阵。
分析师
是的。
安德烈·卡帕西
因此,我们必须从那个像素和亮度值组成的网格,转化到鬣蜥之类的高层次概念。正如你可能想象的那样,这只鬣蜥具有某种亮度值模式。但鬣蜥实际上可以呈现出许多外观。所以它们可以有许多不同的外观、不同的姿势,在不同背景下处于不同的亮度条件中。你可以对那只鬣蜥进行不同的裁剪。
安德烈·卡帕西
所以我们必须在所有这些条件下都保持稳健,而且必须理解,所有那些不同的亮度模式实际上都对应着鬣蜥。现在,你我之所以非常擅长这个,是因为我们的大脑里有一个庞大的神经网络在处理那些图像。所以光线照射到视网膜上,传到你的大脑后部,也就是视觉皮层。视觉皮层由许多相互连接的神经元构成,它们对那些图像进行所有的模式识别。
分析师
图像。
安德烈·卡帕西
实际上,在过去,我会说大约 5 年里,使用计算机处理图像的最先进方法也开始使用神经网络,但在这种情况下,是人工神经网络。但这些人工神经网络——而这只是一张它的示意图——是对你视觉皮层非常粗略的数学近似。我们确实有神经元,而且它们被连接、连接在一起。
安德烈·卡帕西
而这里我只展示了 3 或 4 个神经元,在 4 层中的 3 或 4 个。但一个典型的神经网络会有数千万到数亿个神经元,每个神经元会有 1000 个连接。所以这些实际上是大块的、近乎模拟出来的组织。然后,我们可以做的是,拿来那些神经网络,向它们展示图像。所以,例如,我可以把我的鬣蜥输入这个神经网络,网络会对它看到的东西作出预测。
安德烈·卡帕西
一开始,这些神经网络是完全随机初始化的,所以所有那些不同神经元之间的连接强度都是完全随机的。因此,该网络的预测也会是完全随机的。所以它可能认为你现在看到的其实是一艘船。而它认为这实际上是一只鬣蜥的可能性非常低。在训练期间,在训练过程中,我们真正做的是,我们知道那实际上是一只鬣蜥。
安德烈·卡帕西
我们有一个标签。所以我们所做的,基本上就是说,我们希望这幅图像是鬣蜥的概率变大,而所有其他事物的概率下降。然后,有一个称为反向传播、随机梯度下降的数学过程,让我们可以通过那些连接将该信号反向传播。并更新每一个连接。
埃隆·马斯克
抱歉。
安德烈·卡帕西
并且将每一个连接都稍微更新一点。更新完成后,这幅图像是鬣蜥的概率会稍微上升。所以它可能会变成 14%,而其他事物的概率会下降。当然,我们不只是对这一幅图像这样做。实际上,我们有完整的大型标注数据集。所以我们有许多图像。通常,你可能有数百万幅图像、数千个标签之类的,而且你要一遍又一遍地进行前向、反向传播。
安德烈·卡帕西
所以,你给计算机看:这是一张图像;它有一个判断,然后你告诉它这是正确答案,它就会对自身稍作调整。你把这个过程重复数百万次,有时你也会把图像、同一张图像给计算机看数百次。所以,网络训练通常大约需要几个小时或几天,具体取决于你训练的网络有多大。
安德烈·卡帕西
这就是训练神经网络的过程。现在,神经网络的工作方式中有一点非常违反直觉,我必须真正深入讲讲,那就是它们确实需要大量这样的例子,而且确实是从零开始。它们什么都不知道。这真的很难理解。举个例子,这里有一只可爱的狗。你可能不知道这只狗的品种,但正确答案是,这是一只日本狆。
安德烈·卡帕西
现在,我们所有人都看着这个,看到日本狆,然后会想,好吧,我明白了。我大概知道这种日本狆长什么样了。如果我再给你看几张其他狗的图像,你大概就能从中挑出其他日本狆。所以具体来说,那3只看起来像日本狆,其他的则不像。所以你可以很快做到这一点。而且你只需要一个例子,但计算机不是这样工作的。
安德烈·卡帕西
它们实际上需要大量日本狆的数据。所以这是一个展示日本狆的图像网格,你需要数千个例子,展示它们不同的姿势、不同的亮度条件、不同的背景、不同的裁剪方式。你确实需要从所有不同角度教会计算机这种日本狆长什么样。而要让它奏效,确实需要所有这些数据。否则,计算机无法自动捕捉到这种模式。
安德烈·卡帕西
那么,这一切对于自动驾驶的场景意味着什么?当然,我们不太关心狗的品种。也许将来某个时候会关心,但现在我们真正关心的是车道线标记、物体在哪里、哪里可以行驶,等等。所以,我们的做法是,我们没有像“鬣蜥”这样的图像标签,但我们确实有来自车队的这种图像。而我们感兴趣的,比如说,是车道线标记。
安德烈·卡帕西
所以我们,通常由一个人打开一张图像,用鼠标标注车道线标记。这里是一个标注示例,人可以为这张图像创建一个标签。它表示这就是你应该在这张图像中看到的内容。这些是车道线标记。然后我们可以去车队那里,向车队索取更多图像。如果你向车队索取,如果你只是用一种简单粗糙的方式来做,只是随机索取图像,车队可能会返回这样的图像。
安德烈·卡帕西
通常是在某条高速公路上向前行驶,你可能就会得到这样一组随机集合。我们会标注所有这些数据。现在,如果你不够谨慎,只标注了这些数据的随机分布,你的网络就会在某种程度上捕捉到这种数据的随机分布,并且只在那种情形下工作。所以,如果你给它看一个稍有不同的例子,比如这里有一张图像,实际上道路正在弯曲,而且这里更像是住宅区。
安德烈·卡帕西
那么,如果你把这张图像展示给神经网络,那个网络可能会做出错误的预测。它可能会说,好吧,我在高速公路上见过很多次,车道就是向前延伸的。所以这里有一种可能的预测。当然,这非常不正确,但其实不能责怪神经网络。它不知道左边的火车,那棵树,到底重要不重要。它不知道右边的汽车对车道线而言是否重要。
安德烈·卡帕西
它不知道背景中的建筑物是否重要。它确实完全从零开始。而你我都知道,事实是那些东西都不重要。真正重要的是,在远处一个消失点附近有几条白色车道线标记。而它们略微弯曲这一事实应该把预测拉过去。只不过,我们没有任何机制可以直接告诉神经网络,嘿,那些车道线标记实际上很重要。
安德烈·卡帕西
我们工具箱里唯一的工具就是已标注数据。所以,我们需要做的是,当网络在这种图像上失败时,把它们拿出来并正确标注。因此在这个例子中,我们会让车道向右转,然后需要把大量这样的图像输入神经网络。随着时间推移,神经网络会积累,会基本捕捉到这种模式:那边那些东西并不重要,但那些车道线标记很重要,于是我们学会预测正确的车道。
安德烈·卡帕西
所以,真正关键的不只是数据集的规模。我们不只是想要数百万张图像。实际上,我们需要很好地覆盖汽车在道路上可能遇到的各种情况所构成的空间。所以,我们需要教计算机如何处理夜间和潮湿的场景,那时会有各种不同的镜面反射,而正如你可能想象的,这些图像中的亮度模式会显得非常不同。
安德烈·卡帕西
我们必须教计算机如何处理阴影,如何处理道路分叉,如何处理可能占据图像大部分区域的大型物体,如何处理隧道,或如何处理施工现场。而在所有这些情况下,同样,没有明确的机制可以告诉网络该怎么做。我们只有海量数据。我们希望收集所有这些图像,并标注正确的线条,而网络会从这些如今规模庞大且多样的数据集中捕捉模式,基本上让这些网络运行得非常好。
安德烈·卡帕西
这并不只是我们在 Tesla 这里得到的发现。这是整个行业中一个普遍存在的发现。因此,来自 Google、Facebook、百度、Alphabet 旗下 DeepMind 的实验和研究都展示了相似的图表,其中神经网络确实热爱数据,也热爱规模和多样性。随着你添加更多数据,这些神经网络开始表现得更好,并且不费额外代价就能获得更高的准确率。所以,更多数据就是会让它们运行得更好。
安德烈·卡帕西
现在有许多公司,有许多人曾经指出,我们或许可以利用模拟来真正达到数据集所需的规模。而且这里有很多条件由我们掌控。也许现在我们能在 Tesla 的模拟器中实现一定的多样性。刚才的问题中也有人提到了这一点。现在在 Tesla,这实际上是我们自己的模拟器的一张截图。
安德烈·卡帕西
我们广泛使用模拟。我们用它来开发和评估软件。我们甚至也相当成功地把它用于训练。所以,但是当真正涉及神经网络的训练数据时,确实没有任何东西能替代真实数据。模拟在对外观、物理现象以及你周围所有参与者的行为进行建模时面临很多困难。这里有一些例子,真正说明现实世界确实会向你抛来许多疯狂的东西。
安德烈·卡帕西
所以在这个例子中,比如说,我们面对的是非常复杂的环境,有雪、有树、有风。我们会遇到各种可能很难模拟的视觉伪影。我们会遇到复杂的施工现场,还有可能被吹入其中、可能会随风四处飘动的灌木和塑料袋。复杂的施工现场可能会有许多人、孩子、动物,全都混杂在一起。而模拟这些事物如何相互作用,以及如何穿行于这个施工区域,实际上可能完全、完全不可行。
安德烈·卡帕西
问题不在于其中任何一个行人的移动。问题在于他们如何相互反应,那些车辆如何相互反应,以及当你在那种环境中驾驶时,他们如何对你作出反应。而所有这些实际上都非常难以模拟。就好像仅仅为了在模拟中模拟其他车辆,你就必须先解决自动驾驶问题。所以这真的很复杂。所以我们会遇到狗、珍奇动物,而在某些情况下,甚至不是你无法模拟,而是你根本连想都想不到。
安德烈·卡帕西
所以,比如说,我之前不知道卡车可以像那样,卡车叠在卡车上。但在现实世界中,你会发现这种情况,也会发现许多其他甚至真的很难想象出来的东西。所以,我从车队传回的数据中看到的多样性简直疯狂。就我们的模拟器而言,我们有一个非常好的模拟器。
埃隆·马斯克
我的意思是,我认为模拟从根本上说就是你在给自己的作业打分。所以,如果你知道自己要模拟它,好吧,你当然能针对它找到解决方案。但正如安德烈所说,你不知道自己不知道什么。这个世界非常奇怪,有数百万种边缘情况。
埃隆·马斯克
而且,如果有人能制作出一种与现实准确吻合的自动驾驶模拟,那本身就会是人类能力的一项不朽成就。他们做不到。没有办法。
安德烈·卡帕西
对。所以我认为,到目前为止我真正试图强调的3点是,要让神经网络良好运作,你需要这3个基本要素。你需要一个大型数据集、一个非常数据集,以及一个真实数据集。而如果你具备这些能力,实际上就可以训练神经网络,并让它们非常良好地运作。那么,为什么 Tesla 处在如此独特而有趣的位置,能够真正把这3个基本要素全都做好呢?
安德烈·卡帕西
这个问题的答案,当然、当然是车队。我们确实可以从车队获取数据,让我们的神经网络系统极其良好地运作。那么,让我通过一个具体例子带大家看看,比如说,如何让物体检测器更好地运作,让大家了解我们如何开发这些神经网络、如何对它们进行迭代,以及实际上如何让它们随着时间的推移正常运作。物体检测是我们非常重视的事情。
安德烈·卡帕西
我们希望在比如这里的车辆和物体周围画上边界框,因为我们需要追踪它们,也需要了解它们可能会如何移动。所以,我们还是可能会请人工标注员为这些内容提供一些标注。人工标注员可能会介入,告诉你,好吧,那边的那些图案是汽车、自行车,等等。然后你可以用这些数据训练神经网络。但如果你不小心,神经网络在某些情况下就会作出错误预测。
安德烈·卡帕西
举个例子,如果我们偶然遇到像这样一辆后面载着自行车的汽车,那么在我加入时,神经网络实际上会创建两个检测结果。它会创建一个汽车检测结果和一个自行车检测结果。而这其实也算正确,因为我想这两个物体确实都存在。但对于下游的控制器和规划器而言,你真的不希望去处理这辆自行车可以和汽车一起移动这个事实。
安德烈·卡帕西
事实是,那辆自行车是固定在那辆汽车上的。所以,就道路上的物体而言。这里只有一个物体,一辆汽车。因此,你现在想做的是,可能把大量这类图像都标注成这只是一辆汽车。所以,我们团队内部采用的流程是,拿这张图像或几张呈现这种模式的图像,然后通过一种机制、一种机器学习机制,让车队为我们提供看起来像这样的样本。
安德烈·卡帕西
车队可能会返回包含这些模式的图像。举个例子,这6张图像可能来自车队。它们都包含汽车后面载着自行车的情况。我们会进去,把这些全都标注成只有一辆汽车。然后,那个检测器的性能实际上就会提高。网络会在内部理解,嘿,当自行车只是固定在汽车上时,那实际上就只是一辆汽车。
安德烈·卡帕西
只要有足够多的样本,它就能学会这一点,而这就是我们大致解决那个问题的方式。我想提一下,我谈了很多从车队获取数据的内容。我只想快速说明一点:我们从一开始就在设计中考虑了隐私问题,我们用于训练的所有数据都经过了匿名化处理。现在,车队返回的不只是汽车后面载着自行车的情况。我们会寻找所有东西。我们一直在寻找很多东西。
安德烈·卡帕西
比如,我们会寻找船只,而车队可以返回船只的图像。我们会寻找施工现场,而车队可以从世界各地向我们发送大量施工现场的图像。我们甚至会寻找稍微更罕见的情况。比如,发现道路上的碎片对我们相当重要。这些就是从车队传输给我们的图像示例,其中显示了轮胎、锥桶、塑料袋之类的东西。
安德烈·卡帕西
如果我们能大规模获取这些图像,就可以正确标注它们,而神经网络可以学习如何在现实世界中处理它们。这里还有一个例子。动物当然也是非常罕见的情况和事件。但我们希望神经网络真正理解这里发生了什么,理解这些是动物,并且我们希望正确处理这种情况。总结一下,我们迭代改进神经网络预测的流程大致是这样的。
安德烈·卡帕西
我们从一个可能随机获取的种子数据集开始。我们标注这个数据集,然后用这个数据集训练神经网络,并把它部署到车里。接下来,我们有一些机制,可以在这个检测器可能表现异常时发现车内系统的不准确之处。比如,如果我们检测到神经网络可能不确定,或者如果我们检测到,或者驾驶员在其中任何一种情况下进行了干预,我们就可以建立这种触发基础设施,把这些不准确情况的数据发送给我们。
安德烈·卡帕西
比如,如果我们的车道线检测在隧道中的表现不是很好,那么我们可以发现隧道里存在问题。那张图像会进入我们的单元测试。这样我们就可以验证我们确实在随着时间推移修复这个问题。但现在,要修复这种不准确情况,你需要获取更多看起来像那样的样本。所以我们请求车队向我们发送更多隧道图像。然后我们正确标注所有这些隧道,把它们纳入训练集,再重新训练网络、重新部署,并一遍又一遍地迭代这个循环。
安德烈·卡帕西
因此,我们把这个用来改进这些性能预测的迭代流程称为数据引擎。也就是反复部署某个可能处于影子模式的系统,获取不准确情况,并一遍又一遍地纳入训练集。我们基本上会对这些神经网络的所有预测都这样做。到目前为止,我谈了很多显式标注。就像我提到的,我们会让人们标注数据。
安德烈·卡帕西
这个流程在时间方面成本很高。而且在……方面也是。是的,它就是一个成本很高的流程。因此,要完成这些标注当然可能非常昂贵。所以,我还想谈谈如何真正利用车队的力量。你不会想受制于这种人工标注瓶颈。你会希望直接把数据传输进来,并自动地将其自动化。我们有多种机制可以做到这一点。
安德烈·卡帕西
我们最近开展的一个项目示例是检测车辆加塞。也就是你正在高速公路上行驶,有人在你的左侧或右侧,然后他们加塞到你前方,进入你的车道。这里有一段视频,展示自动辅助驾驶检测到这辆车正在侵入我们的车道。当然,我们希望尽快检测到加塞。因此,我们处理这个问题的方式并不是编写显式代码来判断:左转向灯亮了吗?
安德烈·卡帕西
右转向灯亮了吗?持续追踪键盘,看看它是否在水平移动。实际上,我们采用的是车队学习方法。其工作方式是,每当车队看到一辆汽车从右侧车道转入中间车道,或者从左侧转入中间车道时,就请求车队向我们发送数据。然后我们会把时间倒回去,并且可以自动标注:嘿,那辆汽车将会转向,将在1.3秒后加塞到你前方。
安德烈·卡帕西
然后我们可以用它来训练神经网络。因此,神经网络会自动识别其中的许多模式。例如,车辆通常存在偏航,它们正朝这个方向移动,也许转向灯亮着。所有这些都仅仅通过这些示例在神经网络内部发生。因此,我们要求车队自动把所有这些数据发送给我们。我们可以获得大约50万张图像,而且所有这些图像都会标注车辆切入。
安德烈·卡帕西
然后我们训练网络。接着,我们把这个切入网络部署到了车队中。但我们还不启用它。我们让它以影子模式运行。在影子模式下,网络始终在进行预测。嘿,我认为这辆车要切进来了。从它的样子来看,这辆车要切进来了。然后我们寻找错误预测。举个例子,这是我们从切入网络的影子模式中获得的一段视频。
安德烈·卡帕西
这有点难看清,但网络认为我们正前方右侧的车辆将要切进来。你大概能看到它在稍微试探车道线。它正试图,它有点稍微侵入了。网络兴奋起来,觉得那将会是一次切入。那辆车实际上最终会进入我们的中间车道。这后来被证明是不正确的,因为。而且那辆车实际上并没有那么做。
安德烈·卡帕西
所以我们现在所做的,就是不断运转数据引擎。我们获取了那个在影子模式下运行的。它会进行预测,会产生一些假阳性,也有一些假阴性检测。所以我们有时过于兴奋,有时又会漏掉实际发生的切入。所有这些都会生成一个传给我们的触发信号,而这些现在会免费纳入。标注这些数据的过程中没有人类受到伤害,它们被免费纳入我们的训练集。
安德烈·卡帕西
我们重新训练网络,并重新部署影子模式。因此,我们可以把这个过程重复几次。我们始终关注来自车队的假阳性和假阴性。一旦我们对假阳性、假阴性的比例感到满意,就会真正切换那个比特位,并真正让车辆交由那个网络控制。所以你们可能已经注意到,我们实际上大约在,我想是3个月前,发布了切入检测器的首批版本之一。
安德烈·卡帕西
所以,如果你已经注意到车辆在检测切入方面好得多了。那就是大规模运行的车队学习。是的,它实际上运行得相当好。这就是车队学习。过程中没有人类受到伤害。它只是在基于数据进行大量神经网络训练,并进行大量影子模式运行。观察这些结果,另一个本质上,
埃隆·马斯克
归根结底,就像每个人始终都在训练网络。无论,无论Autopilot是开启还是关闭,网络都在接受训练。对于搭载Hardware 2或更高版本的车辆,车辆行驶的每一英里都在训练网络。
分析师
是的。
安德烈·卡帕西
在车队学习体系中,我们使用它的另一个有趣方式,以及我要谈到的另一个项目,是路径预测。所以当你驾驶车辆时,你实际是在标注数据,因为你在转动方向盘,你在告诉我们如何穿越不同的环境。我们现在看到的是车队中的某个人左转通过一个十字路口。
安德烈·卡帕西
而我们在这里所做的是,我们,我们拥有所有摄像头的完整视频,而且我们知道这个,这个人所走的路径,因为有GPS、初始测量单元、方向盘角度、车轮脉冲。所以我们把所有这些整合起来,并理解这个人在这个环境中所走的路径。当然,这个,这个,我们可以用它来监督网络。所以我们只需从车队中获取大量此类数据。
安德烈·卡帕西
我们用这些,用那些轨迹训练一个神经网络,然后神经网络保护仅根据这些数据预测路径。所以实际上,这通常被称为模仿学习。我们从现实世界中获取人类的轨迹,然后只是尝试模仿人们在现实世界中如何驾驶。我们也可以把同样的数据引擎运转过程应用于这一切,并让它随着时间推移发挥作用。
安德烈·卡帕西
这里是路径预测穿越一种复杂环境的例子。你们现在看到的是一段视频,而我们正在叠加网络的预测,预测结果。所以绿色的是网络会遵循的一条路径,以及一些。
埃隆·马斯克
是的,我是说,疯狂的是,这个网络正以极高的准确率预测它甚至看不到的路径。它看不到拐角后面。但、但它说那条曲线的概率极高。所以那就是路径,而且它预测准了。你们今天会在车里看到这一点。但我们会开启增强视觉,这样你们就能看到叠加在视频上的车道线和车辆的路径预测。
安德烈·卡帕西
是的,底层实际发生的事情甚至比你能看出来的更多。
埃隆·马斯克
我是说,说实话,这有点吓人。
安德烈·卡帕西
当然,还有很多细节我略过了。你可能不想标注所有驾驶员,可能只想模仿更优秀的驾驶员。而我们实际上有许多技术方法来切分和处理这些数据。但这里有意思的是,这个预测实际上是一个 3D 预测,我们把它反向投影到这里的图像上。所以这里向前的路径是一个三维的东西,我们只是把它渲染成 2D。
安德烈·卡帕西
但我们能从这一切中了解地面的坡度,而这实际上对驾驶极其有价值。顺便说一下,路径预测如今实际上已经在车队中实时运行。所以,如果你正行驶在苜蓿叶式立交桥上,如果你在高速公路的苜蓿叶式立交桥上,直到大约 5 个月前,你的车还无法驶过苜蓿叶式立交桥。现在可以了。那就是正在你们的车上实时运行的路径预测。我们前一阵子已经发布了它,而今天你们将能体验到这一点。
安德烈·卡帕西
对于穿过交叉路口,你们今天驾车时我们如何通过交叉路口,很大一部分都来源于基于自动标签的路径预测。
安德烈·卡帕西
所以,我到目前为止讲的其实是我们如何迭代网络预测、以及如何使其随着时间推移发挥作用的 3 个关键组成部分。你需要庞大、多样且真实的数据集。我们在 Tesla 确实能够做到这一点,而我们是通过车队的规模、数据引擎、以影子模式发布功能、迭代这个循环,以及可能甚至使用不会让任何人工标注员在此过程中受到伤害的车队学习,仅仅自动使用数据来做到这一点的。
安德烈·卡帕西
而且我们确实可以大规模做到这一点。
安德烈·卡帕西
那么在我演讲的下一部分,我将特别谈谈仅使用视觉进行深度感知。你们可能知道,车里至少有 2 种传感器。一种是只获取像素的视觉摄像头,另一种是许多公司也使用的激光雷达。激光雷达会为你提供周围距离的这些点测量值。现在,我首先想指出的一点是,你们所有人都来到了这里,你们中的许多人是开车来的,而你们使用的是自己的神经网络和视觉。
安德烈·卡帕西
你们并没有从眼睛里发射激光,但还是来到了这里。
埃隆·马斯克
我们也许有。
安德烈·卡帕西
事情进行得很顺利。
安德烈·卡帕西
所以很明显,人类神经网络仅通过视觉就能推导出距离、所有测量结果以及对世界的 3D 理解。它实际上会使用多种线索来做到这一点。我会简要介绍其中一些,只是为了让你们大致了解内部发生了什么。举个例子,我们有 2 只朝前的眼睛,所以在每一个时间步,你都能对前方世界获得 2 个独立的测量结果。你的大脑会把这些信息拼接起来,得出某种深度估计,因为你可以跨这 2 个视点对任何点进行三角测量。
安德烈·卡帕西
相反,许多动物的眼睛长在两侧,所以它们的视野重叠非常少。因此,它们通常会将结构用于运动。其思路是,它们会上下摆动头部,而由于这种移动,它们实际上能对世界获得多次观察。然后你又可以通过三角测量得到深度。即便闭上一只眼睛并且完全不动,你仍然可以拥有一定的深度感知。
安德烈·卡帕西
如果你这样做,我认为你不会察觉到我向你靠近 2 米或向后退 100 米。那是因为还有许多非常强的单眼线索,你的大脑也会将它们考虑在内。这是一个相当常见的视觉错觉示例,其中,你知道,这 2 根蓝色条是完全相同的,但你的大脑在拼接这个场景时,只是因为这幅图像中的消失线而认为其中一根应该比另一根更大。
安德烈·卡帕西
所以你的大脑会自动完成很多这样的事情。而神经网络,也就是人工神经网络,也可以做到。那么让我给你们举 3 个例子,说明如何仅凭视觉获得深度感知。一个经典方法,以及 2 个依赖神经网络的方法。这是一段向下行驶的视频,我想这是在旧金山,是一辆 Tesla。所以这些是我们的摄像头、我们的感知,而我们正在查看全部。我只展示了主摄像头,但所有摄像头都已开启,也就是 Autopilot 的 8 个摄像头。
安德烈·卡帕西
如果你只有这段 6 秒的视频片段,可以做的是使用多视图立体技术把这个环境拼接成 3D。所以这个。
埃隆·马斯克
哎呀。
安德烈·卡帕西
这应该是一段视频,不是吗?一段视频?哦,我知道它是。好了。所以,这是那辆车沿那条路径行驶的那6秒的3D重建。你可以看到,这些信息纯粹,它非常容易从、从仅仅视频中恢复出来。而这大致是通过三角测量过程,以及我提到的多视图立体视觉。我们也在车上应用了类似的技术,稍微更稀疏、更近似一些。
安德烈·卡帕西
所以,值得注意的是,所有这些信息其实都存在于传感器中,只需把它提取出来。我想简要谈谈的另一个项目是,正如我提到的,这与神经网络无关。神经网络是非常强大的视觉识别引擎。而如果你想让它们预测深度,那么你就需要寻找例如深度标签。然后它们确实可以把这件事做得非常好。
安德烈·卡帕西
所以,除了标注数据之外,没有任何因素会限制网络预测这种单目深度。我们实际上在内部研究过的一个示例项目是,我们使用前向雷达,也就是这里以蓝色显示的部分,该雷达会向外探测并测量物体的深度。我们用该雷达来标注视觉所看到的内容,也就是神经网络输出的边界框。所以,与其让人工标注员告诉你,好,这辆车以及这个边界框大约在25米外,不如使用传感器更好地标注这些数据。
安德烈·卡帕西
所以你使用传感器标注。举例来说,雷达非常擅长测量那个距离。你可以对其进行标注,然后用它来训练神经网络。如果你拥有足够多的此类数据,这个神经网络就会非常擅长预测这些模式。这里是这类预测的一个例子。圆圈中,我展示的是雷达物体,而在。而这里输出的长方体完全来自视觉。所以这里的长方体仅仅来自视觉。
安德烈·卡帕西
而这些长方体的深度,是通过来自雷达的传感器标注学习到的。所以,如果它运行得非常好,那么你会看到俯视图中的圆圈与长方体吻合。它们确实吻合。这是因为神经网络非常擅长预测深度。它们可以在内部学习不同车辆的尺寸,并且知道这些车辆有多大。而你实际上可以据此相当准确地推导出深度。
安德烈·卡帕西
我要非常简短地谈论的最后一种机制略微更复杂,也更技术化一些。不过,这是一种最近出现的机制,过去1年或2年里基本上已有几篇关于这种方法的论文。它叫作自监督。所以,在许多此类论文中,你所做的只是把完全没有任何标签的原始视频输入神经网络。而你仍然可以学习,仍然可以让神经网络学习深度。
安德烈·卡帕西
这有点技术性,所以我无法深入讲解全部细节,但其理念是,神经网络会预测该视频每一帧中的深度。然后,神经网络没有要通过标签回归拟合的明确目标。相反,网络的目标是随时间保持一致。所以,无论你预测出什么深度,都应该在该视频的持续时间内保持一致。
安德烈·卡帕西
而保持一致的唯一方式就是正确。因此,神经网络会自动预测所有像素的正确深度。我们已经在内部复现了其中一些结果。所以这种方法也相当有效。
安德烈·卡帕西
总而言之,人们仅凭视觉驾驶,不涉及激光。这似乎运行得相当好。我想强调的一点是,视觉识别,而且是非常强大的视觉识别,对于自动驾驶绝对必不可少。它不是可有可无的东西,我们必须拥有真正能够切实理解你周围环境的神经网络。而激光雷达点所包含的环境信息要少得多。所以,视觉才能真正理解全部细节。
安德烈·卡帕西
仅仅周围的几个点要少得多。其中包含的信息少得多。举个例子,这里左侧的是塑料袋还是轮胎?激光雷达可能只会给你提供上面的几个点,但视觉可以告诉你这两者中哪一个才是真的,而这会影响你的控制。那个略微向后看的人,是正骑着自行车试图并入你的车道,还是只是在施工现场继续向前骑行?
安德烈·卡帕西
那些标志上写着什么?我在这个世界中应该如何行动?我们为道路建设的全部基础设施,都是为供人类视觉获取信息而设计的。所以,所有标志、所有交通信号灯,一切都是为视觉设计的。因此,所有这些信息都在那里。所以你需要这种能力。那个人是否分心并在看手机,他们会不会走进你的车道?所有这些问题的答案只能通过视觉找到,而且对于第四级、第五级自动驾驶而言是必需的。
安德烈·卡帕西
而这正是我们在Tesla开发的能力,这是通过大规模神经网络训练、数据引擎、让它随时间推移发挥作用,以及利用车队的力量共同实现的。所以从这个意义上说,激光雷达实际上是一条捷径。它绕开了根本问题,也就是自动驾驶所必需的视觉识别这一重要问题。因此,它会造成一种虚假的进展感,最终只是一个拐杖。
安德烈·卡帕西
它确实能让你,比如说,很快做出演示。
安德烈·卡帕西
所以,如果要我总结这场
分析师
完整的
安德烈·卡帕西
我的整场演讲,用一张幻灯片来总结,就是这个,全部的自动驾驶。因为你想要的是能够应对所有可能情况的第四级、第五级系统,在99.99%的情形中,而追逐最后几个夜晚中的一些情况将会非常棘手、非常困难,并且需要一个非常强大的视觉系统。所以,我正在向你展示一些图片,说明你可能在那个9的任意一个切片中遇到什么。最开始,你只有非常简单的汽车。
安德烈·卡帕西
继续往后,那些汽车开始看起来有点奇怪。然后汽车上也许有自行车,再然后汽车上也许有汽车。然后你也许开始遇到真正罕见的事件,比如汽车翻倒,甚至汽车腾空。我们会看到大量来自车队的事物,而且我们会以某种频率看到它们,与我们所有竞争对手相比,是以一种非常高的频率。因此,你真正解决这些问题、迭代软件并真正向神经网络提供正确数据的进展速度。
安德烈·卡帕西
这种进展速度实际上只与你在真实环境中遇到这些情况的频率成正比。而我们遇到它们的频率显著高于其他任何人,这就是我们将会表现得极其出色的原因。谢谢。
分析师
请讲。
安德烈·卡帕西
这一切都让人印象极其深刻。非常感谢。你们平均从每辆车收集多少数据、多少张图片
皮特·班农
在每段时间内?
安德烈·卡帕西
然后,听起来配备双、双主动。主动计算机的新硬件,为你们提供了一些非常有趣的机会,可以在完整模拟中运行神经网络的一个
皮特·班农
副本,同时你在
安德烈·卡帕西
运行另一个。运行另一个。
皮特·班农
驾驶汽车并比较结果
安德烈·卡帕西
来做质量保证。然后我还想知道,是否还有其他机会,可以在汽车停在车库里的时候,利用这些计算机进行训练,因为有90%的时间我没有在驾驶我的时间。
投资者关系
开着Tesla到处走。
安德烈·卡帕西
非常感谢。是的。那么第一个问题,我们能从车队获得多少数据?这里非常重要的一点是,关键不仅仅在于数据集的规模,真正重要的是该数据集的多样性。如果你只有大量某个东西在高速公路上向前行驶的图像,那么到了某个时候,神经网络就掌握它了。你不需要那些数据。所以,我们在如何挑选和取舍方面非常有策略。
安德烈·卡帕西
而且,我们建立的触发基础设施相当复杂,使我们能够只获取眼下所需的数据。因此,数据量并不是非常庞大,只是数据挑选得非常好。关于第二个冗余问题,当然可以。你基本上可以在两边都运行一份网络副本。而它实际上正是以这种方式设计的,以实现具有冗余能力的第四级、第五级系统。
安德烈·卡帕西
所以,情况确实如此。还有你的最后一个问题。抱歉,我没有。
埃隆·马斯克
训练,这辆车是一台针对推理优化的计算机。我们在 Tesla 确实有一个重要项目,今天没有足够时间讨论,叫作 Dojo。那是一台超级强大的训练计算机。Dojo 的目标将是能够接收海量数据,在视频层面进行训练,并通过 Dojo 项目,对海量视频开展无监督的大规模训练。Dojo 计算机。但那是以后再谈的事。
分析师
从某种意义上说,我就像一名试飞员,因为我驾驶 40510,而所有这些非常棘手、真正属于长尾的问题每天都会发生。但有一个挑战,我很好奇你们打算如何解决,那就是变道。因为每当我试图驶入有车流的车道时,所有人都会抢到你前面。所以人的行为非常不理性。当你在洛杉矶开车时,汽车只想安全地完成操作,而你几乎必须以不安全的方式去做。
分析师
所以我想知道你们准备如何解决那个问题。好的。
安德烈·卡帕西
所以,我要指出的一点是,我谈到过将数据引擎用于神经网络的迭代,但我们也在软件层面,以及实际变道时机、激进程度等选择所涉及的所有超参数上做完全相同的事情;我们一直在改变这些参数,可能会以影子模式运行它们,并观察其运行效果。因此,为了调校关于何时可以变道的启发式规则,我们也可能会利用数据引擎和影子模式等等。
安德烈·卡帕西
归根结底,我认为在一般情况下,实际设计所有关于何时可以变道的不同启发式规则,其实有点难以处理。因此,理想情况下,你实际上会希望利用车队学习来指导这些决策。那么,人类会在什么时候变道,在什么场景下变道,又会在什么时候觉得变道不安全?我们就查看大量数据,并训练机器学习分类器,以区分什么时候这样做过于安全。
安德烈·卡帕西
而那些机器学习分类器能够写出比人类好得多的代码,因为背后有海量数据作为支撑。因此,它们确实能够调校所有正确的阈值,与人类保持一致,并采取安全的做法。
埃隆·马斯克
我想,我们可能会推出一种比“疯狂麦克斯”模式更进一步的模式,也就是洛杉矶交通模式。是的,嗯,你知道,我觉得“疯狂麦克斯”在洛杉矶的交通中都有点难办。
安德烈·卡帕西
是的。所以这其实是一种权衡。你不希望制造不安全的情况,但你又希望表现得果断。不过,作为人类,你如何通过那种微妙的配合来做到这一点,其实非常复杂,而且很难写成代码。但我认为,我们确实是。机器学习方法确实似乎有点像是正确的
安德烈·卡帕西
处理方式。
安德烈·卡帕西
我们只需观察人们采取这种做法的大量方式,并尝试模仿。
埃隆·马斯克
我们现在只是更加保守。然后,随着我们的信心越来越高,我们会允许用户选择更激进的模式。
埃隆·马斯克
那将由用户决定。
埃隆·马斯克
但在更激进的模式下,以及试图汇入车流时,确实存在轻微的。无论多少新的,还是有轻微的可能性会发生像小剐蹭这样的事,不是严重事故,但基本上你将面临一个选择:你是否愿意接受在高速公路车流中发生小剐蹭的非零概率,不幸的是,这是穿行于洛杉矶交通中的唯一方式。
分析师
是的。
埃隆·马斯克
对。
分析师
是的。
埃隆·马斯克
我是说,是的。是的。这总让我想起《洛城故事》之类的。这部电影是一部很棒的电影。
安德烈·卡帕西
对,这非常微妙,因为这里正在进行一场胆小鬼博弈。
埃隆·马斯克
对。
埃隆·马斯克
随着时间推移,我们会提供更激进的选项,由用户指定。是的。疯狂麦克斯加强版。没错。
安德烈·卡帕西
哦,对。
分析师
你好。嗨。我是来自 Canaccord Genuity 的杰德·多尔斯海默。谢谢你们,也祝贺你们所开发的一切。当我们审视 AlphaZero 项目时,它。就其参数而言,那是一个定义非常明确且变量有限的项目,这使得学习曲线能够如此之快。
分析师
风险,或者说你们在这里试图做的事情,几乎是通过神经网络在汽车中发展出意识。所以我想,挑战在于,从车队的集中式模型提取信息,到把控制权交给已经拥有足够信息的汽车,在这个过程中,你们如何避免形成循环引用。
分析师
那条界线在哪里?我想,就学习过程进行到什么节点而言,才会把它交给车内已有足够信息、无需再从车队提取信息的汽车。
埃隆·马斯克
嗯,即使汽车与车队完全断开连接,它也可以运行。
埃隆·马斯克
它只是,它会上传训练成果,你知道,随着车队变得越来越好,这些训练成果也会越来越好。所以简单来说,如果从那一刻起让它与车队断开连接,它就不会再继续改进,但仍然可以正常运行。
分析师
在你们分享内容的硬件部分。在前一个版本中,谈到了不存储大量图像所带来的许多功耗优势。而在这一部分,你们谈到通过从车队提取信息来进行学习。我想我很难调和这一点:如果出现这样一种情况,我像你们展示的那样驾车上坡,并预测道路将延伸到哪里,那是来自促成这一预测的所有其他车队变量。
分析师
智能,我怎么没有。我如何通过将摄像头与神经网络结合使用而获得低功耗的益处。我就是在这里无法把这两者联系起来。也许只是我的问题,但我想那就是。
埃隆·马斯克
我是说,全自动驾驶计算机的算力令人难以置信。
埃隆·马斯克
也许我们应该提一下,即使它以前从未见过那条道路,只要那是美国的一条道路,它仍然会做出那些预测。
分析师
这里是3月9日的案例。
分析师
就激光雷达而言,3月9日不是有一个例子吗?我只是想谈谈你对激光雷达的猛烈抨击,因为很明显,你不喜欢这场激光雷达口水战中的激光雷达。
埃隆·马斯克
激光雷达很差劲。
分析师
难道不会有这样的情况吗,在未来某个时候,9 9 9 9 9,激光雷达实际上可能会有帮助?为什么不把它作为某种冗余或备用设备?这是我的第1个问题。还有第2个。所以你仍然可以专注于计算机视觉,只是把它作为冗余。我的第2个问题是,如果这是真的,那么业内其他那些基于激光雷达构建自动驾驶解决方案的公司会怎么样?
埃隆·马斯克
他们全都会抛弃激光雷达。这是我的预测,记住我的话。
埃隆·马斯克
我应该指出,我其实并没有听起来那么超级讨厌激光雷达,但在SpaceX,SpaceX的龙飞船使用激光雷达导航至空间站并与之对接。不仅如此,SpaceX还为此从零开始开发了自己的激光雷达。而且我亲自牵头了这项工作,因为在那种场景下,激光雷达是合理的。而在汽车上,它蠢得要命。它既昂贵又没必要。而且正如安德烈所说,一旦你解决了视觉问题,它就毫无价值。
埃隆·马斯克
所以你的车上装着昂贵却毫无价值的硬件。我们确实有一部前向雷达,它成本低廉,而且尤其有助于应对遮挡情况。所以,如果有雾、灰尘或雪,雷达可以穿透它们进行探测。如果你要使用主动光子生成,就不要使用可见光波长,因为通过被动光学,你已经处理了所有可见光波长的东西。你要使用的是像雷达那样能穿透遮挡的波长。
埃隆·马斯克
所以激光雷达只是在可见光谱内主动生成光子。
埃隆·马斯克
如果你要主动生成光子,那就在可见光谱之外的雷达频谱中进行。所以,比如3.8毫米,相比400、700纳米,你将获得好得多的遮挡穿透能力,这就是为什么我们配有前向雷达;此外,除了8个摄像头和前向雷达,我们还有12个超声波传感器,用于获取近场信息。你只需要在前方使用雷达,因为那是唯一一个真正高速行进的方向。
埃隆·马斯克
所以就是这样。我的意思是,我们已经反复讨论过这件事。比如,我们确定自己的传感器套件选对了吗?还应该添加更多东西吗?不。
分析师
你好。就在这里。你刚才提到,你会向车队索取你在某些视觉方面寻找的信息。对此我有2个问题。听起来,车辆会进行一些计算,以确定要向你们发回哪类信息。这个假设正确吗?它们是实时进行计算,还是基于已存储的信息进行计算?
安德烈·卡帕西
是的。所以车辆绝对会在车上实时进行计算,而我们会等待,基本上是指定我们感兴趣的条件,然后那些车辆就会在那里进行计算。如果它们不这样做,我们就必须发送所有数据,并在我们的后端离线处理。我们不想这么做。所以所有这些竞争都发生在车上。
分析师
所以,基于那个问题,听起来你们处于一个非常有利的位置:目前拥有50万辆车,未来可能会有数百万辆车,它们本质上都是计算机,相当于为你们提供免费、几乎免费的数据中心,让你们进行计算。这对Tesla来说是一个巨大的未来机会吗?这是当前的、当前的机会,而这似乎还完全没有被计入任何东西。太不可思议了。
分析师
谢谢。
埃隆·马斯克
我们有425,000辆配备硬件2及以上版本的汽车,这意味着它们拥有全部8个摄像头、雷达和超声波传感器,并且至少配备了英伟达计算机,这基本上足以判断哪些信息重要、哪些不重要。把重要信息压缩成最显著的要素,并上传到网络用于训练。所以,这是对现实世界数据的大规模压缩。
分析师
你们拥有这种由数百万台计算机构成的网络,本质上就像是大型数据中心,是用于提供计算能力的分布式数据中心。你认为未来它会被用于自动驾驶以外的其他用途吗?
埃隆·马斯克
我想,它或许可以用于自动驾驶以外的某些事情。我们一直高度专注于自动驾驶。所以,你知道,等我们真正把这件事彻底做好之后,也许还会有其他用途,用到,你知道,数百万台、然后数千万台配备硬件3或4驾驶计算机的计算机。
埃隆·马斯克
是的,也许会有。
埃隆·马斯克
有可能。有可能。也许这里会有某种类似AWS的应用思路。这是可能的。
埃隆·马斯克
你好。
马特·乔伊斯
嗨,埃隆。我是Loop Ventures的马特·乔伊斯。我在明尼苏达州拥有一辆Model 3,那里经常下雪。
马特·乔伊斯
由于摄像头和雷达无法透过积雪看到道路标线,你们在技术上
马特·乔伊斯
解决这一挑战的策略是什么?其中是否会用到高精度GPS?
安德烈·卡帕西
是的。
安德烈·卡帕西
所以
安德烈·卡帕西
实际上,就像现在,实际上,Autopilot在雪中也能做得相当、相当不错。即使道路标线被覆盖,即使房东标记褪色、被覆盖,或者上面有大量雨水,我们看起来仍然能开得比较好。我们还没有通过数据引擎专门处理积雪,但我确实认为这完全是可以解决的,因为在许多这样的图像中,即使到处都是雪,当你询问人类标注员车道线在哪里时,他们实际上也能告诉你;他们在绘制这些车道线时其实相对一致。
安德烈·卡帕西
只要标注员对你数据的标注保持一致,那么我认为,存在。神经网络就会捕捉到那些模式,而我们会表现得很好。所以真正的问题只是,即使对人类标注员而言,那里是否存在信号?如果,如果答案是肯定的,那么神经网络就完全可以做到。
埃隆·马斯克
是的,实际上有许多重要信号,正如安德烈所说。车道线就是其中之一,但最重要的信号之一是可行驶空间。也就是什么是可行驶空间,什么不是可行驶空间?而真正最重要的其实是可行驶空间,而不是车道线。对可行驶空间的预测极其准确。我认为,特别是在即将到来的这个冬季过后,它会令人难以置信。
埃隆·马斯克
这就像,它会让人觉得,怎么可能这么好?太疯狂了。
安德烈·卡帕西
另一件需要指出的事情是,也许这甚至并不只与人类标注员有关。只要你作为人类能够驾车通过那个环境,通过车队学习,我们实际上就知道你走过的路径。而你显然是用视觉来引导自己走过那条路径的。你并不只是使用车道线标记,而是使用了整个场景的完整几何结构。所以你会看到道路大致如何弯曲。你会看到周围车辆的位置。
安德烈·卡帕西
神经网络会在其内部自动捕捉所有这些模式。只要你拥有足够多的人穿越这些环境的数据。
埃隆·马斯克
是的,实际上,让事物不与GPS刚性绑定极其重要,因为GPS误差可能会有很大变化,道路的实际状况也可能会有很大变化。所以可能会有施工,可能会有绕行路线,而如果汽车将GPS作为主要依据,这种情况就非常糟糕。这是在自找麻烦。把GPS用于类似提示和技巧是没问题的。所以这就像是,与其他国家的某个街区或本国其他地区的某个街区相比,你在自己家附近能开得更好。
埃隆·马斯克
所以你很熟悉自己的街区,并且会利用某种类似对自己街区的了解,更有信心地驾驶,也许会走一些违反直觉的捷径之类的。但你。
埃隆·马斯克
GPS叠加数据应该只起辅助作用,而绝不能成为主要依据。如果它成为主要依据,那就是个问题。
皮特·班农
那么,请后面角落里的那位提问。
分析师
角落。我只是想部分地接着追问
分析师
这一点,因为你们的几家竞争对手
分析师
在这个领域过去几年里
分析师
已经做出了,你知道,已经谈到
分析师
他们如何增强所有
分析师
那些某种程度上位于汽车平台上的感知和路径规划能力
分析师
方法是使用他们所行驶区域的高清地图。
分析师
这在你们的系统中发挥作用吗?你认为它能增加任何价值吗?
分析师
是否有一些领域是你们希望获得、获得更多数据的,而这些数据并非从车队收集,而是更类似于制图风格的数据?
埃隆·马斯克
我认为高精度、高精度GPS地图和车道是一种非常糟糕的想法。系统会变得极其脆弱。所以像这样的任何变化都可能,系统的任何变化都会使它,它无法适应。所以,如果它锁定GPS和高精度车道线,并且不允许视觉覆盖,事实上,视觉应该是完成一切的东西,然后,像车道线只是指导,但不是主要依据。
埃隆·马斯克
我们曾短暂扩充高精度车道线树,随后意识到那是一个巨大的错误,并将其撤销了。
埃隆·马斯克
这不好。
斯图尔特·鲍尔斯
所以这对于理解非常有帮助
安德烈·卡帕西
标注,也就是物体在哪里,以及汽车如何行驶。但协商方面呢,对于
皮特·班农
停车、环岛以及其他方面,在这些地方
斯图尔特·鲍尔斯
道路上还有其他由人类驾驶的汽车,这更多是艺术而非科学。
埃隆·马斯克
实际上它做得相当不错。比如遇到加塞之类的情况。它表现得真的很好。
发言者
是的。
安德烈·卡帕西
所以就像我提到的,我们目前正在大量使用机器学习,来预测,算是创建一个关于世界是什么样子的显式表征。然后在这个表征之上有一个显式规划器和一个控制器。对于如何通行、协商等等,有很多启发式规则。就像一个……一样,存在长尾。在视觉环境是什么样子这方面。仅仅这些协商中也存在长尾,还有你与其他人进行的一点胆小鬼博弈等等。
安德烈·卡帕西
所以我认为,我们非常确信,最终你实际处理这件事的方式中必然要有某种车队学习组件。因为手工编写所有这些规则将会,将会很快到达平台期,我认为。
埃隆·马斯克
是的,我们已经处理过加塞这个问题,做法大概是允许用户逐步选择更激进的驾驶行为。他们只需把设置调高,设为更激进一些或更不激进一些。你知道,轻松驾驶、悠闲模式、激进。
发言者
是的。
分析师
进展令人难以置信。非常了不起。两个问题。首先,就编队行驶而言,你们的系统是否为此做了准备,因为有人问到道路上有积雪时的情况,但如果你有编队行驶功能,就可以直接跟随前车。你们的系统,你们的系统能做到这一点吗?然后我还有两个后续问题。
安德烈·卡帕西
所以你问的是编队行驶。我认为,我们完全可以构建这些功能。但同样,如果你只是使用,如果你只是训练神经网络,例如让它模仿人类,人类本来就会跟随前车。因此,那个神经网络实际上会在内部纳入这些模式。它就是,它会发现你前方汽车的朝向与你将要走的路径之间存在相关性。
安德烈·卡帕西
但这一切都是在网络内部完成的。所以你只需要关注获得足够的数据和那些棘手的数据。而神经网络训练过程实际上颇为神奇。会自动完成所有其他事情。所以你把所有不同的问题都转化成一个问题。只要收集你的数据集并使用神经网络训练。
埃隆·马斯克
是的,自动驾驶有三个步骤。你知道,先是功能完备,然后是在我们认为车内人员不需要集中注意力的程度上实现未来完备。然后是达到某种可靠性水平,让我们也说服监管机构这确实是真的。所以大概有三个层级。我们预计今年实现自动驾驶功能完备,而且从我们的角度来看,我们预计大概在,我不知道,明年第二季度左右,会有足够的信心说,我们认为人们不需要触碰方向盘,也不需要透过方向盘车窗向外看。
埃隆·马斯克
然后,我们预计到明年年底前后,至少会在一些司法管辖区获得这方面的监管批准。这大致就是我预计事情推进的时间表。而对于卡车,编队行驶可能会先于其他任何事项获得监管机构批准。如果你是在进行长途货运,可能可以让1名司机驾驶最前面的车,然后让4辆半挂卡车以编队方式跟在后面。
埃隆·马斯克
而且我认为,监管机构批准这个的速度可能会比其他事项更快。
分析师
关于。当然,你不必说服我们。在我看来,激光雷达是一项已经有答案的技术。却在寻找一个问题?可能已经死了。
发言人
这个。
分析师
我是说,我们今天看到的内容非常令人印象深刻,而且演示可能还能展示更多。我想知道,在你们的训练或深度学习流水线中,可能使用的矩阵最大维度是多少?
安德烈·卡帕西
大致数字,矩阵的最大估算值。那么,是的,神经网络内部会进行大量矩阵乘法运算。你问的是,这个问题有很多种不同的回答方式,但我并不100%确定它们是否。它们很有用。它们是有用的答案。正如我提到的,这些神经网络通常会有大约1000万到数亿个神经元。每个神经元平均与下层神经元有大约1000个连接。
安德烈·卡帕西
所以,这些就是整个行业通常采用的典型规模,也是我们会采用的规模。
分析师
是的。过去1年里,我的Model 3上的Autopilot改进速度确实令我印象非常深刻。我想听听你们对上周遇到的2种情形有何看法。第1种情形是,我当时在高速公路最右侧车道上,那里有一条高速公路入口匝道。然后我的Model 3实际上能够检测到旁边的2辆车,减速并让1辆车驶到我前面,另1辆车驶到我后面。我当时就想,天啊,这简直太疯狂了。
分析师
我完全没想到我的Model 3能做到这一点。所以那真的非常令人印象深刻。但同一周还有另一种情形,我又是在右侧车道上,但我所在的右侧车道正在并入左侧车道。那不是入口匝道,只是一条普通的高速公路车道。而我的Model 3没能真正检测到那种情况,我没法让它减速或加速,最后不得不进行某种干预。所以,从你们的角度,能否介绍一下其中的背景,比如神经网络会如何处理、Tesla可能会如何对此作出调整,以及,你知道,这种情况如何能随着时间推移得到改善?
安德烈·卡帕西
是的。正如我提到的,我们有一套非常复杂的触发基础设施。如果你进行了干预,我们实际上很可能已经收到了那段片段,这样我们就能分析它,看看发生了什么,并调整系统。因此,它可能会进入某些统计数据。好,我们正确汇入车流的比率是多少?然后我们查看这些数字和片段,看看出了什么问题,并尝试修复这些片段,针对这些基准取得进展。
安德烈·卡帕西
所以,是的,是的。我们可能会经历一个分类阶段,然后查看其中一些最大的类别,它们实际上看起来在语义上与同一个问题相关。接着我们会研究其中一些类别,并尝试针对它们开发软件。
埃隆·马斯克
好的。我们确实还有1场关于软件的演讲。所以,基本上,前面是斯图尔特讲解Autopilot硬件,然后是安德烈讲解某种神经网络视觉,接着斯图尔特将介绍大规模软件工程。谢谢。之后我们还会有机会提问。那么,是的,谢谢。
投资者关系
我只想非常简短地说一下,如果你的航班比较早,又想使用我们最新的开发版软件进行试乘,请与我的同事安联系,或者给她发一封电子邮件,我们可以带你出去试乘。斯图尔特,交给你了。
斯图尔特·鲍尔斯
好的,这实际上取自一段超过30分钟、不间断且没有任何干预的驾驶视频,车辆在高速公路系统上使用Navigate on Autopilot,而这一功能如今已经部署在数十万辆量产车上。我是斯图尔特,今天要讲的是我们如何大规模构建其中一些系统。先非常简短地介绍一下我的背景和工作。所以我待过几家公司,或者更少。
斯图尔特·鲍尔斯
我从事专业软件开发大约已有12年。最令我兴奋、也让我真正充满热情的事情,是把机器学习的前沿成果通过稳健性和规模化真正连接到客户。所以在Facebook,我最初在我们的广告基础设施内部工作,构建部分机器学习系统,一些非常、非常聪明的人。我们实际上试图把它构建成一个统一平台,以便随后将其扩展到业务的其他各个方面,从如何对动态消息进行排序,到如何提供搜索结果,再到如何在整个平台上作出每一项推荐。
斯图尔特·鲍尔斯
后来这就成了应用机器学习组。这是我感到无比自豪的一件事。其中很大一部分不仅仅是核心算法以及在那里发生的真正重要的改进。那些,那很重要,很多实际上是为了大规模构建这些系统而采用的工程实践。我后来去的Snap也是如此,当时我们真的、真的很兴奋,想要以某种方式真正帮助这款产品实现商业化。
斯图尔特·鲍尔斯
但最困难的部分是,我们当时使用的是Google。而实际上,你知道,他们让我们在相当小的规模上运行。我们想构建同样的基础设施。我们获得对这些用户的理解,将其与前沿机器学习连接起来,以超大规模构建这一体系,并以真正稳健的方式每天处理数10亿、继而数万亿次预测和竞价。所以当来到Tesla的机会出现时,这正是让我无比兴奋去做的事情,具体来说,就是把硬件方面以及计算机视觉和AI方面正在发生的惊人成果,与所有规划、控制、测试、操作系统内核修补、我们的全部持续集成和模拟真正整合封装起来,并真正把它构建成一款产品。
斯图尔特·鲍尔斯
我们如今把它部署到人们的量产车辆上。所以我想讲讲我们是如何按照时间线为Navigate on Autopilot做到这一点的,以及随着我们把Navigate on Autopilot从高速公路带到城市街道上,我们将如何做到这一点。
斯图尔特·鲍尔斯
所以Navigate on Autopilot已经行驶了7000万英里,这是一件非常、非常、非常酷的事情。我认为这里有一点值得特别指出,那就是我们正在继续加速,并不断从这些数据中学习。就像安德烈所说的,随着这个数据引擎加速运转,我们实际上会执行越来越果断的变道。我们正在从这些人们进行干预的案例中学习,干预原因要么是系统未能正确检测到汇入情况,要么是他们希望车辆在不同环境下表现得更有劲一些。
斯图尔特·鲍尔斯
而我们只想继续取得这样的进展。所以,要开始这一切,我们首先要设法理解周围的世界。我们谈到了车辆中的不同传感器。但我想再深入一点。这里我们有8个摄像头,此外还有12个超声波传感器、1个雷达、1个惯性测量单元和GPS。还有1件我们忘记的事,就是我们还拥有踏板和方向盘操作数据。
斯图尔特·鲍尔斯
所以,我们不仅可以查看车辆车辆周围正在发生的情况,还可以查看人类如何选择与该环境互动。现在我来讲讲这个片段。它基本上展示了目前车内正在发生的情况,而我们还在继续推动这方面向前发展。所以,我们从单个神经网络开始。我们看到它周围的检测结果。然后,我们用多个神经网络和多项检测结果把这一切整合起来。
斯图尔特·鲍尔斯
我们引入其他传感器,并将其转换成埃隆所说的向量空间,也就是对我们周围世界的理解。而在这方面,随着我们继续、继续做得越来越好,我们正把越来越多的这种逻辑移入神经网络本身。这里显而易见的最终目标是,神经网络查看所有车辆,把所有信息汇集起来,最终直接输出我们周围世界的一个事实来源。
斯图尔特·鲍尔斯
而且,从很多意义上说,这其实并不是艺术家绘制的渲染图。这实际上是我们团队每天使用的一款调试工具的输出,用来了解我们周围的世界是什么样子。所以,还有一件事我认为真的、真的让我感到兴奋,我想,当我听到激光雷达这类传感器时,一个常见问题就是,为什么不增加一些传感器模态,比如为什么不给车辆配置一些冗余?
斯图尔特·鲍尔斯
而我想深入谈谈一个并不。并不总是能从神经网络本身明显看出来的事情。比如说,我们有一个神经网络运行在我们的广角鱼眼摄像头上。这个神经网络并不是只对世界作出一项预测。它会作出许多相互独立的预测,其中一些实际上会互相核查。举一个真实的例子,我们有检测行人的能力。那是一个。我们会非常、非常谨慎地训练这项能力,并为此投入大量工作。
斯图尔特·鲍尔斯
我们还具备检测道路中障碍物的能力,而行人就是障碍物。它以不同的方式呈现给神经网络。它会说,哦,那里有个我不能直接开过去的东西。这些能力结合起来,让我们更清楚地了解在车辆前方什么能做、什么不能做,以及该如何为此进行规划。然后,我们在多个摄像头上这样做,因为车辆周围很多位置的视野相互重叠;在车辆前方,我们的重叠视野数量尤其多。
斯图尔特·鲍尔斯
最后,我们可以把它与雷达和超声波之类的东西结合起来,对汽车前方正在发生的情况形成极其精确的理解。我们既可以利用它来学习非常准确的未来行为,也可以对我们前方事态将如何继续发展作出非常准确的预测。我认为有一个例子非常令人兴奋:我们实际上可以观察骑自行车的人和行人,不只是问,你现在在哪里?
斯图尔特·鲍尔斯
而是你要去哪里?这实际上正是我们新一代自动紧急制动系统的核心所在,它不仅会为处在你行驶路径上的人停车,还会为即将进入你行驶路径的人停车。而它目前正在影子模式下运行。本季度我们会把它推送到车队,我稍后会讲一下影子模式。
斯图尔特·鲍尔斯
所以,当你想为高速公路系统上的自动辅助导航启动这样一项功能时,你可以从数据学习入手。你可以直接观察人类如今是怎么做的。他们的果断程度如何?他们如何变道?什么会导致他们要么接受,要么改变自己的操作?而且你可以看到一些并非立刻显而易见的事情,比如,哦,对,同时并线很少见,但非常复杂,也非常重要。
斯图尔特·鲍尔斯
而且,你可以开始形成对不同场景的判断,比如一辆快速超车的车辆。所以,当我们最初有一些想尝试的算法时,我们就是这样做的。我们可以把它们部署到车队中,看看在现实世界的场景中它们本会怎么做,比如这辆正在非常快速地超过我们的车。这取自我们真实的模拟环境,展示了我们考虑采用的不同、不同路径,以及这些路径如何叠加到一名用户在现实世界中的行为上。
斯图尔特·鲍尔斯
当你把那些算法调校好,并且对它们具体的表现感到满意时,而这实际上就是获取神经网络的输出,把它放进那个向量空间,并在其上构建和调校这些参数。最终,这是我们可以通过越来越多的机器学习来做的一件事。你会进入受控部署阶段,对我们而言就是我们的抢先体验计划。然后,你会把它推送给几千名非常兴奋的人,他们会针对它的表现方式给你高度警觉但有用的反馈,不是在开环中,而是在闭
安德烈·卡帕西
环方式下在现实世界中运行。
斯图尔特·鲍尔斯
然后你观察他们的干预。我们谈过这个,比如当有人接管时,我们实际上可以获取那个片段,尝试理解发生了什么。我们真正能做的一件事是,实际上可以再次以开环方式回放它,并且随着我们构建软件而问:我们是更接近还是更偏离人类在现实世界中的行为方式?还有一件特别酷的事是,借助全自动驾驶计算机,我们实际上正在构建自己的机架和基础设施,所以我们基本上可以把4台全自动驾驶计算机完整地装进机架,将它们构建到我们自己的集群中,并实际运行这个非常复杂的数据基础设施,从而真正理解随着时间推移,当我们调校和修复这些算法时,我们是否正越来越接近人类的行为方式?
斯图尔特·鲍尔斯
最终,我们能否超越他们的能力?所以,一旦我们有了这个,我们就对此感觉非常好。我们想进行广泛推送,但在开始时,我们实际上要求每个人通过拨杆确认来确认汽车的行为。于是,我们开始就应该如何在高速公路上行驶作出大量、大量的预测。我们请人们告诉我们,这是对还是错。这又是一次运转那个数据引擎的机会。
斯图尔特·鲍尔斯
而且,我们确实发现了一些非常棘手且有趣的长尾情况,在这个案例中,我认为一个非常有趣的例子就是这些非常有意思的同时并线情况:你开始行动,然后有人在没有注意到你的情况下移到你后面或前面。这里恰当的应对方式是什么?为了极其精确地实现恰当行为,我们需要对神经网络进行哪些调校?
斯图尔特·鲍尔斯
在这里,我们开展了工作,在后台调校这些内容,让它们变得更好,而随着时间推移,我们完成了900万次被成功接受的变道。我们再次结合持续集成基础设施利用这些数据,来真正理解我们是否认为自己已经准备好了。这也是全自动驾驶让我感到非常兴奋的一点。由于我们拥有整个软件栈,从内核补丁一直到ISO,比如图像信号处理器的调校,我们可以开始收集更多甚至更加准确的数据。
斯图尔特·鲍尔斯
而这让我们能够做得越来越好,通过这些更快的迭代周期进行调优。所以本月早些时候,我们有点觉得已经准备好部署一个更加无缝的高速公路自动辅助导航版本。而这个无缝版本不需要拨杆确认。所以你可以坐在那里,放松下来,把手放在方向盘上,只需监督车辆的行为。在这种情况下,我们实际上看到高速公路系统上每天都有超过100,000次自动变道。
斯图尔特·鲍尔斯
而对我们来说,能够大规模部署这个东西简直超级酷。在这一切当中,我最兴奋的事情大概是它的实际生命周期,以及我们实际上如何能够随着时间推移,越来越快地转动这个数据引擎的曲柄。而且我认为,有一件事确实正变得非常、非常清楚,那就是我们已经建成的基础设施、在其上构建的工具,以及完全自动驾驶计算机的综合力量;我相信,随着我们把高速公路系统上的自动辅助导航转移到城市街道上,我们可以更快地做到这一点。
斯图尔特·鲍尔斯
所以,是的,讲到这里,我把话筒交给埃隆。
埃隆·马斯克
是的,我是说,据我所知,所有这些变道都是在零事故的情况下完成的。
斯图尔特·鲍尔斯
没错。是的,每一起事故我都会看。
埃隆·马斯克
所以。所以它显然是保守的,但实现数十万乃至数百万次变道并且零事故,我认为是Tesla团队取得的一项伟大成就。
斯图尔特·鲍尔斯
谢谢。
埃隆·马斯克
酷。
埃隆·马斯克
那么让我们看看,你知道,还有其他几件可能值得一提的事情。为了拥有一辆自动驾驶汽车或机器人出租车,你确实需要车辆在硬件层面从头到尾都有冗余。所以从,也许是2016年10月开始,Tesla生产的所有汽车都有冗余的动力转向系统。所以我们的动力转向系统上有冗余电机。所以任何1个故障,如果电机发生故障,汽车仍然可以转向,所有电力和数据线路都有冗余,因此你可以切断任何1条指定的电力线路或任何数据线路,汽车仍会继续行驶,辅助供电系统,即使是主电池包,你的主电池包完全失去电力。
埃隆·马斯克
汽车能够利用辅助供电系统进行转向和制动,所以即使主电池包完全失效,汽车也是安全的。
埃隆·马斯克
从硬件角度来看,整个系统基本上从2016年10月起就是按照机器人出租车来设计的。
埃隆·马斯克
所以,当我们推出第2版自动辅助驾驶硬件时,我们不打算升级在那之前生产的汽车。我们认为,制造一辆新车实际上比升级这些汽车成本更高。只是为了让你们了解做这件事有多难。除非它一开始就被设计进去,否则不值得。
埃隆·马斯克
所以我们已经讲过了自动驾驶的未来,很明显,它涉及硬件、视觉,然后还有大量软件,而且这里的软件问题不应该被最小化。这是一个巨大的软件问题,它
分析师
是的,
埃隆·马斯克
管理海量数据,利用数据进行训练。如何根据视觉来控制汽车?这是一个非常困难的软件问题。
埃隆·马斯克
所以接下来回顾一下,就像Tesla总体规划一样,显然我们作出了一堆他们所谓的前瞻性陈述。
埃隆·马斯克
但我们来回顾一下我们作出的其他一些前瞻性陈述。很久以前,当我们创立公司时,我们说会制造Tesla Roadster。他们说这是不可能的,而且即使我们真的造了出来,也不会有人买。
埃隆·马斯克
这就像是普遍的看法,即制造电动汽车是极其愚蠢的,而且会失败。
埃隆·马斯克
我同意他们所说的失败概率很高,但这件事很重要。所以我们在2008年投产了Tesla Roadster并交付了那款车,它现在成了收藏品。
埃隆·马斯克
我们用Model S制造了一款价格更实惠的汽车。我们又做到了。我们被告知那是不可能的。我被称为骗子和说谎者。那不可能发生。这些全都不是真的。好吧,现在说句著名的事后不祥之言,我们在2012年投产了Model S,超出了所有预期。直到2019年,仍然没有汽车能够与2012年的Model S竞争。已经过去7年了,仍在等待。
埃隆·马斯克
所以我们会制造一款价格实惠的汽车,也许是非常实惠。它价格实惠。Model 3更加实惠。我们买了Model 3,我们正在生产。我说过Model 3的产量会超过每周5,000辆。到目前这个时候,每周5,000辆对我们来说轻而易举,甚至一点都不难。
埃隆·马斯克
所以我们做大规模太阳能,我们通过收购SolarCity做到了这一点,而且我们开发和部署太阳能屋顶,这进展得非常顺利。我们现在已经做到第3版太阳能瓦片屋顶,并且预计今年晚些时候会显著地溢出太阳能瓦片屋顶的产量。
埃隆·马斯克
我家里就装着它,而且很棒。
埃隆·马斯克
而且我算是制造了 Powerwall 和 Powerpack。我们制造了 Powerwall 和 Powerpack。事实上,Powerpack 现已部署于世界各地的大规模电网级公用事业系统中,包括全球规模最大的在运电池项目,功率超过100兆瓦。到明年,或者大概明年、最迟2年后,我们预计将完成一个吉瓦级电池项目。所以所有这些事情,我说过我们会做,我们做到了。
埃隆·马斯克
说过我们会做,我们做到了。我们也会做无人驾驶出租车这件事。唯一的批评是,有时我不能按时完成,而这个批评是公允的,但我会把事情做成,Tesla 团队也会把事情做成。
埃隆·马斯克
所以我们今年要做的是,让 S、X 和 3 的合计产量达到每周10,000辆。对此非常有信心,而且我们非常有信心在自动驾驶方面做到未来完整。
埃隆·马斯克
明年,我们将通过 Model Y 和 Semi 扩充产品线,而且我们预计明年会有首批投入运营的无人驾驶出租车,明年车里不会有任何人。
埃隆·马斯克
当事物呈指数式、以指数级速度改善时,总是很难理解。人们很难真正想明白,因为我们习惯于进行线性外推。但当你拥有海量的,比如硬件,路上有海量硬件时,累积数据会呈指数增长。软件也在以指数级速度变得更好。
埃隆·马斯克
我非常有信心预测,Tesla 明年会有自动驾驶出租车。不是在一个较老的州,也不是在所有司法管辖区,因为我们不会在所有地方都获得监管批准。但我有信心,实际上就在明年,我们至少会在某个地方获得监管批准。
埃隆·马斯克
所以任何客户都可以把自己的车加入 Tesla 网络,或者从中移除。我们预计它的运营方式可能会像 Uber 和 Airbnb 模式的结合。所以如果你拥有这辆车,就可以把它加入 Tesla 网络或从中移除,而 Tesla 会收取收入的25%或30%。而在没有足够多人共享汽车的地方,我们就会直接配备专用的 Tesla 车辆。
埃隆·马斯克
所以当你使用这辆车时,我们会向你展示我们的拼车应用。这样你就能从停车场召唤车辆,上车,然后开车出行。
埃隆·马斯克
这真的很简单。你只需使用目前已有的同一个 Tesla 应用。我们会更新该应用,加入召唤、召唤 Tesla,或者把你的车投入车队的功能。所以就是召唤你的车,或者召唤一辆 Tesla,或者把你的车加入车队或从车队中移除。你可以通过手机完成这些操作。
埃隆·马斯克
所以我们看到了平滑需求分布曲线的潜力。
埃隆·马斯克
并且让一辆车以比普通汽车高得多的利用率运行。通常,一辆车每周的使用时间约为10至12小时。也就是说,大多数人每天会开1.5至2小时,通常每周总计驾驶10至12小时。但如果你有一辆能够自动驾驶的汽车,那么很可能你大概可以。很可能你会让那辆车每周运行三分之一的时间或更久。一周有168小时。
埃隆·马斯克
所以每周的运行时间大概在55、60小时左右,也许还会更久一些。
埃隆·马斯克
所以一辆车的基础效用提高了5倍。你可以从宏观经济的角度来看这件事,然后说,就像这是某种。如果我们在运行某种大型模拟,而你可以升级模拟,使汽车的效用提高5倍,那将大幅提升这个模拟的经济效率。提升幅度极其巨大。
埃隆·马斯克
所以我们会做 Model 3s3 和过剩出租车。但我们对租赁做了一项重要调整。所以如果你租赁一辆 Model 3,租期结束时你没有购买它的选择权。我们希望把它们收回来。如果你购买这辆车,你可以留着它,但如果你租赁它,就必须归还。
埃隆·马斯克
正如我所说,在任何没有足够共享供给的地方,Tesla 都会直接制造自己的汽车,并将它们加入当地的网络。
埃隆·马斯克
所以目前 Model 3 无人驾驶出租车的成本低于38,000美元。我们预计这个数字会随着时间推移而改善,以及辞职。目前制造的汽车全都按行驶100万英里来设计。驱动单元按运行100万英里进行设计、测试和验证。目前的电池包大约能行驶300,000至500,000英里。可能会在明年投产的新电池包,则明确按运行100万英里来设计。
埃隆·马斯克
整车电池包,它被设计为只需极少维护便可运行100万英里。所以我们实际上还会调整轮胎设计,真正针对超高效率的无人驾驶出租车优化汽车。到某个时候,你将不再需要方向盘或踏板,我们就会直接删掉它们。所以随着、随着、随着这些东西变得越来越不重要,我们就会直接删掉这些部件。就是不会再有了。
埃隆·马斯克
比如说,可能从现在起2年后,我们会制造一辆没有方向盘或踏板的汽车;如果我们需要加快这个时间进度,随时可以直接删掉这些部件。很容易。
埃隆·马斯克
是的,从长期来看,大概3年后,删减了部件的无人驾驶出租车,最终价格也许会是25,000美元或更低。
埃隆·马斯克
而且我们想要一辆效率极高的汽车。这样耗电量就非常低。我们目前能做到每千瓦时行驶4.5英里。但我们可以,我们会把它提高到5英里乃至更高。
埃隆·马斯克
确实没有哪家公司拥有完整的全栈整合能力。我们拥有车辆设计和制造,但计算机硬件是内部开发的。我们拥有内部软件开发和 AI,而且我们的车队规模遥遥领先。当 Tesla 每天的行驶里程是其他所有公司总和的100倍时,想要追赶极其困难,也许并非不可能,但极其困难。
埃隆·马斯克
这是运行一辆汽油车或一辆……的成本。在美国运行一辆汽车的平均成本取自 AAA。所以目前约为每英里62美分。
埃隆·马斯克
15,000,000辆汽车各行驶13,500英里,加起来每年达到2万亿。这些数据确实就是取自 AAA 网站。
埃隆·马斯克
根据 Uber 和 Lyft 的数据,共享出行的成本是每英里2至3美元。我们认为,运营一辆无人驾驶出租车的成本低于每英里18美分。
埃隆·马斯克
而且还在下降。
埃隆·马斯克
这会是目前的。这是目前的成本。未来的成本会更低。你会问,一辆无人驾驶出租车可能产生多少毛利润?
埃隆·马斯克
我们认为每年大概在30,000美元左右。
埃隆·马斯克
而且我们预期会如此。我们确实是在设计,我们设计汽车的方式,与商用半挂车、半挂卡车的设计方式相同。商用半挂卡车全都是按100万英里寿命设计的,而我们也在按100万英里寿命设计汽车。
埃隆·马斯克
所以按名义美元计算,那会是,你知道,略高于300,000美元。在11年期间可能会更高。我认为这些消耗量实际上相对保守。而这假设行驶里程的50%是。没有什么不是有用的。所以这只按50%的利用率计算。
埃隆·马斯克
到明年年中,我们将有超过100万辆Tesla汽车上路,配备功能完整的完全自动驾驶硬件,其可靠性达到我们认为任何人都不需要留意的水平。意思是你可以去睡觉。从我们的角度来看,如果你把时间快进1年,也许1年,也许1年零3个月。
埃隆·马斯克
但可以肯定的是,明年我们将有超过1,000,000辆无人驾驶出租车上路。
埃隆·马斯克
整个车队通过一次无线更新被唤醒。仅此而已。
埃隆·马斯克
你会问,一辆无人驾驶出租车的净现值是多少?可能大约是20万美元。所以购买一辆Model 3很划算。
埃隆·马斯克
有什么问题吗?
埃隆·马斯克
嗯,我是说,在我们自己的车队中,我不知道,我想长期来看,我们可能会有大约1,000万辆车。
埃隆·马斯克
我是说我们通常的生产速率。如果你看看我们自2012年以来的复合年生产率,那就像是。那是我们全面生产Model Model S的第1个完整年份。我们从2013年生产23,000辆车,发展到去年生产约250,000辆车。所以在5年期间,我们将产量提高了倍数,10倍。我预计未来5年或6年会发生类似的情况。
埃隆·马斯克
至于共享,共享还是,我不知道。好处在于,基本上是客户在预先把车款给我们。这很棒。
斯图尔特·鲍尔斯
所以就那一件事而言,就是蛇形充电器,我对此很好奇。
安德烈·卡帕西
还有,你们是如何确定定价的?
斯图尔特·鲍尔斯
看起来,看起来你们的价格比平均的Lyft或Uber行程低了大约50%。所以我很好奇,你能否谈谈
安德烈·卡帕西
一点定价策略。
埃隆·马斯克
当然。我们预计这要解决。解决蛇形充电器的问题,从视觉道具的角度来看相当直接。这就像一种已知情形。对视觉来说,任何一种已知情形,比如充电口,都很简单。
埃隆·马斯克
所以。所以,是的,汽车会自动停车并自动接入充电。不会有任何人,也不需要人工监督。
埃隆·马斯克
是的。所以抱歉,定价是多少来着?我们只是在上面随便放了几个数字。我的意思是,我认为肯定包括插电。你觉得什么定价合理就用什么定价。我们只是有点随意地说,好吧,也许1美元。
埃隆·马斯克
而问题是,全世界大约有20亿辆汽车和卡车。所以在非常长的时间里,Robotaxi 的需求都会极其旺盛。而我到目前为止的观察是,汽车行业的适应速度非常慢。我的意思是,就像我说的,如今你能买到的上路汽车中,仍然没有一辆能像2012年的 Model S 那么好。
埃隆·马斯克
所以这表明汽车行业的适应速度相当缓慢。因此未来10年里,1美元可能还是保守的,因为人们多少会觉得,实际上大家对制造的难度认识不足。制造难得不可思议。但我交谈过的很多人都认为,只要有正确的设计,你就能立刻制造出全世界想要的那么多产品。
埃隆·马斯克
这不是真的。
埃隆·马斯克
为新技术设计一套新的制造系统极其困难。
埃隆·马斯克
我的意思是,奥迪在制造 E Tron 时遇到了重大问题,而,而且他们极其擅长制造。如果他们都遇到问题,其他人又会怎样?
埃隆·马斯克
所以,这个,你知道,全世界大约有20亿辆汽车和卡车,汽车年产能大约为1亿辆,但只能生产旧设计的汽车。
埃隆·马斯克
要把所有这些都转变为完全自动驾驶汽车,需要非常长的时间。而且它们确实需要是电动的,因为汽油柴油汽车的运营成本远高于电动汽车。所以任何、任何、任何不是电动的机器人出租车都绝对不具备竞争力。
科林·拉什
埃隆,我是这边奥本海默的科林·拉什。你知道,显然,我们理解客户正在为这支车队的组建预付一部分现金,但听起来随着时间推移,这需要该组织在资产负债表上作出巨额投入。你能稍微谈谈
科林·拉什
这会是什么样子,你们的预期
科林·拉什
是什么,就未来的融资而言,在接下来
科林·拉什
就算3年吧,3、4年
科林·拉什
里组建这支车队,并开始通过你们的客户群体将其变现?
埃隆·马斯克
嗯,我们的目标是在车队组建阶段实现大致的现金流中性。然后,我预计一旦机器人出租车启用,我们将产生极其可观的正现金流。但我不想谈融资轮次。在这个场合谈论融资轮次会很困难。不过我认为我们会采取正确的行动。我认为我们会采取你们认为我们应该采取的行动。
分析师
我有一个问题。如果我是 Uber,我为什么不直接买下你们所有的车?你知道,我为什么要让你们把我挤出市场?
埃隆·马斯克
我们的车里有一项,有一项条款。我想那是大约3年或4年前加进去的。它们只能在 Tesla 网络中使用。
分析师
所以即使是个人,比如我出去买10辆 Model 3,我不能,我可以在这个网络上运营。这现在是一门生意。对吧。
埃隆·马斯克
你只有权使用 Tesla 网络。
分析师
对。但如果我使用 Tesla 网络,理论上,我可以用我的10辆 Model 3 经营一家汽车共享机器人出租车业务。
埃隆·马斯克
可以,但它就像 App Store。
埃隆·马斯克
你只能通过 Tesla 网络添加或移除它们,然后 Tesla 会获得收入分成。
分析师
不过这和 Airbnb 类似,我有这套房子、我的车,现在我可以直接把它们租出去,这样我就能通过拥有多辆车并把它们租出去赚取额外收入。比如我有一辆 Model 3。我希望你们造出这里这辆 Roadster 后,下一辆就买它。而我会直接把我的 Model 3 租出去。我为什么要把它还给你们?你知道,
埃隆·马斯克
我想你可以运营一支租车车队,但我认为这样非常难以管理。是的,我不知道。看起来很容易。好吧,试试看。
分析师
要运营机器人出租车网络,听起来你们必须解决某些问题,比如说,现在使用 Autopilot 时,如果你过度转动方向盘,它会让你接管。但如果它是,你知道,如果它是一款让其他人坐进乘客座位的共享出行产品,那么像转动方向盘这种操作,方向盘不能让那个人接管汽车,例如,因为他们甚至可能不在驾驶座上。所以,要让它成为机器人出租车,硬件已经具备了吗?
分析师
而且它可能遇到警察示意它靠边停车之类的情况,届时可能需要某个人进行干预,比如使用一个中央操作员团队,由他们远程以某种方式与人互动,或者,我的意思是,每辆车中是否已经内置了所有这类基础设施?
埃隆·马斯克
这样说清楚了吗?我认为会有某种呼叫总部的机制,汽车如果被困住了,就会直接呼叫 Tesla 总部并寻求解决方案。比如被离岸警察拦下之类的事情。对我们来说,这很容易通过编程实现。这不是问题。
埃隆·马斯克
至少在一段时间内,某些……某个人可以使用方向盘接管。然后可能再往后,我们会直接把方向盘封住,这样就无法进行转向控制。
埃隆·马斯克
我们会直接拆下方向盘,从长远来看装上一个盖子。给它大约几年时间。硬件
分析师
对汽车进行改装,以便让它实现这一点。
埃隆·马斯克
或者,是的,我们实际上只要把方向盘拆下来,再在目前方向盘手柄所在的位置装一个盖子。
分析师
但是,但是那,那是一种你们未来会推出的汽车。但今天这些汽车怎么办?方向盘是接管 Autopilot 的一种机制,比如说,如果它处于机器人出租车模式,有人能不能只要简单转动一下方向盘之类的就接管它?
埃隆·马斯克
是的,我认为会有一个过渡期,人们将能够接管自动驾驶出租车,也应该能够接管。然后,一旦监管机构能够接受我们不设方向盘,我们就会直接把它去掉。对于那些已经加入车队的车辆,你知道,如果车辆归其他人所有,显然要征得车主的许可,我们会直接拆下方向盘,并在目前安装方向盘的位置盖上盖子。
分析师
所以,自动驾驶出租车可能会有两个阶段。一个阶段是提供这项服务,你上车后坐在驾驶员的位置,但有可能接管。然后未来可能不再提供驾驶员选项。你也是这样看的吗
埃隆·马斯克
或者说未来呢?未来会的。未来方向盘被拆掉的概率是100%。人们。消费者会要求这样做。
分析师
但是,但是一开始你会叫来。
埃隆·马斯克
这不是。这是,我要说明白,并不符合专业地规定一种关于世界的观点。这是我在预测消费者会要求什么。未来,消费者会要求不允许人们驾驶这些3吨重的死亡机器。
分析师
我完全同意。但为了让现在的一辆Model 3成为自动驾驶出租车网络的一部分,当你叫车时,你随后基本上会坐进驾驶座,因为只是为了保持一致。
埃隆·马斯克
好的,这就说得通了。
分析师
谢谢。
埃隆·马斯克
没错。就有点像,你知道,曾经有两栖动物,你知道,但后来差不多那些东西就变成了陆地生物。会有一个短暂的两栖阶段。
分析师
你好。
埃隆·马斯克
抱歉。我能看出那个。好的。
发言人
是的。
安德烈·卡帕西
我们从自动驾驶出租车领域的其他参与者那里听到的策略,是选择某个特定的市辖区域,创建有地理围栏的自动驾驶。这样一来,你就可以利用高清地图,在一个更受限制、稍微更安全的区域内运行。
发言人
一个。
安德烈·卡帕西
今天我们没有听到太多关于高清地图重要性的内容,对你们来说,高清地图在多大程度上是必要的?其次,我们也没有听到太多关于在特定市辖区部署这项技术的内容,也就是你们与市政府合作,以获得
分析师
他们的认同,而且你们还
安德烈·卡帕西
能得到一个界定更明确的区域。那么,高清地图的重要性是什么?你们在多大程度上考虑在特定市辖区推出?
埃隆·马斯克
我认为 HTMAPs 是个错误。其实我们有一段时间用过 HTMAPs。其实不能能那样做,因为要么你需要 HTMAPs,而在这种情况下,只要环境有任何变化,汽车就会出故障;要么你不需要 HTML,而在这种情况下,你为什么还要浪费时间做高清地图?所以高清地图这件事,就像两个不该使用的主要拐杖一样,事后回看会发现它们显然是错误而愚蠢的,那就是激光雷达和高清地图。
埃隆·马斯克
记住我的话。
分析师
你好。
埃隆·马斯克
如果你需要地理围栏区域,那你就不是真正的自动驾驶。
安德烈·卡帕西
只是听起来,电池供应可能是实现这一愿景仅存的瓶颈。另外,你能否说明一下,如何让电池包能使用100万英里?
埃隆·马斯克
我认为电芯会成为一个制约因素。那是一个完全独立的。那是一个完全独立的话题。
埃隆·马斯克
而且我认为,实际上我们会希望更多地推动某种标准续航升级版电池,而不是长续航电池,因为长续航电池包的能量含量按千瓦时计要高50%。
埃隆·马斯克
所以本质上,如果你,如果你只是。要是它们全都用某种标准续航升级版,而不是长续航电池包,你就能多生产三分之一的汽车。所以一个大约是50千瓦时,另一个大约是75千瓦时。因此,我们实际上可能会有意让销售偏向较小的电池包,以便拥有更多的,基本上你显然要做的是最大限度增加自动驾驶单元的数量,或者最大限度增加产出,这将替代性地带来未来最大的自动驾驶下泄。
埃隆·马斯克
所以我们正在这方面做很多事情,但这不属于今天会议的内容。
埃隆·马斯克
100万英里的寿命,基本上就是要把电池包的循环寿命做到,你知道,你大致需要。比如说,做个基本计算,如果你的电池包续航为250英里,你知道你会需要4,000次循环。
埃隆·马斯克
所以非常容易实现。我们的固定式储能已经做到了。我们的一些固定式储能解决方案,比如 Powerpack,我们已经准备好部署具备4,000次循环寿命能力的 Powerpack。
说话人
对。
分析师
我能问一下吗。
分析师
抱歉,对,我本来想。
埃隆·马斯克
这就像腹语术。
分析师
不,这显然影响重大。如果你能大幅提高全自动驾驶选装项的选购率,对利润率显然有非常积极的影响。我只是想知道,你能否说明一下目前这些选购率大致处于什么水平,以及你们预计将如何让消费者了解 Robotax 的情景,从而使选购率随着时间推移得到实质性提高提高。
埃隆·马斯克
抱歉,你的问题有点听不清。
分析师
对,只是想知道,就全自动驾驶选购率及其财务影响而言,我们目前处于什么水平。我认为,如果这些选购率显著上升,会带来巨大的好处,因为会有更高的毛利额流入。就人们确实订购完整 FSD 而言,只是想知道你认为它会如何增长
安德烈·卡帕西
或者选购
分析师
率目前是多少,相比之下,你知道,你们预计何时。你们预计如何教育消费者,让他们意识到自己在购车时应该选购 FSD?
埃隆·马斯克
今天之后,我们会大幅加大这方面的力度。
埃隆·马斯克
对,我的意思是,消费者今天应该领会的根本性、真正根本性的信息是:购买 Tesla 以外的任何东西,在财务上都是疯狂的。
埃隆·马斯克
3年后,这就会像拥有一匹马。我的意思是,如果你想拥有一匹马,那没问题,但你应该抱着这种预期去做。
埃隆·马斯克
如果你买了一辆不具备全自动驾驶所需硬件的汽车,那就像买了一匹马,而唯一具备全自动驾驶所需硬件的汽车就是 Tesla。
埃隆·马斯克
人们真的应该认真考虑自己购买任何其他车辆这件事。基本上,购买 Tesla 以外的任何汽车都是疯狂的。
埃隆·马斯克
我们需要把这个,把这个论点清楚地传达出去,今天之后我们会这样做。
分析师
太好了。感谢你们把未来带到当下,今天的信息量很大。我想问的是,你们没有怎么谈到 Tesla 皮卡,我来说明一下相关背景。我可能是错的,但在我看来,Tesla 网络是作为一个早期采用者,以及某种测试面包。我认为 Tesla 的皮卡可能是将车辆接入网络的第一阶段,因为 Tesla 皮卡的用途基本上会面向那些要装载很多东西的人,或者从事建筑行业的人,又或者处理这里那里零碎物品的人,比如从家得宝取东西。
分析师
我会说,你知道,也许需要一个两阶段流程,先让皮卡专供 Tesla 网络使用,作为起点。然后像我这样的人以后再购买。但你对此有什么看法?
埃隆·马斯克
嗯,今天其实只讨论自动驾驶。我们可以谈很多事情,比如电芯生产、皮卡和未来的车辆车辆。但今天只聚焦自动驾驶。不过我同意,这是件大事。我非常期待今年晚些时候发布 Tesla 皮卡。它会很棒。
斯图尔特·鲍尔斯
瑞银的 Colin Lang。为了确保我们理解这些定义,当你提到功能完整的自动驾驶时,听起来你说的是第5级、没有地理围栏。这是预计在今年年底前实现的吗?
埃隆·马斯克
只是为了让我们大家都。
斯图尔特·鲍尔斯
然后是监管流程,我的意思是
皮特·班农
你们就此和监管机构谈过吗?从
斯图尔特·鲍尔斯
其他人已经推出的东西。我是说,他们是不是,你知道,需要克服哪些障碍,获得批准的时间表又是什么?
安德烈·卡帕西
你们是否还需要诸如在
斯图尔特·鲍尔斯
加利福尼亚州,而且他们是否在追踪里程,你知道,有操作员在其后?你们需要这些东西吗?这个流程会是什么样的?
埃隆·马斯克
是的,我是说,随着我们推出诸如自动辅助驾驶导航之类的附加功能,我们一直都在与世界各地的监管机构沟通。
埃隆·马斯克
这需要逐个司法管辖区获得监管批准。
埃隆·马斯克
所以,不过根据我的经验,我认为从根本上说,数据能够说服监管机构。因此,如果你拥有海量数据,表明自动驾驶是安全的,他们就会听取这些数据。他们可能需要时间来消化他们所处理的信息。可能需要一点时间,但就我所见,他们最终总是会得出正确的结论。
安德烈·卡帕西
哦,这边有个问题。
埃隆·马斯克
我拿到了许可证和支柱。好的。
安德烈·卡帕西
我只是想,只是想,你知道,根据我们为更好地了解网约车市场所做的一些工作。看起来它高度集中在主要的高密度城市中心。那么,应该这样理解吗:无人驾驶出租车可能会更多地部署到这些地区,而个人拥有车辆所增加的完全自动驾驶功能则会用于郊区?
埃隆·马斯克
我认为大概是的,Tesla 自有的无人驾驶出租车会和客户车辆一起出现在高密度城市地区。然后到了中等和低密度地区,更多的情况往往会是人们拥有汽车,并偶尔把它借出去。
埃隆·马斯克
是的,曼哈顿以及比如旧金山市中心有很多边缘案例,但那些,你知道,而且世界各地有各种城市拥有颇具挑战性的开放环境。
埃隆·马斯克
但我们预计这不会是一个重大问题。而且当我说“未来完整”时,我的意思是,它今年将在旧金山市中心和曼哈顿市中心运行。
安德烈·卡帕西
你好,我有一个关于神经网络架构的问题。比如说,你们是否会为路径规划和感知使用不同的模型,或者使用不同类型的人工智能?你们大致是如何把这个问题拆分到自动驾驶的不同组成部分中的?
埃隆·马斯克
嗯,基本上目前人工智能或神经网络实际上是用于物体识别的。而且我们基本上仍然只是将其用于静止帧,也就是识别静止帧中的物体,然后在之后的感知路径规划层中把这些联系起来。但正在发生的情况是,神经网络正在逐渐越来越多地蚕食软件基础。因此随着时间推移,我们预计神经网络会做得越来越多。
埃隆·马斯克
现在,从计算成本的角度来看,有些事情对于启发式方法非常简单,而对于神经网络却非常困难。因此,在系统中保留某种程度的启发式方法可能是合理的,因为从计算角度看,它们比神经网络容易整整1000倍。比如,神经网络就像巡航导弹,而如果你想拍死一只苍蝇,就用苍蝇拍,不要用巡航导弹。
埃隆·马斯克
所以,不过随着时间推移,我预计它真的会转向仅用视频进行训练,然后输入视频,输出车内的转向和踏板操作,或者基本上几乎完全是输入视频,输出横向和纵向加速度。
埃隆·马斯克
这就是我们要使用 Dojo 系统来做的事。目前还没有任何系统能够做到这一点。
安德烈·卡帕西
也许这边。
斯图尔特·鲍尔斯
回到传感器套件的讨论,埃隆。有一个方面我想谈谈
埃隆·马斯克
的是缺少侧向雷达。
斯图尔特·鲍尔斯
在有停车标志的十字路口这种情况下,那里
安德烈·卡帕西
可能有时速35、40英里的
斯图尔特·鲍尔斯
横向车流,你是否确信
安德烈·卡帕西
传感器套件和侧向摄像头能够
分析师
应对这种情况?
斯图尔特·鲍尔斯
能不能简单谈谈这个问题?
埃隆·马斯克
可以,没问题。
埃隆·马斯克
从本质上说,汽车会采取类似人类的做法。你可以把人类看作基本上安装在慢速云台上的摄像头。人们能以现在这种方式驾驶汽车,着实相当了不起,因为你无法同时观察所有方向。汽车则确实可以通过多个摄像头同时观察所有方向。所以,人类只需像这样看看这边、看看那边,就能够开车。
埃隆·马斯克
他们实际上被固定在驾驶座上。他们其实无法离开驾驶座。所以这就像云台上的一个摄像头,却能够驾驶。尽职的驾驶员可以以非常高的安全性驾驶。车内摄像头的视角比人更好。它们位于 B 柱较高的位置或后视镜前方。它们确实拥有绝佳的视角。所以,如果你要转入一条有大量高速车流的道路,只需像人一样做。
埃隆·马斯克
只需稍微往前转一点。不要完全驶入道路。让摄像头看看情况。如果看起来没问题,而且后置摄像头也没有显示任何驶来的车辆,那就出发。如果看起来不太稳妥,你可以稍微往后退一点。就像人一样。这种行为非常像。它开始变得非常栩栩如生。其实相当诡异。汽车就开始像这边的人一样行动。
埃隆·马斯克
那我们开始吧。
斯图尔特·鲍尔斯
棘手的问题就在这里。
埃隆·马斯克
好的。
斯图尔特·鲍尔斯
鉴于你们通过将所有这些技术围绕自身整合,在汽车业务中创造了如此多的价值,我想知道,你们为什么仍会拿出一部分电芯产能投入 Powerwall 和 Powerpack?把你们能生产的每一个电芯都投入这部分业务,不是更合理吗?
埃隆·马斯克
我们已经抢走了几乎所有原本要用于 Powerwall 和 Powerpack 的电芯生产线,并将它们用于 Model 3。我的意思是,去年,为了完成 Model 3 的生产并避免自己陷入短缺,我们不得不把超级工厂的所有2170生产线都转用于汽车销售。
埃隆·马斯克
所以,以总吉瓦时计算,我们固定式储能的实际产出与汽车相比相差一个数量级。而对于固定式储能,我们基本上可以使用市面上一大批各种各样的电芯。因此,我们可以从世界各地的多家供应商那里收集电芯,而且不会像汽车那样存在认证问题或安全问题。所以基本上,我们的固定式电池业务很长时间以来一直都只是在捡零碎的东西吃。
埃隆·马斯克
所以。
埃隆·马斯克
但是,真的要把生产想成是。有很多很多因素会制约一个庞大的生产系统。制造业供应链被低估的程度令人震惊。存在一整套制约因素。而某一周的制约因素,到另一周可能就不再是制约因素。
埃隆·马斯克
制造一辆汽车难得不可思议,尤其是制造一辆正在快速演进的汽车。所以。
分析师
是的。
埃隆·马斯克
不过我再回答几个问题,然后我想我们就休息四,这样你们就可以试驾这些车了。
皮特·班农
你好,埃隆,亚当,乔纳斯。
斯图尔特·鲍尔斯
关于安全性的问题。
分析师
什么。
斯图尔特·鲍尔斯
你们今天能与我们分享什么数据?
皮特·班农
这项技术有多安全,这将
安德烈·卡帕西
显然在监管或保险讨论中很重要。
埃隆·马斯克
嗯,我们每个季度都会公布每英里事故数。我们目前看到的是,Autopilot 的安全性大约是普通驾驶员,你知道,平均而言普通驾驶员的2倍。我们预计这一数字会随着时间推移大幅提高。
埃隆·马斯克
就像我说的,未来会是这样。消费者会想要禁止,而且我是在说他们会成功,我也不是在说我赞同这一立场,但在未来,消费者会希望禁止人们自己驾驶汽车,因为那不安全。如果你想想电梯,电梯过去是通过一根大操纵杆来操作的,比如上下楼层,还有一个大继电器,也有电梯操作员,但他们有时会疲倦或喝醉之类的,然后会在错误的时机扳动操纵杆,把某个人拦腰切成两半。
埃隆·马斯克
所以现在已经没有电梯操作员了。而且,如果你走进一部装有一根大操纵杆、可以在楼层之间任意移动的电梯,那会相当令人不安。
分析师
所以。
埃隆·马斯克
所以就只有按钮,而从长远来看,再说一次,这并不是价值判断。我不是说我希望世界变成这样。我是说,消费者很可能会要求不允许人们驾驶汽车。
皮特·班农
还有,埃隆,追问一下,你能否
斯图尔特·鲍尔斯
和我们分享一下 Tesla 花费了多少
安德烈·卡帕西
每年在 Autopilot 或自动驾驶技术上的支出大致是什么数量级?谢谢。
埃隆·马斯克
这基本上就是我们的整个费用结构。
埃隆·马斯克
关于Tesla网络的经济性有一个问题。我只是想确认自己理解了,看起来是这样。所以你拿到一辆租赁期满的Model 3,25,000美元计入资产负债表,会是一项资产,然后你。它每年会产生大约30,000美元的现金流。应该这样理解吗。对,差不多是这样,对。
安德烈·卡帕西
然后仅就其融资而言
分析师
之前有一个问题,你
埃隆·马斯克
提到你会这么做。这对 Robotaxi 项目而言是现金流中性,还是对
分析师
Tesla 整体而言?
埃隆·马斯克
抱歉,现金流中性是就什么而言。
安德烈·卡帕西
他问了一个关于为
埃隆·马斯克
机器人税融资的问题,但在我看来
分析师
它们像是自我融资的。
埃隆·马斯克
但你提到它们基本上会实现现金流中性。你指的是这个吗?我只是说,从现在到Robotaxi在全世界全面部署之时,对我们来说明智的做法是最大化速率,并推动公司实现现金流中性。
埃隆·马斯克
一旦Robotaxi车队投入运营,我预计现金流会极其充裕。所以你说的是生产,对。把它们全部生产出来。
安德烈·卡帕西
好的,谢谢。
埃隆·马斯克
最大化所生产的自动驾驶车辆数量。
分析师
谢谢。
埃隆·马斯克
好的,也许就再来一个。这里最后一个问题。你好。
分析师
如果我。
安德烈·卡帕西
如果我把我的Tesla加入Robotaxi网络,谁。
埃隆·马斯克
谁要为事故负责?
分析师
是Tesla,还是我?
皮特·班农
如果车辆发生事故并造成伤害。
埃隆·马斯克
可能是Tesla。可能是Tesla。
说话人
如果。
埃隆·马斯克
是的。
埃隆·马斯克
我认为正确的做法就是确保事故非常、非常少。好了,谢谢大家。请享受试乘。
斯图尔特·鲍尔斯
谢谢。
安德烈·卡帕西
谢谢你,非常。
说话人
多。Sam。
说话人
萨。
说话人
Sam。
说话人
它。
说话人
Sam。
分析师
萨。
说话人
它。
说话人
它。
说话人
它是。
说话人
萨。
说话人
萨姆。
说话人
它。
说话人
拉。
Speaker
It.
Speaker
Sa.
Speaker
Sam.
Speaker
It.
Speaker
Range sam it.
Speaker
Sa.
Pete Bannon
Sa.
Speaker
Race.
Speaker
Sa.
Speaker
Sa.
Speaker
Sa.
Speaker
Sam.
Speaker
Sa.
Speaker
Sam.
Analyst
Sa.
Speaker
It.
Speaker
Sam.
Speaker
Sa.
Speaker
Sam.
Speaker
It.
Speaker
Sa.
Speaker
Sam sa.
Speaker
Sam.
Speaker
Sa.
Speaker
Sa.
Speaker
Sa.
Speaker
Sa.
Speaker
It.
Speaker
Sam.
Speaker
Sa.
Speaker
Sa.
Speaker
Sam.
Speaker
Sa.
Speaker
Sa.
Speaker
Sa.
Speaker
It.
Speaker
Sam.
Analyst
Sa.
Speaker
Sam.
Speaker
Sa.
Speaker
It.
Speaker
Sa.
Investor Relations
Hi everyone. I'm sorry for being late. Welcome to our very first analyst day for Autonomy. I really hope that this is something we can do a little bit more regularly now to keep you posted about the, the development we're doing with regards to autonomous driving. About three months ago we were getting prepped up for our Q4 earnings call with Elon and quite a few other executives. And one of the things that I told the group is that from all the conversations that I keep having with investors on regular basis, the biggest gap that I see with what I see inside the company and what the outside perception is is our ability of autonomous driving.
Investor Relations
And it kind of makes sense because for the past couple of years we've been really talking about Model 3 ramp. And you know, a lot of the debate has revolved around Model 3, but in reality a lot of things have been happening in the background. We've been working on the new full self driving chip. We've had a complete overhaul of our neural net for vision recognition, et cetera. So now that we finally started to produce our full self driving computer, we thought it's a good idea to just open the veil, invite everyone in and talk about everything that we've been doing for the past two years.
Investor Relations
So about three years ago we wanted to use, we wanted to find the best possible chip for Full Autonomy. And we found out that there's no chip that's been designed from ground up for neural nets. So we invited my colleague Pete Bannon, the VP of Silicon Engineering, to design such chip for us. He's got about 35 years of experience of building chips and designing chips. About 12 of those years were for a company called PA Semi, which was later acquired by Apple.
Investor Relations
So he worked on dozens of different architectures and designs and he was the lead designer, I think for Apple iPhone 5 just before joining Tesla. And he's going to be joined on the stage by Elon Musk. Thank you.
Elon Musk
Actually, I was going to introduce Pete, but Martin Stunt. So he's just the best chip and system architect that I know in the world. And it's an honor to have you and your team at Tesla and take away. Just tell them about the incredible work that you and your team have done.
Pete Bannon
Thanks, Elon. It's a pleasure to be here this morning and a real treat really to tell you about all the work that my colleagues and I've been doing here at Tesla for the last three years.
Pete Bannon
I think we'll tell you a little bit about how the whole thing got started and then I'll introduce you to the full self driving computer and tell you a little bit about how it works. We'll dive into the chip itself and go through some of those details. I'll describe how the custom neural network accelerator that we designed works and then I'll show you some results and hopefully you'll all still be awake by then.
Pete Bannon
I was hired in February of 2016. I asked Elon if he was willing to speak all the money it takes to do full custom system design. And he said, well, are we going to win? And I said, well, yeah, of course. So he said I'm in. And so that got us started. We hired a bunch of people and started thinking about what a custom designed chip for full autonomy would look like. We spent 18 months doing the design and in August of 2017 we released the design for manufacturing.
Pete Bannon
We got it back in December at pack powered up and it actually worked very, very well on the first try. We made a few changes and released a B0 Rev in April of 2018. In July of 2018, the chip was qualified and we started full production of production quality parts. In December of 2018, we had the autonomous driving stack running on the new hardware and we were able to start retrofitting employee cars and testing the hardware and software out in the real world.
Pete Bannon
Just last March we started shipping the new computer in the Model S and X. And just earlier in April we started production in the Model 3. So this whole program, from the hiring of the first few employees to having it in full production in all three of our cars, is just a little over three years and is probably the fastest system development program I've ever been associated with. And it really speaks a lot to the advantages of having a tremendous amount of vertical integration to allow you to do concurrent engineering and speed up deployment.
Pete Bannon
In terms of goals, we were totally focused exclusively on Tesla requirements and that makes life a lot easier. If you have one and only one customer, you don't have to worry about anything else. One of those goals was to keep the power under 100 watts so that we could retrofit fit the new machine into the existing cars.
Pete Bannon
We also wanted a lower part cost so we could enable full redundancy for safety. At the time, we had a thumb in the wind estimate that it would take at least 50 trillion operations. A second of neural Network performance to drive a car. And so we wanted to get at least that much and really as much as we possibly could. Batch size is how many items you operate on at the same time. So for example, Google's TP has a batch size of 256 and you have to wait around until you have 256 things to process before you can get started.
Pete Bannon
We didn't want to do that, so we designed our machine with a batch size of one. So as soon as an image shows up, we process it immediately. To minimize latency, which maximizes safety, we needed a GPU to run some post processing. At the time we were doing quite a lot of that. But we speculated that over time the amount of post processing on the GPU would decline as the neural networks got better and better. And that has actually come to pass.
Pete Bannon
So we took a risk by putting a fairly modest GPU in the design, as you'll see, and that turned out to be a good bet. Security is super important. If you don't have a secure car, you can't have a safe car. So there's a lot of focus on security and then of course, safety in terms of actually doing the chip design. As Elon alluded earlier, there was really no ground up neural network accelerator in existence in 2016. Everybody out there was adding instructions to their CPU or GPU or DSP to make it better for inference, but nobody was really just doing it natively.
Pete Bannon
So we set out to do that ourselves. And then for other components on the chip we purchased industry standard IP for CPUs and GPUs. That allowed us to minimize the design time and also the risk to the program.
Pete Bannon
Another thing that was a little unexpected when I first arrived was our ability to leverage existing teams at Tesla. Tesla had wonderful power supply design teams, signal integrity analysis, package design, system software, firmware board designs, and a really good system validation program that we were able to take advantage of to accelerate this program. Here's what it looks like.
Pete Bannon
Over there on the right you see all the connectors for the video that comes in from the eight cameras that are in the car. You can see the two self driving computers in the middle of the board and then on the left is the power supply and some control connections. And so I really love it when a solution is boiled down to its barest elements. You have video computing and power and it's straightforward and simple. Here's the original hardware 2.5 enclosure that the computer went into and we've been shipping for the last two years.
Pete Bannon
Here's the new design for the FSD computer. It's basically the same, and that of course is driven by the constraints of having a retrofit program for the cars. I'd like to point out that this is actually a pretty small computer. It fits behind the glove box. Between the glove box and the firewall in the car. It does not take up half your trunk.
Pete Bannon
As I said earlier, there's two fully independent computers on the board. You can see them there highlighted in blue and green. To either side of the large SoC you can see the DRAM chips that we use for storage. And then below left you see the flash chips that represent the file system. So these are two independent computers that boot up and run their own operating system.
Elon Musk
Yeah. If I can add something, the general principle here is that any part of this could fail and the car will keep driving. So you could have cameras fail, you could have power circuits fail, you could have one of the Tesla full self driving computer chips fail, car keeps driving. The probability of this computer failing is substantially lower than somebody losing consciousness. That's the key metric, at least in order of magnitude.
Analyst
Yep.
Pete Bannon
So one of the things that we additional thing we do to keep the machine going is to have redundant power supplies in the car. So one, one machine's running on one power supply and the other one's on the other. The cameras are the same. So half of the cameras run on the blue power supply, the other half run on the green power supply, and both chips receive all of the video and process it independently. So in terms of driving the car, the basic sequence is collect lots of information from the world around you.
Pete Bannon
Not only do we have cameras, we also have radar, GPS maps, the imus, ultrasonic sensors around the car. We have wheel ticks, steering angle. We know what the acceleration and deceleration of the car is supposed to be. All of that gets integrated together to form a plan. Once we have a plan, the two machines exchange their independent version of the plan to make sure it's the same. And assuming that we agree, we then act and drive the car.
Pete Bannon
Now, once you've driven the car with some new control, you want to validate it. So we validate that what we transmitted was what we intend to transmit to the other actuators in the car. And then you can use the sensor suite to make sure that it happens. So if you ask the car to accelerate or brake or steer right or left, you can look at the accelerometers and make sure that you are in fact doing that. So there's a tremendous amount of redundancy and overlap in both our data acquisition and our data monitoring capabilities here.
Pete Bannon
Moving on to talk about the full self driving chip a little bit. It's packaged in a 37.5 millimeter BGA with 1600 balls. Most of those are used for power and ground, but plenty for signal as well. If you take the lid off, it looks like this. You can see the package substrate and you can see the die sitting in the center there. If you take the die off and flip it over, it looks like this. There's 13,000 C4 bumps scattered across the top of the die and then Underneath that are 12 metal layers which is obscuring all the details of the design.
Pete Bannon
So if you strip that off, it looks like this.
Pete Bannon
This is a 14 nanometer FinFET solution DMOS process. It's 260 millimeters in size, which is a modest sized die. So for comparison, a typical cell phone chip is about 100 millimeters square, which so we're quite a bit bigger than that. But a high end GPU would be more like 600 to 800 millimeters square. So we're sort of in the middle. I would call it the sweet spot. It's a comfortable size to build. There's 250 million logic gates on there and a total of 6 billion transistors, which even even though I work on this all the time, that's mind boggling to me.
Pete Bannon
The chip is manufactured and tested to AEC Q100 standards, which is a standard automotive criteria. Next, I'd like to just walk around the chip and explain all the different pieces to it. And I'm sort of going to go in the order that a pixel coming in from the camera would visit all the different pieces. So up there in the top left you can see, see the camera serial interface. We can ingest 2.5 billion pixels per second, which is more than enough to cover all the sensors that we know about.
Pete Bannon
We have an on chip network that distributes data from the memory system. So the pixels would travel across the network to the memory controllers on the right and left edges of the chip. We use industry standard LPDDR4 memory running at 4266 gigabits per second, which gives us a peak bandwidth to 68 gigabytes a second, which is a pretty healthy bandwidth. But again, this is not like ridiculous. So we're sort of trying to stay in the comfortable sweet spot for cost reasons.
Pete Bannon
The image signal processor has a 24 bit internal pipeline that allows us to do Take full advantage of the HDR sensors that we have around the car. It does advanced tone mapping which helps to bring out details and shadows. And then it has advanced noise reduction which just improves, improves the overall quality of the images that we're using in the neural network. The neural network accelerator itself. There's two of them on the chip.
Pete Bannon
They each have 32 megabytes of SRAM to hold temporary results and minimize the amount of data that we have to transmit on and off the chip, which helps reduce power. Each array has a 96 by 96 multiply add array with in place accumulation which allows us to do almost 10,000 multiply ads per cycle. There's dedicated RELU hardware, dedicated pooling hardware, and each of these deliver 306. Excuse me, each one delivers 36 trillion operations per second and they operate at 2 gigahertz.
Pete Bannon
The two of them together on a die deliver 72 trillion operations a second. So we exceeded our goal of 50 teraops by a fair bit.
Pete Bannon
There's also a video encoder. We encode video and use it in a variety of places in the car, including the backup camera display. There's optionally a user feature for dash cam and also for clip logging data to the cloud, which Stuart and Andre will talk about more later. There's a GPU on the chip. It's modest performance. It has support for both 32 and 16 bit floating point. And then we have 12 A72 64 bit C CPUs for general purpose processing.
Pete Bannon
They operate at 2.2 gigahertz. And this represents about 2 1/2 times the performance available in the current solution.
Pete Bannon
There's a safety system that contains two CPUs that operate in lockstep. This system is the final arbiter of whether it's safe to actually drive the actuators in the car. So this is where the two plans come together and we decide whether it's safe or not to move forward. And lastly there's a safety system. And basically the job of the safety system is to ensure that this chip only runs software that's been cryptographically signed by Tesla.
Pete Bannon
If it's not been signed by Tesla, then the chip does not operate.
Pete Bannon
Now, I've told you a lot of different performance numbers and I thought it'd be helpful maybe to put it into perspective a little bit. So throughout this talk, I'm going to talk about a neural network from our narrow camera. It uses 35 billion operations, 35 giga ops, and if we use all 12 CPUs to process that network, we could do one and a half frames per second, which is super slow, not nearly adequate to drive the car.
Pete Bannon
If we use the 600 gigaflop GPU, the same network, we'd get 17 frames per second, which is still not good enough to drive the car. With eight cameras, the neural network accelerators on the channel chip can deliver 2100 frames per second. And you can see from the scaling as we moved along that the amount of computing in the CPU and GPU are basically insignificant to what's available in the neural network accelerator.
Pete Bannon
It really is night and day.
Pete Bannon
So, moving on to talk about the neural network accelerator, we're just going to stop for some water.
Pete Bannon
On the left, there's a cartoon of a neural network just to give you an idea of what's going on. The data comes in at the top and visits each of the boxes. And the data flows along the arrows to the different boxes. The boxes are typically convolutions or deconvolutions with relus. The green boxes are pooling layers. And the important thing about this is that the data produced by one box is then consumed by the next box, and then you don't need it anymore.
Pete Bannon
You can throw it away. So all of that temporary data that gets created and destroyed as you flow through the network, there's no need to store that off chip in dram. So we keep all that data in sram. And I'll explain why that's super important in a few minutes. If you look over on the right side of this, you can see that in this Network, of the 35 billion operations, almost all of them are convolution, which is based on dot products.
Pete Bannon
The rest are deconvolution, also based on dot product, and then relu and pooling, which are relatively simple operations. So if you were designing some hardware, you'd clearly target doing dot products, which are based on multiply, add, and really kill that. But imagine that you sped it up by a factor of 10,000. So 100% all of a sudden turns into 0.1%, 0.01%, and suddenly the relu and pooling operations are going to be quite significant.
Pete Bannon
So our hardware doesn't. Our hardware design includes dedicated resources for processing, relu and pooling as well.
Pete Bannon
Now, this chip is operating in a thermally constrained environment, so we had to be very careful about how we burn that power. We want to maximize the amount of arithmetic we can do. So we picked integer add. It's 9 times less energy than the corresponding floating point Add and we picked 8 bit by 8 bit integer multiply, which is significantly less power than other multiply operations and is probably enough accuracy to get good results.
Pete Bannon
In terms of memory, we chose to use SRAM as much as possible. And you can see there that going off chip to Dram is approximately 100 times more expensive, expensive in terms of energy consumption than using local sram. So clearly we want to use local SRAM as much as possible. In terms of control, this is data that was published in a paper by Mark Horowitz at ISSCC where he sort of critiqued how much power it takes to execute a single instruction on a regular integer cpu.
Pete Bannon
And you can see that the add operation is only 0.15 centimeter percent of the total power. All the rest of the power is control overhead and bookkeeping. So in our design we sought to basically get rid of all that as much as possible, because what we're really interested in is arithmetic. So here's the design that we finished. You can see that it's dominated by the 32 megabytes of SRAM. There's big banks on the left and right and in the center bottom.
Pete Bannon
And then all the computing is done in the upper middle. Every single clock, we read 256 bytes of activation data out of the SRAM array, 128 bytes of weight data out of the SRAM array, and we combine it in a 96 by 96 mulad array which performs 9000 multiply adds per clock at 2 gigahertz. That's a total of 3.63 36.8 teraops.
Pete Bannon
Now, when we're done with the dot product, we unload the engine so that we shift the data out across the dedicated RELU unit, optionally across a pooling unit, and then finally into a write buffer where all the results get aggregated up. And then we write out 128 bytes per cycle back into the SRAM. And this whole thing cycles along all the time, continuously. So we're doing dot products while we're unloading previous results, doing pooling and writing back into the memory.
Pete Bannon
If you add it all up at 2 gigahertz, you need 1 terabyte per second of SRAM bandwidth to support all that work. And so the hardware supplies that. So one terabyte per second of bandwidth per engine. There's two on the chip. Two terabytes per second.
Pete Bannon
The accelerator has a relatively small instruction set. We have a DMA read operation to bring data in from memory. We have a DMA Write operation to push results back out to memory. We have three dot product based instructions, instructions, convolution, deconvolution and inner product. And then two relatively simple scale is one input, one output operation and outwise is two inputs and one output. And then of course, stop when you're done.
Pete Bannon
We had to develop a neural network compiler for this. So we take the neural network that's been trained by our vision team as it would be deployed in the older cars, and we take that and compile it from for use on the new accelerator.
Pete Bannon
The compiler does layer fusion, which allows us to maximize the computing each time we read data out of the SRAM and put it back. It also does some smoothing so that the demands on the memory system aren't too lumpy. And then we also do channel padding to reduce bank conflicts. And we do bank aware SREM allocation. And this is a case where, where we could have put more hardware in the design to handle bank conflicts.
Pete Bannon
But by pushing it into software, we save hardware and power at the cost of some software complexity. We also automatically insert DMAs into the graph so that data arrives just in time for computing without having to stall the machine. And then at the end, we generate all the code, we generate all the weight data, we compress it and we add a CRC checksum for reliability.
Pete Bannon
To run a program, all the neural network descriptions, programs are loaded into SRAM at the start and then they sit there ready to go all the time. So to run a network, you have to program the address of the input buffer, which presumably is a new image that just arrived from a camera. You set the output buffer address, you set the pointer to the network weights and then you set, set, go, and then the machine goes off and will sequence through the entire neural network all by itself, usually running for a million or 2 million cycles.
Pete Bannon
And then when it's done, you get an interrupt and can post process the results. So moving on to results, we had a goal to stay under 100 watts. This is measured data from cars driving around, running the full autopilot stack. And we're dissipating 72 watts, which is a little bit more power than the previous design. But with the dramatic improvement in performance, it's still a pretty good answer. Of that 72 watts, about 15 watts is being consumed running the neural networks.
Pete Bannon
In terms of cost, the silicon cost of this solution is about 80% of what we were paying before. So we are saving money by switching to this solution. And in terms of performance, we took the narrow camera neural network, which I've been talking about that has 35 billion operations in it. We ran it on the old hardware in a loop as quick as possible and we delivered 110 frames per second. We took the same data, the same network compiled it for hardware for the new FSD computer.
Pete Bannon
And using all four accelerators we can get 2,300 frames per second processed. So a factor of 21.
Elon Musk
I think this is perhaps the most significant slide. It's night and day.
Pete Bannon
I've never worked on a project where the performance increase was more than three, so this was pretty fun.
Pete Bannon
If you compare it to say Nvidia's Drive Xavier solution, A single chip delivers 21 teraops. Our full stop performance driving computer with two chips is 144 teraops.
Pete Bannon
So to conclude, I think we've created a design that delivers outstanding performance. 144 teraops for neural network processing. It has outstanding power performance. We managed to jam all of that performance into the thermal budget that we had. It enables a fully redundant computing solution. It has a modest cost and really the important thing is that this FSD computer will enable a new level of safety and autonomy in Tesla's vehicles without impacting their cost or range.
Pete Bannon
Something that I think we're all looking forward to.
Elon Musk
I think why don't we do Q and A after each segment so if people have questions about the hardware, they can ask right now.
Elon Musk
The reason I asked Pete to do just a detailed, far more detailed than perhaps most people would appreciate, dive into the Tesla full self driving computer is because at first it seems improbable. How could it be that Tesla, who has never designed a chip before, would design the best chip in the world? But that is objectively what has occurred. Not, not best by a small margin, best by a huge margin. It's in the cars right now.
Elon Musk
All Teslas being produced right now have this computer. We switched over from the Nvidia solution for SNX about a month ago and we switched over Model 3 about 10 days ago. All cars being produced have all the hardware necessary, compute and otherwise for full self driving.
Elon Musk
I'll say that again. All Tesla cars being produced right now have everything necessary for full self driving. All you need to do is improve the software and later today you will drive the cars with the development version of the improved software and you will see for yourself themselves.
Elon Musk
Questions for Pete?
Analyst
Y.
Analyst
Questions. I saw Trip Chaudhary Global equities research. Very, very impressive in every shape and form. I was wondering, like I. I took some notes. You are using Activation function relu, the rectify linear unit. But if you think about the deep neural network, it has multiple layers and some algorithms may use different activation functions for different hidden layers like softmax or tanh. Do you have flexibility for incorporating different activation functions rather than LU in your platform?
Analyst
Then I have a follow up.
Pete Bannon
Yes, we do. We have implementations of Tanh and Sigmoid for example.
Analyst
Beautiful. One last question. Like in the nanometers you mentioned 14nm. As I was wondering, wouldn't it make sense to come little lower? Maybe 10nm, two years down or maybe 7?
Pete Bannon
At the time we started the design, not all the IP that we wanted to purchase was available in 10 nanometers. So we finished the design in 14.
Elon Musk
It's maybe worth pointing out that we finished this design like maybe one and a half two years ago and began design of the next generation. We're not talking about the next generation today, but we're about halfway through it.
Elon Musk
That will all the things that are obvious for next generation chip we're doing.
Speaker
Yep.
Pete Bannon
Hi.
Analyst
You talked about the software as the piece now. You did a great job. I was blown away. Understood 10% of what you said, but I trust that it's in good hands.
Elon Musk
Thanks.
Analyst
So it feels like you got the hardware pieces done and that was really hard to do and now you have to do the software piece. Now maybe that's outside of your expertise, but how should we think about that software piece?
Elon Musk
Well, couldn't ask for a better introduction to Andre and Stuart. Are there any questions for the chip part before the next part of the presentation is neural nets and software.
Analyst
So maybe on the chip side, the last slide was 144 trillions of operations per second versus was it Nvidia 21?
Pete Bannon
That's right.
Analyst
And maybe can you just contextualize that for a finance person, why that's so significant, that gap? Thank you.
Pete Bannon
Well, I mean it's a factor of seven in performance delta. So that means you can do seven times as many frames. You can run neural networks that are seven times larger and more sophisticated. So it's a very big current currency that you can spend on lots of interesting things to make the car better.
Elon Musk
I think that Xavier power usage is higher than ours. Xavier powers higher than ours, I think. Or comparable.
Pete Bannon
I don't know that I believe it's
Analyst
like
Elon Musk
to best my knowledge, the power requirements would increase at least to the same degree, a factor of seven and costs would also increase by a factor of seven.
Elon Musk
Great. So yeah, power is a real problem because it also Reduces range. So it has the penalty for power is very high and then you have to get rid of that power by the thermal problem becomes really significant because you got to get rid of all that power. So
Investor Relations
thank you very much. I think we have, you know, a
Elon Musk
lot of quite a bit ask the questions. If you guys don't mind the day running a bit long. Just we're going to do the drive demos afterwards. So if you've got, if you, if you, if anybody needs to pop out and do drive demos a little sooner, you're welcome to do that. But I do want to make sure we answer your questions. Yep.
Elon Musk
Pradeep Ramani from ubs, intel and AMD to some extent have started moving towards a chiplet based architecture. I did not notice a chiplet based design here. Do you think that looking forward that
Stuart Bowers
would be something that might be of
Elon Musk
interest to you guys from an architecture standpoint?
Pete Bannon
A chiplet based architecture?
Analyst
Yes.
Pete Bannon
We're not currently considering anything like that. I think that's mostly useful when you need to use different styles of technology. So if you want to integrate silicon germanium or DRAM technology on the same silicon substrate, that gets pretty interesting. But until the die size gets obnoxious, I wouldn't go there.
Elon Musk
To be clear, the strategy here, and this started basically three, a little over three years ago, was design and build a computer that is fully optimized and aiming for full self driving. Then write software that is designed to work specifically on that computer and get the most out of that computer. So you have tailored hardware that is a master of one trade, self driving.
Elon Musk
Nvidia is a great company but they have many customers and so when as they apply their resources, they need to do a generalized solution.
Elon Musk
We care about one thing, self driving. So it was designed to do that incredibly well. The software is also designed to run on that hardware incredibly well. And the combination of the software and the hardware I think is unbeatable.
Investor Relations
Hi, the chip is designed to process video input.
Andrej Karpathy
In case you use, let's say lidar,
Elon Musk
would it be able to process that as well or is that, is it primarily for video? What we're going to explain to you today is that LIDAR is a fool's errand and anyone relying on LIDAR is doomed.
Elon Musk
Doomed. Expensive, expensive sensors that are unnecessary. It's like having a whole bunch of expensive appendices. One appendix is bad. Well, now they want to put a whole bunch of them. That's ridiculous. You'll see.
Pete Bannon
There's somebody up here.
Speaker
Hi.
Analyst
Hi.
Pete Bannon
Oh, there's a Gentleman. Hi.
Analyst
Hi.
Andrej Karpathy
So just two questions just on the power consumption. Is there a way to maybe give us like a rule of thumb on, you know, every watt is, reduces range by certain percent or a certain amount
Stuart Bowers
just so we can get a sense
Andrej Karpathy
of how much of an improvement a
Pete Bannon
model three, the, the target consumption is 250 watts per mile.
Elon Musk
It depends on the nature of the driving as to how many miles that that affects in city. It would have a much better, bigger effect than on highway. So if you're driving for an hour in a city and you had a solution, Hypothetically that was a kilowatt, you'd lose four miles on a Model 3. So if you're only going say 12 miles an hour, then that would be a 20, 25% impact on range in city. It's basically power is the power of the system has a massive impact on city range, which is where we think most of the robo taxi market will be.
Elon Musk
So power is extremely important.
Elon Musk
Tasha.
Pete Bannon
I'm sorry, I didn't hear you.
Analyst
Thank you.
Andrej Karpathy
What's the primary design objective of the next generation chip?
Elon Musk
We don't want to talk too much about the next generation chip, but it's
Pete Bannon
safety.
Elon Musk
It'll be at least let's say three times better than the current system.
Pete Bannon
Go ahead.
Elon Musk
About two hours away.
Analyst
To develop this chip is the chip being you don't manufacture the chip, you contract that out. And how much cost reduction does that save in the overall vehicle cost?
Pete Bannon
The 20% cost reduction I cited was the piece cost per vehicle reduction. That wasn't a development costs, that was just the actual.
Analyst
No, I'm saying. But like if I'm manufacturing these in mass, is this saving money in doing it yourself?
Pete Bannon
Yes, a little bit.
Elon Musk
I mean most chips are made. Most people don't make chips with their own fab. It's pretty unusual.
Analyst
I think you don't see any supply issues with getting the chip mass produced.
Pete Bannon
The cost saving pays for the development. I mean the basic strategy going to Elon was we're going to build this chip, it's going to reduce the cost. And Elon said times a million cars a year.
Analyst
Deal.
Elon Musk
That's correct, yes.
Pete Bannon
Sorry.
Elon Musk
If there are really chip specific questions, we can answer them. Otherwise there will be a Q and A opportunity after Andre talks and after Stuart talks. So there will be two other Q& A opportunities. This is if it's very chip specific,
Pete Bannon
then also I'll be here all afternoon.
Elon Musk
Yeah, and exactly. And Pete will be here at the end as well. So go ahead.
Andrej Karpathy
Oh yeah, thanks.
Stuart Bowers
That Dye photo you had.
Pete Bannon
There's.
Stuart Bowers
The neural processor takes up quite a bit of the dye. I'm curious, is that your own design or is there some external IP there?
Pete Bannon
Yes, that was a custom design by Tesla.
Speaker
Okay.
Stuart Bowers
And then I guess the follow on would be there's probably a fair amount of opportunity to reduce that footprint as you tweak the design.
Pete Bannon
It's actually quite dense. So in terms of reducing it, I don't think so. It'll greatly enhance the functional capabilities in the next generation.
Stuart Bowers
Okay, and then last question. Can you share where you're fabbing this part?
Elon Musk
Where what?
Pete Bannon
Where are we fabbing it?
Elon Musk
Oh, Samsung.
Pete Bannon
Samsung, yes. Austin, Texas.
Elon Musk
Thank you.
Investor Relations
There's one at the back.
Pete Bannon
Grant Tanaka, Tanaka Capital. Just curious how defensible your chip technologies
Analyst
and design is from a, from a IP point of view and hoping that
Pete Bannon
you won't be offering a lot of
Andrej Karpathy
the IP to the outside for free.
Elon Musk
Thanks.
Pete Bannon
We have filed on the order of a dozen patents on this technology.
Pete Bannon
Fundamentally it's linear algebra, which I don't think you can patent. I'm not sure.
Elon Musk
I think if somebody started today and they were really good, they might have some something like what we have right now in three years, but in two years we'll have something three times better.
Analyst
Talking about the intellectual property protection, you have the best intellectual property and some people just steal it for the fun of it. I was wondering if we look at few interactions with Aurora that companies industry believes they stole your intellectual property. I think the key ingredient that you need to protect is the weights that associate to various parameters. Do you think your chip can do something to prevent anybody?
Analyst
Maybe encrypt all the weights so that even you don't know what the weights are at the chip level so that your intellectual property remains inside it and nobody knows about it and nobody can just steal it.
Elon Musk
Ben, I'd like to meet the person that could do that because I would hire them in a heartbeat. Yeah, so that'd be a hard problem.
Elon Musk
Yeah. Do you want to, I mean we do encrypt the.
Elon Musk
It's a hard chip to crack, so if they can crack it, it's very good. If they can then crack it and then also also figure out the software and the neural net system and everything else. They can design it from scratch like that's all.
Pete Bannon
It's our intention to prevent people from stealing all that stuff. And if they do, we hope it at least takes a long time.
Elon Musk
It will definitely take them a long time. Yeah. I mean I just don't think if it was our goal to do that. How would we do it? It would be very difficult. But the thing that's, I think a very powerful sustainable advantage for us is the fleet. Nobody has the fleet. Those weights are constantly being updated and improved. Based on billions of miles driven, Tesla has 100 times more cars with the full self driving hardware than everyone else combined.
Elon Musk
You know, we, we have, by the end of this quarter we'll have 500,000 cars worth of the full 8 camera setup, 12 ultrasonics, some of them will still be on hardware too, but we still have the data gathering ability. And then a year from now we'll have over a million cars with full self driving computer hardware, everything.
Elon Musk
Yeah, so we have founders. It's just a massive data advantage. It's similar to like, you know how like the Google search engine has a massive advantage because people use it and people are programming effectively program Google with their queries and their results.
Analyst
May I just press you on that and please reframe the question because I'm a tech layman, if it's appropriate. But you know, when we talk to Waymo or Nvidia, they do speak with equivalent conviction about their leadership because, because of their competence in simulating miles driven. Can you talk about the advantage of having real world miles versus simulated miles? Because I think they express that, you know, by the time you get a million miles, they can simulate a billion.
Analyst
And no Formula one race car driver, for example, could ever successfully complete a real world track without driving in a simulator. Can you talk about the advantages it sounds like that you perceive to have associated with having data ingestion coming from real world miles versus simulated miles?
Elon Musk
Absolutely. The simulator, we have quite a good simulation too, but it just does not capture the long tail of weird things that happen in the real world. If the simulation fully captured the real world. Well, I mean that would be proof that we're living in a simulation, I think. Yeah, it doesn't, I wish.
Elon Musk
But simulations do not capture the real world. The real world is really weird and messy. You need the cars on the road. We're actually going to get into that in Andre and Stuart's presentation. So. Okay, why don't we move on to Andre?
Investor Relations
Great, thanks. Thank you.
Pete Bannon
Thank you everybody.
Investor Relations
Thank you very much.
Investor Relations
The last question was actually a very good segue because one thing to remember about our FSD computer is that it can run much more complex neural nets for much more precise image recognition. And to talk to you about how we actually get that image data and how we analyze them, we have our Senior Director of AI, Andre Karpathi, who's going to explain all of that to you. Andrej has a PhD from Stanford University where he studied computer science, focusing on vision recognition and deep learning.
Elon Musk
Andre, why don't you just talk? Do your own intro. There's a lot of PhDs from Stanford. That's not important. Yes, okay, we don't care.
Investor Relations
Come on in.
Pete Bannon
Thank you.
Elon Musk
Andre started the computer vision class at Stanford. That's much more significant. That's what matters.
Andrej Karpathy
Just.
Elon Musk
So can you please talk about your background in a way that is not bashful? Just tell me about the stuff you've done and then.
Investor Relations
Sure.
Elon Musk
Yeah.
Andrej Karpathy
So yeah, I think I've been training neural networks basically for what is now a decade. And these neural networks were not actually really used in the industry until maybe five or six years ago. So it's been some time that I've been training these neural networks and that included institutions at stanford, at, at OpenAI, at Google, and really just training a lot of neural networks not just for images, but also for natural language and designing architectures that couple those two modalities.
Andrej Karpathy
For my PhD, so.
Elon Musk
And the computer computer science class.
Andrej Karpathy
Oh yeah. And at Stanford I actually taught the convolutional neural networks class. And so I was the primary instructor for that class. I actually started the course and designed the entire curriculum. So in a beginning it was about 150 students and then it grew to 700 students over the next two or three years. So it's a very popular class. It's one of the largest classes at Stanford right now. So that was also really successful.
Elon Musk
I mean, Andre is like really one of the best computer vision people in the world. Arguably the best.
Andrej Karpathy
Okay, thank you.
Analyst
Yeah. So.
Andrej Karpathy
Hello everyone. So Pete told you all about the chip that we've designed that runs neural networks in the car. My team is responsible for training of these neural networks and that includes all of data collection from the fleet neural network training and then some of the deployment onto that chip.
Andrej Karpathy
So what do the neural networks do exactly in the car? So what we are seeing here is a stream of videos from across the vehicle, across the car. These are eight cameras that send us videos. And then these neural networks are looking at those videos and are processing them and making predictions about what they're seeing. And so some of the things that we're interested in and some of the things you're seeing on this visualization here are lane line markings, other objects, the distances to those objects, what we call drivable space, shown in blue, which is where the car is allowed to go, and a lot of other predictions like traffic lights, traffic signs, and so on.
Andrej Karpathy
Now for my talk, I will talk roughly in three stages. So first I'm going to give you a short primer on neural networks and, and how they work and how they're trained. And I need to do this because I need to explain in the second part why it is such a big deal that we have the fleet and why it's so important and why it's a key enabling factor to really train these neural networks and making them work effectively on the roads.
Andrej Karpathy
And in the third stage, I'll talk about vision and lidar and how we can estimate depth just from vision alone. So the core problem that these networks are solving in the car is that of visual recognition. So for you and I, these are very, this is a very simple problem. You can look at all of these four images and you can see that they contain a cello, a boat, an iguana or scissors. So this is very simple and effortless for us.
Andrej Karpathy
This is not the case for computers. And the reason for that is that these images are, to a computer, really just a massive grid of pixels. And at each pixel you have the brightness value at that point. And so instead of just seeing an image, a computer really gets a million numbers in a, a grid that tell you the brightness values at all the positions.
Elon Musk
A matrix, if you will. It really is the matrix.
Analyst
Yeah.
Andrej Karpathy
And so we have to go from that grid of pixels and brightness values into high level concepts like iguana and so on. And as you might imagine, this iguana has a certain pattern of brightness values. But iguanas actually can take on many appearances. So they can be in many different appearances, different poses and different brightness conditions against different backgrounds. You can have a different crops of that iguana.
Andrej Karpathy
And so we have to be robust across all those conditions, and we have to understand that all those different brightness patterns actually correspond to iguanas. Now, the reason you and I are very good at this is because we have a massive neural network inside our heads that is processing those images. So light hits the retina, travels to the back of your brain, to the visual cortex. And the visual cortex consists of many neurons that are wired together and that are doing all the pattern recognition on top of those images.
Analyst
Images.
Andrej Karpathy
And really over the last, I would say about five years, the state of the art approaches to processing images using computers have also started to use neural networks, but in this case, artificial neural networks. But these artificial neural networks, and this is just a cartoon diagram of it are a very rough mathematical approximation to your visual cortex. We really do have neurons, and they are connected, connected together.
Andrej Karpathy
And here I'm only showing three or four neurons in three or four in four layers. But a typical neural network will have tens to hundreds of millions of neurons, and each neuron will have a thousand connections. So these are really large pieces of almost simulated tissue. And then what we can do is we can take those neural networks and we can show them images. So, for example, I can feed my iguana into this neural network, and the network will make predictions about what it's seeing.
Andrej Karpathy
Now, in the beginning, these neural networks are initialized completely randomly, so the connection strengths between all those different neurons are completely random. And therefore the predictions of that network are also going to be completely random. So it might think that you're actually looking at a boat right now. And it's very unlikely that this is actually an iguana. And during the training, during the training process, really what we're doing is we know that that's actually an iguana.
Andrej Karpathy
We have a label. So what we're doing is we're basically saying we'd like the probability of iguana to be larger for this image and the probability of all the other things to go down. And then there's a mathematical process called backpropagation, stochastic gradient descent that allows us to back propagate that signal through those connections. And update every one of those connections.
Elon Musk
Sorry.
Andrej Karpathy
And update every one of those connections just a little amount. And once the update is complete, the probability of iguana for this image will go up a little bit. So it might become 14%, and the probability of the other things will go down. And of course, we don't just do this for this single image. We actually have entire large data sets that are labeled. So we have lots of images. Typically, you might have millions of images, thousands of labels or something like that, and you are doing forward, backward, passes over and over again.
Andrej Karpathy
So you're showing the computer, here's an image, it has an opinion, and then you're saying this is the correct answer, and it tunes itself a little bit. You repeat this millions of times, and you sometimes you show images, the same image to the computer, you know, hundreds of times as well. So the network training typically will take on the order of a few hours or a few days, depending on how big of a network you're training.
Andrej Karpathy
And that's the process of training a neural network. Now, there's something very unintuitive about the way neural Networks work that I have to really get into, and that is that they really do require a lot of these examples and they really do start from scratch. They know nothing. And it's really hard to wrap your head around this. So as an example, here's a cute dog. And you probably may not know the breed of this dog, but the correct answer is that this is a Japanese spaniel.
Andrej Karpathy
Now, all of us are looking at this and we're seeing Japanese spaniel, and we're like, okay, I got it. I understand kind of what this Japanese spaniel looks like. And if I show you a few more images of other dogs, you can probably pick out other Japanese spaniels here. So in particular, those three look like a Japanese spaniel and the other ones do not. So you can do this very quickly. And you need one example, but computers do not work like this.
Andrej Karpathy
They actually need a ton of data of Japanese spaniels. So this is a grid of Japanese spaniels showing them, you need thousands of examples showing them in different poses, different brightness conditions, different backgrounds, different crops. You really need to teach the computer from all the different angles what this Japanese spaniel looks like. And it really requires all that data to get that to work. Otherwise the computer can't pick up on that pattern automatically.
Andrej Karpathy
So what does all this imply about the setting of self driving? Of course, we don't care about dog breeds too much. Maybe we will at some point, but for now we really care about lane line markings, objects where they are, where we can drive, and so on. So the way we do this is we don't have labels like iguana for images, but we do have images from the fleet like this. And we're interested in, for example, lane line markings.
Andrej Karpathy
So we, a human typically goes into an image and using a mouse, annotates the lane line markings. So here's an example of an annotation that a human could create a label for this image. And it's saying that that's what you should be seeing in this image. These are the lane line markings. And then what we can do is we can go to the fleet and we can ask for more images from the fleet. And if you ask the fleet, if you just do a naive job of this and you just ask for images at random, the fleet might respond with images like this.
Andrej Karpathy
Typically, going forward on some highway, this is what you might just get like a random collection like this. And we would annotate all that data. Now, if you're not careful, and you only annotate a random distribution of this data, your network will kind of pick up on this random distribution on data and work only in that regime. So if you show it slightly different example, for example, here is an image that actually the road is curving and it is a bit of a more residential neighborhood.
Andrej Karpathy
Then if you show the neural network this image, that network might make a prediction that is incorrect. It might say that, okay, well, I've seen lots of times on highways, lanes just go forward. So here's a possible prediction. And of course this is very incorrect, but the neural network really can't be blamed. It does not know that the train on the, the tree on the left, whether or not it matters or not. It does not know if the car on the right matters or not towards the lane line.
Andrej Karpathy
It does not know that the buildings in the background matter or not. It really starts completely from scratch. And you and I know that the truth is that none of those things matter. What actually matters is that there are a few white lane line markings over there in a vanishing point. And the fact that they curl a little bit should pull the prediction. Except there's no mechanism by which we can just tell the neural network, hey, those lane line markings actually matter.
Andrej Karpathy
The only tool in the toolbox that we have is labeled data. So what we do is we need to take images like this when the network fails and we need to label them correctly. So in this case, we will turn the lane to the right and then we need to feed lots of images of this to the neural net. And neural net, over time will accumulate, will basically pick up on this pattern that those things there don't matter, but those lane line markings do, and we learn to predict the correct lane.
Andrej Karpathy
So what's really critical is not just the scale of the data set. We don't just want millions of images. We actually need to do a really good job of covering the possible space of things that the car might encounter on the roads. So we need to teach the computer how to handle scenarios where it's night and wet, you have all these different specular reflections, and as you might imagine, the brightness patterns in these images will look very different.
Andrej Karpathy
We have to teach the computer how to deal with shadows, how to deal with forks in the road, how to deal with large objects that might be taking up most of that image, how to deal with tunnels or how to deal with construction sites. And in all these cases, there's no, again, explicit mechanism to tell the network what to do. We only have massive amounts of data. We want to source all those images and we want to annotate the correct lines and the network will pick up on the patterns of those now large and varied data sets basically make these networks work very well.
Andrej Karpathy
This is not just a finding for us here at Tesla. This is a ubiquitous is finding across the entire industry. So experiments and research from Google, from Facebook, from Baidu, from Alphabet's DeepMind all show similar plots where neural networks really love data and love scale and variety. As you add more data, these neural networks start to work better and get higher accuracies for free. So more data just makes them work better.
Andrej Karpathy
Now a number of companies have, a number of people have kind of pointed out that potentially we could use simulation to actually achieve the scale of the data sets. And we're in charge of a lot of the conditions here. And maybe we can achieve some variety in a simulator now at Tesla. And that was also kind of brought up in the questions just before this. Now at Tesla, this is actually a screenshot of our own simulator.
Andrej Karpathy
We use simulation extensively. We use it to develop and evaluate the software. We've also even used it for training quite successfully. So but really when it comes to training data for neural networks, there really is no substitute for real data. The simulations have a lot of trouble with modeling appearance, physics and the behaviors of all the agents around you. So there are some examples to really drive that point across the real world really throws a lot of crazy stuff at you.
Andrej Karpathy
So in this case, for example, we have very complicated environments with snow, with trees, with wind. We have various visual artifacts that are hard to simulate potentially. We have complicated construction sites, bushes and plastic bags that can go in, that can kind of go around with the wind. Complicated construction sites that might feature lots of people, kids, animals, all mixed in. And simulating how those things interact and flow through this construction zone might actually be completely, completely intractable.
Andrej Karpathy
It's not about the movement of any one pedestrian in there. It's about how they respond to each other and how those cars respond to each other and how they respond to you driving in that setting. And all of those are actually really tricky to simulate. It's almost like you have to solve the self driving problem to just simulate other cars in your simulation. So it's really complicated. So we have dogs, exotic animals, and in some cases it's not even that you can't simulate, it is that you can't even come up with it.
Andrej Karpathy
So for example, I didn't know that you can have truck, truck on truck like that. But in the real world you find this and you find lots of other things that are Very hard to really even come up with. So really the variety that I'm seeing in the data coming from the fleet is just crazy. With respect to what we have in the simulator, we have a really good simulator.
Elon Musk
I mean, I think simulation, you're fundamentally grading your own homework. So if you know that you're going to simulate it, ok, you can definitely solve for it. But as Andre is saying, you don't know what you don't know. The world is very weird and has millions of corner cases.
Elon Musk
And if somebody can produce a self driving simulation that accurately matches reality, that in itself would be a monumental achievement of human capability. They can't. There's no way.
Andrej Karpathy
Yep. So I think the three points that I really tried to drive home until now are to get neural networks to work well, you require these three essentials. You require a large data set, a very data set and a real data set. And if you have those capabilities, you can actually train neural networks and make them work very well. And so why is Tesla in such a unique and interesting position to really get all these three essentials right?
Andrej Karpathy
And the answer to that, of course, course, is the fleet. We can really source data from it and make our neural network systems work extremely well. So let me take you through a concrete example of, for example, making the object detector work better to give you a sense of how we develop these neural networks, how we iterate on them, and how we actually get them to work over time. So object detection is something we care a lot about.
Andrej Karpathy
We'd like to put bounding boxes around, say the cars and the objects here, because we need to track them and we need to understand how they might move around. So again, we might ask human annotators to give us some annotations for these. And humans might go in and might tell you that, okay, those patterns over there are cars and bicycles and so on. And you can train a neural network on this. But if you're not careful, the neural network will make mispredictions in some cases.
Andrej Karpathy
So as an example, if we stumble by a car like this that has a bike on the back of it, then the neural network actually, when I joined, would actually create two detections. It would create a car detection and a bicycle detection. And that's actually kind of correct because I guess both of those objects actually exist. But for the purposes of the controller and the planner downstream, you really don't want to deal with the fact that this bicycle can go with the car.
Andrej Karpathy
The truth is that that bike is attached to that car. So in terms of like Just objects on the road. There's a single object, a single car. And so what you'd like to do now is you'd like to just potentially annotate lots of those images, as this is just a single car. So the process that we go through in of terms internally in the team is that we take this image or a few images that show this pattern, and we have a mechanism, a machine learning mechanism, by which we can ask the fleet to source us examples that look like that.
Andrej Karpathy
And the fleet might respond with images that contains those patterns. So as an example, these six images might come from the fleet. They all contain bikes on backs of cars. And we would go in and we would annotate all those as just a single car. And then the performance of that detector actually improves. And the network internally understands that, hey, when the bike is just attached to the car, that's actually just a single car.
Andrej Karpathy
And it can learn that given enough examples, and that's how we sort of fix that problem. I will mention that I talk quite a bit about sourcing data from the fleet. I just want to make a quick point that we've designed this from the beginning with privacy in mind, and all the data that we use for training is anonymized. Now, the fleet doesn't just respond with bicycles on backs of cars. We look for all the things. We look for lots of things all the time.
Andrej Karpathy
So, for example, we look for boats, and the fleet can respond with boats. We look for construction sites, and the fleet can send us lots of construction sites from across the world. We look for even slightly more rare cases. So, for example, finding debris on the road is pretty important to us. So these are examples of images that have streamed to us from the fleet that show tires, cones, plastic bags and things like that.
Andrej Karpathy
If we can source these at scale, we can annotate them correctly and the neural network can learn how to deal with them in the world. Here's another example. Animals, of course, also a very rare occurrence and event. But we want the neural network to really understand what's going on here, that these are animals and we want to deal with that correctly. So to summarize, the process by which we iterate on neural network predictions looks something like this.
Andrej Karpathy
We start with a seed data set that was potentially sourced at random. We annotate that data set and then we train neural networks on that data set and put that in the car. And then we have mechanisms by which we notice inaccuracies in the car when this detector may be misbehaving. So for Example, if we detect that the neural network might be uncertain, or if we detect that, or if there's a driver intervention on any of those settings, we can create this trigger infrastructure that sends us data of those inaccuracies.
Andrej Karpathy
And so, for example, if we don't perform very well on lane line detection on tunnels, then we can notice that there's a problem in the tunnels. That image would enter our unit test. So we can verify that we've actually fixing the problem over time. But now what you do is to fix this inaccuracy, you need to source many more examples that look like that. So we ask the fleet to please send us many more tunnels. And then we label all those tunnels correctly, we incorporate that into the training set, and we retrain the network, redeploy and iterate the cycle over and over again.
Andrej Karpathy
And so we refer to this iterative process by which we improve these performances predictions as the data engine. So iteratively deploying something potentially in shadow mode, sourcing inaccuracies, and incorporating the training set over and over again. And we do this basically for all the predictions of these neural networks. Now, so far I talked about a lot of explicit labeling. So like I mentioned, we ask people to annotate data.
Andrej Karpathy
This is an expensive process in time. And also with respect to. Yeah, it's just an expensive process. And so these annotations of course, can be very expensive to achieve. So what I want to talk about also is really to utilize the power of the fleet. You don't want to go through this human annotation bottleneck. You want to just stream in data and automate it automatically. And we have multiple mechanisms by which we can do this.
Andrej Karpathy
So as one example of a project that we recently worked on is the detection of cut ins. So you're driving down the highway, someone is on the left or on the right, and they cut in in front of you into your lane. So here's a video showing the autopilot detecting that this car is intruding into our lane. Now, of course, we'd like to detect a cut in as fast as possible. So the way we approach this problem is we don't write explicit code for is the left blinker on?
Andrej Karpathy
Is the right blinker on? Track the keyboard over time and see if it's moving horizontally. We actually use a fleet learning approach. So the way this works is we ask the fleet to please send us data whenever they see a car transition from a right lane to the center lane or from left to center. And then what we do is we rewind time backwards and we automatically can annotate that, hey, that car will turn, will in 1.3 seconds cut in in front of you.
Andrej Karpathy
And then we can use that for training the neural net. And so the neural net will automatically pick up on a lot of these patterns. So for example, the cars are typically yawed, they're moving this way, maybe the blinker is on. All that stuff happens internally inside the neural net just from these examples. So we asked the fleet to automatically send us all this data. We can get half a million or so images and all of these would be annotated for cut ins.
Andrej Karpathy
And then we train the network. And then we took this cut in network and we deployed it to the fleet. But we don't turn it on yet. We run it in shadow mode. And in shadow mode, the network is always making predictions. Hey, I think this vehicle is going to cut in. From the way it looks, this vehicle is going to cut in. And then we look for mispredictions. So as an example, this is a clip that we had from shadow mode of the cut in network.
Andrej Karpathy
And it's kind of hard to see, but the network thought that the vehicle right ahead of us on the right was going to cut in. And you can sort of see that it's slightly flirting with the lane line. It's trying to, it's sort of encroaching a little bit. And the network got excited and it felt that that was going to be cut in. That vehicle will actually end up in our center lane. That turns out to be incorrect because. And the vehicle did not actually do that.
Andrej Karpathy
So what we do now is we just churn the data engine. We source that ran in the shadow mode. It's making predictions, it makes some false positives and there are some false negative detections. So we got overexcited in sometimes and sometimes we missed a cut in when it actually happened. All those create a trigger that streams to us and that gets incorporated now for free. There's no humans harmed in the process of labeling this data incorporated for free into our training set.
Andrej Karpathy
We retrained the network and redeployed the shadow mode. And so we can spin this a few times. And we always look at the false positives and negatives coming from the fleet. And once we're happy with the false positive, false negative ratio, we actually flip the bit and actually let the car control to that network. And so you may have noticed we actually shipped one of our first versions of a cut in detector approximately, I think three months ago.
Andrej Karpathy
So if You've noticed that the car is much better at detecting cut ins. That's fleet learning operating at scale. Yes, it actually works quite nicely. So that's fleet learning. No humans were harmed in the process. It's just a lot of neural network training based on data and a lot of shadow mode. And looking at those results, another essentially,
Elon Musk
like everyone's training the network all the time is what it amounts to. Whether the, whether autopilot is on or off, the network is being trained. Every mile that's driven for the car, that's hardware 2 or above is training the network.
Analyst
Yeah.
Andrej Karpathy
Another interesting way that we use this in the scheme of fleet learning and the other project that I will talk about is a path prediction. So while you are driving the car, what you're actually doing is you are annotating the data because you are steering the wheel, you're telling us how to traverse different environments. So what we're looking at here is some person in the fleet who took a left through an intersection.
Andrej Karpathy
And what we do here is we, we have the full video of all the cameras and we know that the, the path that this person took because of the gps, the initial measurement unit, the wheel angle, the wheel ticks. So we put all that together and we understand the path that this person took through this environment. And then of course, this, this, we can use this for supervision for the network. So we just source a lot of this from the fleet.
Andrej Karpathy
We train a neural network on the, on those trajectories and then the neural network protection predicts paths just from that data. So really what this is referred to typically is called imitation learning. We're taking human trajectories from the real world and we're just trying to imitate how people drive in real worlds. And we can also apply the same data engine crank to all of this and make this work over time.
Andrej Karpathy
So here's an example of path prediction going through a kind of a complicated environment. So what you're seeing here is a video and we are overlaying the prediction, the predictions of the network. So this is a path that the network would follow in green and some.
Elon Musk
Yeah, I mean, the crazy thing is the network is predicting paths it can't even see with incredibly high accuracy. It can't see around the corner. But, but it's saying the probability of that curve is extremely high. So that's the path and it nails it. You will see that in the cars today. But we're going to turn on augmented vision so you can see the lane lines and the path Predictions of the cars overlaid on the video.
Andrej Karpathy
Yeah, there's actually more going on under the hood that you can even tell.
Elon Musk
I mean, it's kind of scary, to be honest.
Andrej Karpathy
And of course there's a lot of details I'm skipping over. You might not want to annotate all the drivers, you might want to just imitate the better drivers. And there's many technical ways that we actually slice and dice that data. But the interesting thing here is that this prediction is actually a 3D prediction that we project back to the image here. So the path here forward is a three dimensional thing that we're just rendering in 2D.
Andrej Karpathy
But we know about the slope of the ground from all this and that's actually extremely valuable for driving. So path prediction actually is live in the fleet today, by the way. So if you're driving cloverleafs, if you're in a cloverleaf on the highway until maybe five months ago or so, your car would not be able to do cloverleaf. Now it can. That's path prediction running live on your cars. We shipped this a while ago and today you are going to get to experience this.
Andrej Karpathy
For traversing intersections, a large component of how we go through intersections in your drives today is all sourced from path prediction from automatic labels.
Andrej Karpathy
So what I talked about so far is really the three key components of how we iterate on the predictions of the network and how we make it work over time. You require large, varied and real data set. We can really achieve that here at Tesla and we do that through the scale of the fleet, the data engine, shipping things in shadow mode, iterating that cycle, and potentially even using fleet learning where no human annotators are harmed in the process, and just using data automatically.
Andrej Karpathy
And we can really do that at scale.
Andrej Karpathy
So in the next section of my talk, I'm going to especially talk about depth perception using vision only. So you might be familiar that there are at least two sensors in the car. One is vision cameras just getting pixels, and the other is LiDAR that a lot of companies also use. And LiDAR gives you these point measurements of distance around you. Now, one thing I'd like to point out, first of all is you all came here, you drove here, many of you, and you used your neural net and vision.
Andrej Karpathy
You were not shooting lasers out of your eyes and you still ended up here.
Elon Musk
We might have.
Andrej Karpathy
Things went well.
Andrej Karpathy
So clearly the human neural net derives distance and all the measurements and the 3D understanding of the world just from vision. It actually uses multiple cues to do so. I'll just briefly go over some of them just to give you a sense of roughly what's going on inside. As an example, we have two eyes pointed out, so you get two independent measurements at every single time step of the world ahead of you. And your brain stitches this information together to arrive at some depth estimation because you can triangulate any points across those two viewpoints.
Andrej Karpathy
A lot of animals instead have eyes that are positioned on the sides so they have very little overlap in their visual fields. So they will typically use structure for motion. And the idea is that they bob their heads and because of the movement, they actually get multiple observations of the world. And you can triangulate again depths. And even with one eye closed and completely motionless, you can still have some sense of depth perception.
Andrej Karpathy
If you did this, I don't think you would notice me coming 2 meters towards you or 100 meters back. And that's because there are a lot of very strong monocular cues that your brain also takes into account. This is an example of a pretty common visual illusion where you have, you know, these two blue bars are identical, but your brain, the way it stitches up this scene is it just expects one of them to be larger than the other because of the vanishing lines of this image.
Andrej Karpathy
So your brain does a lot of this automatically. And neural nets, artificial neural nets can as well. So let me give you three examples of how you can arrive at depth perception from vision alone. A classical approach and two that rely on neural networks. So here's a video going down, I think this is San Francisco of a Tesla. So these are our cameras, our sensing, and we're looking at all. I'm only showing the main camera, but all the cameras are turned on, the eight cameras of the autopilot.
Andrej Karpathy
And if you just have this six second clip, what you can do is you can stitch up this environment into 3D using multi view stereo techniques. So this.
Elon Musk
Oops.
Andrej Karpathy
This is supposed to be a video, is it not? A video? Oh, I know it's. There we go. So this is the 3D reconstruction of those 6 seconds of that car driving through that path. And you can see that this information is purely, it's very well recoverable from, from just videos. And roughly that's through process of triangulation and as I mentioned, multi view stereo. And we've applied similar techniques slightly more sparse and approximate also in the car.
Andrej Karpathy
So it's remarkable all that information is really there in the sensor and just a matter of extracting it. The other project that I want to briefly talk about is as I Mentioned, there's nothing about neural network. Neural networks are very powerful visual recognition engines. And if you want them to predict depth, then you need to, for example, look for labels of depth. And then they can actually do that extremely well.
Andrej Karpathy
So there's nothing limiting networks from predicting this monocular depth except for labeled data. So one example project that we've actually looked at internally is we use the forward facing radar, which is shown in blue, and that radar is looking out and measuring depths of objects. And we use that radar to annotate the what vision is seeing the bounding boxes that come out of the neural networks. So instead of human annotators telling you, okay, this car and this bounding box is roughly 25 meters away, you can annotate that data much better using sensors.
Andrej Karpathy
So you use sensor annotation. So as an example, radar is quite good at that distance. You can annotate that and then you can train a neural network on it. And if you just have enough data of it, this neural network is very good at predicting those patterns. So here's an example of predictions of that. So in circles, I'm showing radar objects and in. And the cuboids that are coming out here are purely from vision. So the cuboids here are just coming out of vision.
Andrej Karpathy
And the depth of those cuboids is learned by a sensor annotation from the radar. So if this is working very well, then you would see that the circles in the top down view would agree with the cuboids. And they do. And that's because neural networks are very competent at predicting depths. They can learn the different sizes of vehicles internally and they know how big those vehicles are. And you can actually derive depth from that quite accurately.
Andrej Karpathy
The last mechanism I will talk about very briefly is slightly more fancy and gets a bit more technical. But it is a mechanism that has recently, there's a few papers basically over the last year or two on this approach. It's called self supervision. So what you do in a lot of these papers is you only feed raw videos into neural networks with no labels whatsoever. And you can still learn, you can still get neural networks to learn depth.
Andrej Karpathy
And it's a little bit technical, so I can't go into the full details, but the idea is that the neural network predicts depth at every single frame of that video. And then there are no explicit targets that the neural network is supposed to regress to with the labels. But instead, the objective for the network is to be consistent over time. So whatever depth you predict should be consistent over the duration of that video.
Andrej Karpathy
And the Only way to be consistent is to be right. And so the neural network automatically predicts the correct depths for all the pixels. And we've reproduced some of these results internally. So this also works quite well.
Andrej Karpathy
So in summary, people drive with vision only, no lasers are involved. This seems to work quite well. The point that I'd like to make is that visual recognition and very powerful visual recognition is absolutely necessary for autonomy. It's not a nice to have like we must have neural networks that actually really understand the environment around you. And LIDAR points are much less information rich environment. So vision really understands the full details.
Andrej Karpathy
Just a few points around are much. There's much less information in those. So as an example on the left here, is that a plastic bag or is that a tire? Lidar might just give you a few points on that, but vision can tell you which one of those two is true and that impacts your control. Is that person who is slightly looking backwards, are they trying to merge in into your lane on the bike or are they just going forward in the construction sites?
Andrej Karpathy
What do those signs say? How should I behave in this world? The entire infrastructure that we have built up for roads is all designed for human visual consumption. So all of the signs, all the traffic lights, everything is designed for vision. And so that's where all that information is. And so you need that ability. Is that person distracted and on their phone, are they going to walk into your lane? Those answers to all these questions are only found in vision and are necessary for level four, level five autonomy.
Andrej Karpathy
And that is the capability that we are developing at Tesla and through this is done through combination of large scale neural network training through data engine and getting that to work over time and using power of the fleet. And so in this sense LIDAR is really a shortcut. It sidesteps the fundamental problems that the important problem of visual recognition that is necessary for autonomy. And so it gives a false sense of progress and is ultimately a crutch.
Andrej Karpathy
It does give like really fast demos.
Andrej Karpathy
So if I was to summarize the
Analyst
entire
Andrej Karpathy
my entire talk in one slide, it would be this, all of autonomy. Because you want level four, level five systems that can handle all the possible situations in, in 99.99% of the cases and chasing some of the last few nights is going to be very tricky and very difficult and is going to require a very powerful visual system. So I'm showing you some images of what you might encounter in any one slice of that nine. So in the beginning you just have very simple cars.
Andrej Karpathy
Going forward then those Cars start to look a little bit funny. Then maybe you have bikes on cars, then maybe you have cars on cars. Then maybe you start to get into really rare events like cars turned over or even cars airborne. We see a lot of things coming from the fleet and we see them at some rate, at like a really good rate compared to all of our competitors. And so the rate of progress at which you can actually address these problems iterate on the software and really feed the neural networks with the right data.
Andrej Karpathy
That rate of progress is really just proportional to how often you encounter these situations in the wild. And we encounter them significantly more frequently than anyone else, which is why we're going to do extremely well. Thank you.
Analyst
Go ahead.
Andrej Karpathy
It's all super impressive. Thank you so much. How much data, how many pictures are you collecting on average from each car
Pete Bannon
per period of time?
Andrej Karpathy
And then it sounds like the new hardware with the dual, dual active. Active computers gives you some really interesting opportunities to run in full simulation one
Pete Bannon
copy of the neural net while you're
Andrej Karpathy
running the other one. Running the other one.
Pete Bannon
Drive the car and compare the results
Andrej Karpathy
to do quality assurance. And then I was also wondering if there are other opportunities to use the computers for training when they're parked in the garage for the 90% of the time that I'm not driving my time.
Investor Relations
Tesla around.
Andrej Karpathy
Thank you very much. Yep. So for the first question, how much data do we get from the fleet? So it's really important to point out it's not just the scale of the data set, it really is the variety of that data set that matters. If you just have lots of images of something going forward on the highway, at some point the neural just gets it. You don't need that data. So we are really strategic in how we pick and choose.
Andrej Karpathy
And the trigger infrastructure that we've built up is quite sophisticated and allows us to get just the data that we need right now. And so it's not a massive amount of data, it's just very well picked up data. For the second question with respect to redundancy, absolutely. You can run basically the copy of the network on both. And that is actually how it's designed to achieve level four, level five system that is redundant.
Andrej Karpathy
So that's absolutely the case. And your last question. I'm sorry, I did not.
Elon Musk
Training the car is an inference optimized computer. We do have a major program at Tesla which we don't have enough time to talk about today, called Dojo. That's a super powerful training computer. The goal of Dojo will be to Be able to take in vast amounts of data and train at a video level and do unsupervised, massive training of vast amounts of video with the dojo program. Dojo computer. But that's for another day.
Analyst
I'm like a test pilot in a way because I drive the 40510 and all these really tricky, really long tail things happen every day. But the one challenge that I'm curious to how you're going to solve is changing lanes. Because whenever I try to get into a lane with traffic, everybody cuts you off. And so human behavior is very irrational. When you're driving in LA and the car just wants to do it safely and you almost have to do it unsafely.
Analyst
So I was wondering how you're going to solve that problem. Yeah.
Andrej Karpathy
So one thing I will point out is I spoke about the data engine as iterating on neural networks, but we do the exact same thing on the level of software and all the hyper parameters that go into the choices of when we actually lane change, how aggressive we are, we're always changing those, potentially running them in shadow mode and seeing how well they work. And so to tune our heuristics around when it's okay to lane change, we would also potentially utilize the data engine and the shadow mode and so on.
Andrej Karpathy
Ultimately, actually designing all the different heuristics for when it's okay to lane change is actually a little bit intractable, I think, in the general case. And so ideally you actually want to use fleet learning to guide those decisions. So when do humans lane change, in what scenarios and when do they feel it's not safe to lane change? And let's just look at a lot of the data and train machine learning classifiers for distinguishing when it is too safe to do so.
Andrej Karpathy
And those machine learning classifiers can write much better code than humans because they have the mass amount of data backing that. So they can really tune all the right thresholds and agree with humans and do something safe.
Elon Musk
I think we'll probably have a mode that goes beyond Mad Max mode to LA traffic mode. Yeah, well, you know, Mad Max sort of have a hard time in LA traffic, I think.
Andrej Karpathy
Yeah. So really it's a trade off. Like you don't want to create unsafe situations, but you want to be assertive. But that little dance of how you make that work as a human is actually very complicated and it's very hard to write in code. But I think we really do. It really does seem like machine learning approach is kind of like the right
Analyst
way to go about it.
Andrej Karpathy
Where we just look at a lot of ways that people do this and try to imitate that.
Elon Musk
We're just being like more conservative right now. And then as we gain higher and higher confidence, we'll allow users to select a more aggressive mode.
Elon Musk
That'll be up to the user.
Elon Musk
But in the more aggressive modes and trying to merge in traffic, there is a slight. No matter how many new, there's a slight chance of like a fender bender, not a serious accident, but you basically will have a choice of, do you want to have a non zero chance of a fender bender on freeway traffic, which is unfortunately the only way to navigate LA traffic.
Analyst
Yes.
Elon Musk
Yeah.
Analyst
Yes.
Elon Musk
I mean, yes. Yes. It always reminds me of like LA Story. This movie is a great movie.
Andrej Karpathy
Yeah, it's very subtle because there's this game of chicken that's going on.
Elon Musk
Yeah.
Elon Musk
We'll offer more aggressive options over time that will be user specified. Yes. Mad Max plus. Exactly.
Andrej Karpathy
Oh, yeah.
Analyst
Hello. Hi. Jed Doersheimer from Canaccord Genuity. Thank you and congratulations on everything that you've developed. When we look at the Alphazero project, it. It was a very defined and limited variable in terms of the parameters on that, which allowed for the learning curve to be so quick.
Analyst
The risk or what you're trying to do here is almost develop consciousness in the cars through the neural network. And so I guess the challenge is how do you not create a circular reference in terms of. Of the pulling from the centralized model of the fleet to that handoff where the car has enough information.
Analyst
Where is that line? I guess in terms of the point of the learning process to handing it off where there's enough information in the car and not having to pull from the fleet.
Elon Musk
Well, the car can operate if it's completely disconnected from the fleet.
Elon Musk
It just, it uploads the training that's, you know, better and better as the fleet gets better and better. So simply, if you disconnected it from the fleet from that point onwards, it would stop getting better, but it would still function fine.
Analyst
In the hardware portion of your share. In the previous version, it talked about a lot of the power benefits of not storing a lot of the images. And so in this portion, you're talking about the learning that's going on by pulling from the fleet. I guess I'm having a hard time reconciling how if there was a situation where I'm driving up the hill, as you showed, and I'm predicting where the road is going to go, that's coming from all of the other fleet variables that led to that.
Analyst
Intelligence, how I'm not. How I'm getting the benefit of the low power using the cameras with the neural network. That's where I'm losing the the two. Maybe it's just me, but I guess that's.
Elon Musk
I mean the compute power in the full self driving computer is incredible.
Elon Musk
And maybe we should mention that if it had never seen that road before, it would still have made those predictions provided it was a road in the United States.
Analyst
March of 9 case here.
Analyst
In the case of LiDAR, the March of 9th isn't there an example I want to just get to your slam on LiDAR because it's pretty clear you don't like LiDAR in this LiDAR flame.
Elon Musk
LiDAR is lame.
Analyst
Isn't there like a case where at some point nine nine nine nine nine down the road where actually LIDAR may be helpful and why not have it as some sort of a redundancy or backups? That's my first question. And the second. So you can still have your focus on computer vision but just have it as a redundant. My second question is if that is true, what happens to the rest of the industry that's building their Autonomy Solutions on LiDAR?
Elon Musk
They're all going to dump LiDAR. That's my prediction, mark my words.
Elon Musk
I should point out that I don't actually super hate lidar as much as may Sound, but at SpaceX, SpaceX Dragon uses LiDAR to navigate to the space station and dock. Not only that, SpaceX developed its own LIDAR from scratch to do that. And I spearheaded that effort personally because in that scenario lidar makes sense. And in cars it's pretty friggin stupid. It's expensive and unnecessary. And as Andre was saying, once you solve vision it's worthless.
Elon Musk
So you have expensive hardware that's worthless on the car. We do have a forward radar which is low cost and is helpful especially for occlusion situations. So if there's like fog or dust or snow, the radar can see through that. If you're going to use active photon generation, don't use visible wavelength because once with passive optical you've taken care of all visible wavelength stuff. You want to use a wavelength that is occlusion penetrating like radar.
Elon Musk
So LiDAR is just active photon generation in the visual spectrum.
Elon Musk
If you're going to do active photon generation, do it outside of the visual spectrum in the radar spectrum. So like at 3.8 millimeters versus 400, 700 nanometers you're going to be a much better occlusion penetration and that's why we have a forward radar and then we also have 12 ultrasonics for near field information in addition to the eight cameras and the forward radar. You only need the radar in the forward direction because that's the only direction going real fast.
Elon Musk
So it's. I mean, we've gone over this multiple times. Like, are we sure we have the right sensor suite? Should we add anything more? No.
Analyst
Hi. So right here. So you had mentioned that you asked the fleet for the information that you're looking for for some of the vision. I have two questions about that. It sounds like the cars are doing some computation to determine what kind of information to send back to you. Is that a correct assumption? Are they doing that in real time or are they doing based on stored information?
Andrej Karpathy
Yep. So they absolutely do computation in real time on the car and we wait to basically specify condition that we're interested in and then those cars do that computation there. If they did not, then we'd have to send all the data and do that offline in our backend. We don't want to do that. So all that competition happens on the car.
Analyst
So it's based on that question. It sounds like you guys are in a really good position to have currently half a million cars in the future, potentially millions of cars that are essentially computers representing free, almost free data centers for you to do computational. Is that a huge future opportunity for Tesla? It's current, current opportunity and that's not really factored in for anything yet. That's incredible.
Analyst
Thank you.
Elon Musk
We have 425,000 cars with hardware two and beyond, which means they've got all eight cameras, the radar and ultrasonics and they've got at least the Nvidia computer, which is enough to essentially figure out what information is important, what is not. Compress the information that is important to the most salient elements and upload it to the network for training. So it's a massive compression of real world data.
Analyst
You have these sort of network of millions of computers which is like massive data centers essentially that are distributed data centers for computational capacity. Do you see it being used for other things besides self driving in the future?
Elon Musk
I suppose it could possibly be used for something besides self driving. We've been super focused on self driving. So, you know, as we get that really nailed, maybe there's going to be some other use for, you know, millions and then tens of millions of computers with hardware three or four driving computer.
Elon Musk
Yeah, maybe there would be.
Elon Musk
It could be. It could be. Maybe there's like some sort of aws angle here. It's possible.
Elon Musk
Hello.
Andrej Karpathy
Hi, Elon. Matt Joyce, Loop Ventures. I own a Model 3 in Minnesota where it snows a lot.
Stuart Bowers
Since camera and radar cannot see road markings through snow, what is your technical
Andrej Karpathy
strategy to solve this challenge? Does it involve high precision GPS at all?
Andrej Karpathy
Yeah.
Analyst
So
Andrej Karpathy
actually, like today, actually, autopilot will do a decent, decent job in snow. Even when landmarkings are covered, even when landlord markings are faded covered, or when there's lots of rain on them, we still seem to drive relatively well. We didn't specifically go after snow yet with our data engine, but I actually think this is, this is completely tractable because in a lot of those images, even when things are snowy, when you ask a human annotator where are the lane lines, they actually could tell you they actually are relatively consistent in creating those lane lines.
Andrej Karpathy
As long as the annotators are consistent on your data, then I have, there's. The neural network will pick up on those patterns and we'll do just fine. So it's really just about is the signal there even for the human annotator? If, if the answer to that is yes, then the neural network can do it just fine.
Elon Musk
Yeah, there's actually, there are a number of important signals, as Andre was saying. So lane lines are one of those things, but one of the most important signals is drive space. So what is drivable space and what is not drivable space? And what actually really matters the most is drivable space more than lane lines. And the prediction of drivable space is extremely good. And I think especially after this upcoming winter will be incredible.
Elon Musk
It's like, it will be like, how could it possibly be that good? That's crazy.
Andrej Karpathy
The other thing to point out is maybe it's not even only about human annotators. As long as you as a human can drive through that environment through fleet learning, we actually know the path you took. And you obviously used vision to guide you through that path. You did not just use the lane line markings, you use the entire geometry of that entire scene. So you see how the world is roughly curling. You see how the cars are positioned around you.
Andrej Karpathy
Neural network will pick up on all those patterns automatically inside it. If you just have enough of the data people traversing those environments.
Elon Musk
Yeah, it's actually extremely important that things not be rigidly tied to gps, because GPS error can vary quite a bit and the actual situation for a road can vary quite a bit. So there could be construction, there could be a detour, and if the car is using GPS as primary this is a real bad situation. It's asking for trouble. It's fine to use GPS for like tips and tricks. So it's like you can drive your home neighborhood better than a neighborhood in some other country or some other part of the country.
Elon Musk
So you know your own neighborhood well and you use kind of like the knowledge of your neighborhood to drive with more confidence, to maybe have counterintuitive shortcuts and that kind of thing. But you.
Elon Musk
The GPS overlay data should only be helpful, but never primary. If it's ever primary, it's a problem.
Pete Bannon
So question back here in the back corner.
Analyst
Corner. I just wanted to follow up partially
Stuart Bowers
on that because several of your competitors
Pete Bannon
in the space over the past few
Stuart Bowers
years have made, you know, have talked
Andrej Karpathy
about how they are augmenting all of
Stuart Bowers
their perception and path planning capabilities that are kind of on the car platform
Andrej Karpathy
with high definition maps of the areas that they are driving.
Pete Bannon
Does that play a role in your system? Do you see it adding any value?
Stuart Bowers
Are there areas where you would like to get, get more data that is not collected from the fleet but is more kind of mapping style data?
Elon Musk
I think the high precision, high precision GPS maps and lanes are a really bad idea. The system becomes extremely brittle. So any change like this might, any change to the system makes it, it can't adapt. So if it locks onto GPS and high precision lane lines and does not allow vision override, in fact, vision should be the thing that does everything and then like lane lines are a guideline, but they're not the main thing.
Elon Musk
We briefly bulked up the tree of high precision lane lines and then realized that was a huge mistake and reversed it out.
Elon Musk
It's not good.
Stuart Bowers
So this is very helpful for understanding
Andrej Karpathy
annotation, where the objects are and how the car drives. But what about the negotiation aspect for
Pete Bannon
parking and roundabouts and other things where
Stuart Bowers
there are other cars on the road that are human driven, where it's more art than science.
Elon Musk
It does pretty good actually. Like with cut ins and stuff. It's doing really well.
Speaker
Yeah.
Andrej Karpathy
So like I mentioned, we're using a lot of machine learning right now in terms of predicting kind of creating an explicit representation of what the world looks like. And then there's an explicit planner and a controller on top of that representation. And there's a lot of heuristics for how to traverse and negotiate and so on. There is a long tail just like a. In what visual environments look like. There's a long tail in just those negotiations and a little game of chicken that you play with other people and so on.
Andrej Karpathy
And so I think we have a lot of confidence that eventually there must be some kind of a fleet learning component to how you actually do that. Because writing all those rules by hand is going to, is going to quickly plateau, I think.
Elon Musk
Yeah, we've dealt with this issue with cut ins and it's like we'll allow gradually more aggressive behavior on the part of the user. They can just dial the setting up and say be more aggressive, be less aggressive. You know, drive easy, chill mode aggressive.
Speaker
Yeah.
Analyst
Incredible progress. Phenomenal. Two questions. First, in terms of platooning, do you think the system is geared because somebody asked about when there is snow on the road, but if you have platooning feature, you can just follow the car in front. Does your system, is your system capable of doing that? Then I have two follow ups.
Andrej Karpathy
So you're asking about platooning. So I think like we could absolutely build those features. But again if you just use, if you just train neural networks, for example on imitating humans, humans already follow the car ahead. And so that neural network actually incorporates those patterns internally. It's just, it figures out that there's a correlation between the way the car ahead of you faces and the path that you are going to take.
Andrej Karpathy
But that's all done internally in the net. So you're just concerned with getting enough data and the tricky data. And the neural network training process actually is quite magical. Does all the other stuff automatically. So you turn all the different problems into just one problem. Just collect your data set and use neural network training.
Elon Musk
Yeah, there's three steps to self driving. You know, there's been feature complete, then there's being future complete to the degree that where we think that the person in the car does not need to pay attention. And then there's at a reliability level where we've also convinced regulators that that is true. So there's kind of like three levels. We expect to be feature complete in self driving this year and we expect to be confident enough from our standpoint to say that we think people do not need to touch the wheel, look out of the wheel window sometime probably around, I don't know, second quarter of next year.
Elon Musk
And then we start to expect to get regulatory approval, at least in some jurisdictions for that towards the end of next year. That's roughly the timeline that I expect things to go on. And probably for trucks, the platooning will be approved by regulators before anything else. And you could have like maybe if you're long haul doing long haul freight, you can have one driver in the front and then have four semis trailing behind in a platooning manner.
Elon Musk
And I think that probably the regulators will be quicker to approve that than other things.
Analyst
Regarding. Of course, you don't have to convince us. LIDAR is a technology, in my opinion, which has an answer. Looking for a question? Probably dead.
Speaker
The.
Analyst
I mean this is very impressive what we saw today and probably demo could show something more. I was wondering what is the maximum dimension of a matrix that you may be having in your training or in your deep learning pipeline?
Andrej Karpathy
Ballpark figure, max intimation of the matrix. So yeah, doing a lot of matrix multiply operations inside the neural network. You're asking about the like there's many different ways to answer that question, but I'm not 100% sure if they're. They're useful. They're useful answers. These neural networks will typically have, like I mentioned, about tens to hundreds of millions of neurons. Each of them on average have about a thousand connections to neurons below.
Andrej Karpathy
So those are the typical scales that are kind of used across the industry and also that we would use as well.
Analyst
Yeah. I've been actually very impressed by the rate of improvement on Autopilot the past year on my Model 3. The two scenarios I wanted your feedback on last week. The first scenario was I was on the right hand most lane of the freeway and there was a highway on ramp. And then my Model 3 actually was able to detect two cars on the side slow down and let the car go in front of me and one car go behind me. And I was like, oh my gosh, this is like insane.
Analyst
Like I didn't think my Model 3 could do that. So that was like super impressive. But the same week another scenario which is I was on the right hand lane again, but my right hand lane was merging with the left lane. It wasn't an on ramp, it's just a normal highway freeway lane. And my Model 3 wasn't able to detect really that situation and I wasn't able to slow down or speed up and I had to intervene kind of. So can you from your perspective, kind of share kind of the background on how a neural net would, how Tesla might adjust for that and you know, like how that could be improved over time?
Andrej Karpathy
Yeah. So like I mentioned, we have a very sophisticated trigger infrastructure. If you have intervened, it's actually potentially likely that we received that clip and that we can actually analyze it and see what happened and tune the system. So it probably enters some statistics over. Okay, at what rate are we correctly merging the traffic? And we look at those numbers and we look at the clips and we see what's wrong and we try to fix those clips and make progress against those benchmarks.
Andrej Karpathy
So yeah, Yeah. So we would potentially go through a phase of categorization and then we look at some of the biggest kind of categories that actually seem to, to semantically be related to the same problem. And then we will look at some of those and then try to develop software against that.
Elon Musk
Okay. We do have one more presentation which is the software. So it's like essentially the autopilot hardware with Stuart, there's the sort of neural net vision with Andre, and then there's the software engineering at scale that's going to be presented by Stuart. So thanks. And we'll have opportunity afterwards to ask questions. So yeah, thanks.
Investor Relations
I just wanted to very briefly say, if you have an early flight and you want to do a test ride with our latest development software, if you could please speak to my colleague Ann or drop her an email and we can take you out for a test ride. And Stuart, over to you.
Stuart Bowers
All right, so that's actually from a clip of a longer than 30 minute uninterrupted drive with no interventions navigate an autopilot on the highway system which is in production today on hundreds of thousands of cars. So I'm Stuart and I'm here to talk about how we build some of these systems at scale. Just like a really short introduction on kind of where I'm coming from, what I do. So I've been in a couple companies or less.
Stuart Bowers
I've been writing Software professional for about 12 years. The thing that excites me most and I'm really passionate about is taking the cutting edge of machine learning and actually connecting that with customers through robustness and scale. So at Facebook, I worked initially inside of our ads infrastructure to build some of the machine learning some really, really smart people. And we actually tried to build it into a single platform that was we could then scale to all the other aspects of the business, from how we rank the news feed to how we deliver search results to how we make every recommendation across the platform.
Stuart Bowers
And that became the Applied Machine Learning Group. That's something I'm incredibly proud of. And a lot of that wasn't just the core algorithms and the really important improvements that happened there. Those that matters a lot of actually the engineering practices to build these systems at scale. And the same thing was true at Snap, where I went, where we were really, really excited to sort of actually help to monetize this product.
Stuart Bowers
But the hardest part, we were using Google at the time. And they were effectively, you know, running us on a fairly small scale. And we wanted to build that same infrastructure. We take understanding of these users, connect that with cutting edge machine learning, build that at massive scale, and handle billions and then trillions of both predictions and auctions every day in a way which is really robust. And so when the opportunity came to come to Tesla, that's something I'm just like incredibly excited to do, which is specifically take the amazing things that are happening both in the hardware side and the computer vision and AI side and actually package that together with all the planning, the controls, the testing, the kernel patching of the operating system, all of our continuous integration, our simulation, and actually build that into a product.
Stuart Bowers
We get onto people's cars in production today. And so I want to talk about the timeline for how we did that with navigate on Autopilot and how we're going to do that as we get navigate on autopilot off the highway and onto city streets.
Stuart Bowers
So we're at 70 million miles already for Navigate on Autopilot is something really, really, really cool. And I think one thing that is worth kind of calling out on this is that we're continuing to accelerate and keep learning from this data. Like Andre talked about, this data engine, as this accelerates up, we actually do make more and more assertive lane changes. We are learning from these cases where people intervene either because they fail to detect a merge correctly or because they wanted the car to be a little more peppy in different environments.
Stuart Bowers
And we just want to keep making that progress. So to start all of this, we begin with trying to understand the world around us. And we talked about the different sensors in the vehicle. But I wanted to dig in a little bit more. Here we have eight cameras, but then we also have additionally 12 ultrasonic sensors, a radar, an inertial measurement unit, GPS. And then one thing we forget about is we also have the pedal and steering actions.
Stuart Bowers
So not only can we look at what's happening around the vehicle vehicle, we can look at how humans chose to interact with that environment. And so I'll talk to this clip right now. This basically is showing what's happening today in the car, and we're continuing to push this forward. So we start with a single neural network. We see the detections around it. We then build all that together with multiple neural networks and multiple detections.
Stuart Bowers
We bring in the other sensors and we convert that into what Elon calls a vector space, an understanding of the world around us. And this is something where, as we continue, continue to get better and better at this, we're moving more and more of this logic into the neural networks themselves. And the obvious end game here is that the neural network looks across all the cars, brings in all the information together, and just ultimately outputs a source of truth for the world around us.
Stuart Bowers
And this is actually not like an artist rendering, in many senses. This is actually the output of one of the debugging tools that we use on the team every day to understand what the world looks like around us. So another thing that I think is really, really exciting to me, I think when I do hear about sensors like lidar, a common question is around just having extra sensor modalities like why not have some redundancy on the vehicle?
Stuart Bowers
And I want to dig in on one thing that's not. Is not always obvious with neural networks themselves. So we have a neural network running on our, say, wide fisheye camera. That neural network is not making one prediction about the world. It's making many separate predictions, some of which actually audit each other. So as a real example, we have the ability to detect a pedestrian. That's a. Something we train very, very carefully on and put a lot of work into.
Stuart Bowers
We also have the ability to detect obstacles in the roadway, and a pedestrian is an obstacle. And it's shown differently to the neural network. It says, oh, there's a thing I can't drive through. And these together combine to give us an increased sense of what we can and can't do in front of the vehicle and how to plan for that. We then do this across multiple cameras because we have overlapping fields of view in many places around the vehicle in front, we have a particularly large number of overlapping fields of view.
Stuart Bowers
Lastly, we can combine that with things like the radar and the ultrasonics to build these extremely precise understandings of what's happening in front of the car. We can use that both to learn future behaviors that are very accurate. We can also build very accurate predictions of how things will continue to happen in front of us. So one example I think is really exciting is we can actually look at bicyclists and people and not just ask, where are you now?
Stuart Bowers
But where are you going? And this is actually the heart of what we're doing for our new next generation automatic emergency braking system, which will not just stop for people in your path, it'll stop for people who are going to be in your path. And that's running in shadow mode right now. We'll go out to the fleet this quarter and I'll talk about shadow mode in A second.
Stuart Bowers
So when you want to start a feature like this for navigate on autopilot on the highway system, you can start by learning from data. And you can just look at how humans do things today. What is their assertiveness profile? How do they change lanes? What causes them to either absorb or change their maneuvers? And you can see things that are not immediately obvious, like, oh yeah, simultaneous merging is rare, but very complicated and very important.
Stuart Bowers
And you can start to build opinions about different scenarios, such as a fast overtaking vehicle. So this is what we do when we initially have some algorithms we want to try out. We can put them on the fleet and we can see what they would have done in a real world scenario, such as this car that's overtaking us very quickly. This is taken from our actual simulation environment showing different, different paths that we have considered taking and how those overlay on the real world behavior of a user.
Stuart Bowers
When you get those algorithms tuned up and you feel good about them specifically, and this is really taking that output of the neural network, putting it in that vector space, and building and tuning these parameters on top of it. Ultimately a thing we can do through more and more machine learning. You go into a controlled deployment, which for us is our early access program. And then you get this out to a couple thousand people who are really excited to give you highly vigilant but useful feedback about how this behaves not in open loop, but in a closed
Andrej Karpathy
loop way in the real world.
Stuart Bowers
And you watch their interventions. And we talked about this like when somebody takes over, we can actually get that clip, try to understand what happens. And one thing we can really do is we can actually play this back again in an open loop way and ask as we build our software, are we getting closer or further from how humans behave in the real world? And one thing which is super cool, with the full self driving computers, we're actually building our own racks and infrastructure so we basically can fit four of our full self driving computers fully racked up, build these into our own cluster, and actually run this very sophisticated data infrastructure to actually understand over time, as we tune and fix these algorithms, are we getting closer and closer to how humans behave?
Stuart Bowers
And ultimately can we exceed their capabilities? And so once we had this, we felt really good about it. We wanted to do our wide rollout, but to start, we actually asked everybody to confirm the car's behavior via stock confirm. And so we started making lots and lots of predictions about how we should be navigating the highway. We asked people to tell us, is this right or is this wrong. And this is again a chance to churn that data engine.
Stuart Bowers
And we did spot some really tricky and interesting long tails of in this case, I think a really fun example like these very interesting cases of simultaneous merging where you start going and then somebody moves either behind or before you not noticing you. And what is the approach appropriate behavior here and what are the tunings of the neural network we need to do to be super precise about the appropriate behaviors?
Stuart Bowers
Here we worked, we tuned these in the background, we made them better, and over the course of time we got 9 million successfully accepted lane changes. And we use these again with our continuous integration infrastructure to actually understand do we think we're ready. And this is one thing where full self driving is also really exciting to me. Since we own the entire software stack straight from the kernel patching all the way to the ISO like the tuning on the image signal processor, we can start to collect even more data that is even more accurate.
Stuart Bowers
And this allows us to do even better and better tuning these faster iteration cycles. And so earlier this month we were kind of thought we're ready to deploy an even more seamless version of Navigate on autopilot on the highway system. And that seamless version does not require a stock confirm. So you can sit there, relax, put your hand on the wheel and just oversee what the car is doing. And in this case, we're actually seeing over 100,000 automated lane changes every single day on the highway system.
Stuart Bowers
And this is something that's just like super cool to us to deploy at scale. And the thing that I'm kind of most excited about from all this is the actual life cycle of this and how we actually able to turn that data engine crank faster and faster and faster with time. And I think one thing that's really, really becoming very clear is the combination of the infrastructure we have built, the tooling we built on top of that, and the combined power of the full self driving computer, I believe we can do this even faster as we move navigate on autopilot from the highway system onto city streets.
Stuart Bowers
And so yeah, with that I'll hand off to Elon.
Elon Musk
Yeah, I mean to the best of my knowledge, all those lane changes have occurred with zero accidents.
Stuart Bowers
That is correct. Yeah, I watch every single accident.
Elon Musk
So. So it's conservative obviously, but it's to have hundreds of thousands going to millions of lane changes and zero accidents is I think a great achievement by the Tesla team.
Stuart Bowers
Thank you.
Elon Musk
Cool.
Elon Musk
So let's see, you know, a few other things that are maybe worth mentioning. The in order to have a self driving car or robo taxi, you really need redundancy throughout the vehicle at the hardware level. So starting in Maybe it was October 2016, all cars made by Tesla have redundant power steering. So we have redundant motors on the power steering. So any one failure of the if the motor fails, the car can still steer all of the power and data lines have redundancy so you can sever any given power line or any data line and the car will keep driving the auxiliary power system even if the main pack, you lose complete power in the main pack.
Elon Musk
The car is capable of steering and braking using the auxiliary power system, so you can completely lose the main pack and the car is safe.
Elon Musk
The whole system from a hardware standpoint has been designed for to be a RoboTaxi since basically October 2016.
Elon Musk
So when we rolled out hardware autopilot version 2, we do not expect to upgrade cars made before that. We think it would actually cost more to make a new car than to upgrade the cars. Just to give you a sense of how hard it is to do this. Unless it's designed in, it's not worth it.
Elon Musk
So we've gone through the future of self driving where it's clear it's hardware, it's vision and then there's a lot of software and the software problem here should not be minimizing. It's a massive software problem that
Analyst
yeah,
Elon Musk
managing vast amounts of data training against the data. How do you control the car based on the vision? It's a very difficult software problem.
Elon Musk
So going after going over just like Tesla master plan, obviously we've made a bunch of forward looking statements as they call it.
Elon Musk
But let's go through some of our other forward looking statements that we've made. Way back when we created the company, we said we'd build the Tesla Roadster. They said it was impossible and that even if we did build it, nobody would buy it.
Elon Musk
This is like universal opinion was that building an electric car was extremely dumb and would fail.
Elon Musk
I agreed with them that probability of failure was high, but that this was important. So we built the Tesla roadster production in 2008 and shipping that car, it's now a collector's item.
Elon Musk
We build a more affordable car with the Model S. We did that again. We were told that's impossible. I was called a fraud and a liar. It was not going to happen. This is all untrue. Okay, famous last words now is we went into production with the Model S in 2012, exceeded all expectations. There is still in 2019 no car that can Compete with the Model S of 2012. It's seven years later, still waiting.
Elon Musk
So we'd build an affordable car, maybe highly affordable. It's affordable. More affordable with the Model 3. We bought the Model 3, we're in production. I said we'd get over 5,000 cars a week for Model 3. At this point, 5,000 cars a week is a walk in the park for us. It's not even hard.
Elon Musk
So we do large scale solar, which we did through the solar city acquisition and that we develop and deploy the solar roof which is going really well. We're now on version three of the solar tile roof and we expect to spill a production of the solar tile roof significantly later this year.
Elon Musk
I have it on my house and it's great.
Elon Musk
And I sort of make the powerwall and the power pack. We made the power wall and power pack. In fact the power pack is now deployed in massive grid scale utility systems around the world, including the largest operating battery projects in the world that above 100 megawatts. And in the next or probably by next, next year, two years at the most, we expect to have a gigawatt scale battery project completed. So all these things, I said we'd do them, we did it.
Elon Musk
Said we'd do it, we did it. We're going to do the robotaxi thing too. Only criticism and it's a fair one and sometimes I'm not on time, but I get it done and the Tesla team gets it done.
Elon Musk
So what we're going to do this year is we're going to reach combined production of 10,000 a week between SX and 3. Feel very confident about that and we feel very confident about being future complete with self driving.
Elon Musk
Next year we'll expand the product line with Model Y and semi and we expect to have the first operating robotaxis next year with no one in them next year.
Elon Musk
It's always difficult to like when things are on an exponential, at an exponential rate of improvement. It's very difficult to kind of wrap one's mind around it because we're used to extrapolating on a linear basis. But when you've got massive amounts of like as the hardware, massive amounts of hardware on the road, the cumulative data is increasing exponentially. The software is getting better at an exponential rate.
Elon Musk
I feel very confident predicting autonomous robotaxis for Tesla next year. Not an older state, not in all jurisdictions because we won't have regulatory approval everywhere. But I'm confident we'll have at least regulatory approval somewhere literally next year.
Elon Musk
So any customer will be able to add or remove their car to the Tesla network. So expect this to operate like a combination of maybe the Uber and Airbnb model. So if you own the car, you can add or subtract it to the Tesla network, and Tesla would take 25 or 30% of the revenue. And then in places where there aren't enough people sharing their cars, we would just have dedicated Tesla vehicles.
Elon Musk
So when you use the car, we'll show you our ride sharing app. So you're able to summon the car from the parking lot, get in, and go for a drive.
Elon Musk
It's really simple. So you just take the same Tesla app that you currently have. We'll update the app and add a summon, summon Tesla, or commit your car to the fleet. So it's either summon your car or summon a Tesla or add or subtract your car to the fleet. You'll be able to do that from your phone.
Elon Musk
So we see potential for smoothing out the demand distribution curve.
Elon Musk
And having a car operates at a much higher utility than a normal car operates. So typically the use of a car is about 10 to 12 hours a week. So most people will drive one and a half to two hours a day, typically 10 to 12 hours a week of total driving. But if you have a car that can operate autonomously, then most likely you could probably. Most likely you'd have that car operate for a third of the week or longer. So there are 168 hours in a week.
Elon Musk
So probably you've got something on the order of 55, 60 hours a week of operation, maybe a bit longer.
Elon Musk
So the fundamental utility of a vehicle increases by a factor of five. So you can look at this from a macroeconomic standpoint and say, just if this was like some. If we were operating some big simulation, if you could upgrade your simulation to increase the utility of cars by a factor of five, that would be a massive increase in the economic efficiency of the simulation. Just gigantic.
Elon Musk
So we'll do model 3s3 and excess taxis. But we made an important change to our leases. So if you lease a Model 3, you don't have the option of buying it at the end of the lease. We want them back. If you buy the car, you can keep it, but if you lease it, you have to give it back.
Elon Musk
And as I said, in any locations where there's not enough supply for sharing, Tesla will just make its own cars and add them to the network in that place.
Elon Musk
So the current cost of Model 3 Robo Taxi is less than $38,000. We expect that number to improve over time and resigning. The cars, the cars currently being built are all designed for a million miles of operation. The drive units designed and tested and validated for a million miles of operation. The current battery pack is about maybe 300 to 500,000 miles. The new battery pack that probably go into production next year is designed explicitly for a million miles of operation.
Elon Musk
The entire vehicle battery pack, it's designed to operate for a million miles with minimal maintenance. So we'll actually be adjusting tire design and really optimizing the car for a hyper efficient robotaxi. And at some point you won't need steering wheels or pedals and we'll just delete those. So as, as, as these things become less and less important, we'll just delete parts. Just they won't be there.
Elon Musk
If you say like probably two years from now, we make a car that has no steering wheels or pedals and if we need to accelerate that time, we can always just delete parts. Easy.
Elon Musk
Yeah, probably say long term, three years Robotaxis with eliminated parts, maybe it ends up being $25,000 or less.
Elon Musk
And we want a super efficient car. So the electricity consumption is very low. So we're currently at four and a half miles per kilowatt hour. But we can, we'll improve that to five and beyond.
Elon Musk
And there's just really no company that has the full stack integration. We've got the vehicle design and manufacturing, but the computer hardware in house. We've got the in house software development and AI and we've got by far the biggest fleet. It's extremely difficult, not impossible perhaps, but extremely difficult to catch up when Tesla has 100 times more miles per day than everyone else combined.
Elon Musk
This is the cost of running a gasoline car or a. The average cost of running a car in the US is taken from AAA. So it's currently about 62 cents a mile.
Elon Musk
13 and a half thousand miles from 15 million vehicles adds up to 2 trillion a year. These are literally just taken from the AAA website.
Elon Musk
Cost of ride sharing is according to Uber and Lyft is $2 to $3amile. The cost to run a Robotaxi we think less than 18 cents a mile.
Elon Musk
And dropping.
Elon Musk
This would be current. This is current cost. Future cost will be lower, You say. What would be the probable gross profit from a single Robotaxi?
Elon Musk
We think probably something on the order of $30,000 per year.
Elon Musk
And we expect that. We're literally designing, we're designing the cars the same way that commercial semi trailer, semi trucks are designed. Commercial semi trucks Are all designed for a million mile life and we're designing the cars for a million mile life as well.
Elon Musk
So in nominal dollars that would be, you know, a little over $300,000. Over the course of 11 years might be higher. I think these consumptions are actually relatively conservative. And this assumes that 50% of the miles driven are. There's nothing are not useful. So this is only at 50% utility.
Elon Musk
By the middle of next year we'll have over a million Tesla cars on the road with full self driving hardware feature complete at a reliability level that we would consider that no one needs to pay attention. Meaning you could go to sleep. From our standpoint, if you fast forward a year, maybe a year, maybe a year and three months.
Elon Musk
But next year for sure, we will have over a million robotaxis on the road.
Elon Musk
The fleet wakes up with an over the air update. That's all it takes.
Elon Musk
You say what is the net present value of a robotaxi? Probably on the order of a couple hundred thousand dollars. So buying a Model 3 is a good deal.
Elon Musk
Any questions?
Elon Musk
Well, I mean in our own fleet, I don't know, I guess long term we have probably on the order of 10 million vehicles.
Elon Musk
I mean our production rates generally. If you look at our compound annual production rate since 2012 which is like the. That's our first full year of model model S production. We went from 23,000 vehicles produced in 2013 to around 250,000 vehicles produced last year. So in the course of five years we increased output by a factor, factor of 10. I would expect that something similar occurs over the next five or six years.
Elon Musk
As for sharing, sharing versus I don't know. The nice thing is that essentially customers are fronting us the money for the car. It's great.
Stuart Bowers
So in terms of the one thing is the Snake Charger, I'm curious about that.
Andrej Karpathy
And also how did you determine the pricing?
Stuart Bowers
Looks like like you're undercutting the average lift or Uber ride by about 50%. So I'm curious if you could talk
Andrej Karpathy
a little bit about the pricing strategy.
Elon Musk
Sure. We expect the to solving. Solving for the Snake Charger is pretty straightforward from a vision prop standpoint. It's like a known situation. Any kind of known situation with Vision is like a charge port. It's trivial.
Elon Musk
So. So yeah, the car was just automatically park and automatically plug in. There would be no one, no human supervision required.
Elon Musk
Yeah. So sorry, what was pricing? We just threw some numbers on there. I mean I think definitely plug in. Whatever pricing you think makes sense. We just Kind of randomly said, okay, maybe a dollar.
Elon Musk
And the thing is like there's like on the order of 2 billion cars and trucks in the world. So Robotaxis will be in extremely high demand for a very long time. And my observation thus far is that the auto industry is very slow to adapt. I mean, like I said, there's still not a car on the road that you can buy today that is as good as the Model s was in 2012.
Elon Musk
So that suggests a pretty slow rate of adaptation for the car industry. And so probably a dollar is conservative for the next 10 years because people sort of think like there's like actually not enough appreciation for the difficulty of manufacturing. Manufacturing is insanely difficult. But a lot of people I talk to think like if you just have the right design, you can instantly make as much of that thing as the world wants.
Elon Musk
This is not true.
Elon Musk
It's extremely hard to design a new manufacturing system for new technology.
Elon Musk
I mean, Audi is having major problems manufacturing E Tron and, and they are extremely good at manufacturing. And if they're having problems, what about others?
Elon Musk
So the, you know, on the order of 2 billion cars and trucks in the world, on the order of about 100 million units per year of production capacity of vehicles, but only of the old design.
Elon Musk
It will take a very long time to convert all of that to full self driving cars. And they really need to be electric because the cost of operation of a gasoline diesel car is much higher than electric car. So any, any, any robo tax that isn't electric will absolutely not be competitive.
Stuart Bowers
Elon, it's Colin Rush from Oppenheimer over here. You know, obviously we appreciate that the customers are fronting some of the cash for this, this fleet built up, but it sounds like a massive balance sheet commitment from the organization over the course of time. Can you talk a little bit about
Andrej Karpathy
what that looks like, what your expectations
Stuart Bowers
are in terms of financing over the
Andrej Karpathy
next, call it three years, three, four
Stuart Bowers
years for building up this fleet and starting to monetize it with your customer base?
Elon Musk
Well, we're aiming to be approximately cash flow neutral during the, the fleet buildup phase. And then I would expect to be extremely cash flow positive once the robo taxis are enabled. But I don't want to talk about financing rounds. It would be difficult to talk about financing rounds in this venue. But I think we'll make the right moves. I think we'll make the moves you think we should make.
Analyst
I have a question. If I'm Uber, why wouldn't I just buy all your cars? You Know, why would I let you put me out of business?
Elon Musk
There's a, there's a clause that we put into our cars. I think it was about three or four years ago. They can only be used in the Tesla Network.
Analyst
So even a private person, like if I go out and buy 10 model threes, I can't, I can run on the network. That's a business now. Right.
Elon Musk
You're only right to use Tesla Network.
Analyst
Right. But if I use the Tesla network, in theory, I could run a car sharing robo taxi business with my 10 model threes.
Elon Musk
Yes, but it's like the App Store.
Elon Musk
You can only add or remove them through the Tesla Network and then Tesla gets revenue share.
Analyst
But similar to Airbnb though, in that I have this home, my car, and now I can just rent them out so I can make an extra income from owning multiple cars and just renting them out. Like I have a Model 3. I aspire to get this roadster here next when you build it. And I'm going to just rent my Model 3 out. Why would I give it back to you? You know,
Elon Musk
I guess you could operate a rental car fleet, but I think this is very unwieldy. Yeah, I don't know. Seems easy. Okay, try it.
Analyst
In order to operate a robotaxi network, it sounds like you have to solve certain problems, like for example, autopilot today, if you oversteer it, it lets you take over. But if it's, you know, if it's a ride sharing product that someone else is getting in the passenger seat, like moving the steering, steering can't let that person take over the car, for example, because they might not even be in the driver's seat. So is the hardware already there for it to be a robo taxi?
Analyst
And it might get into situations such as a cop pulling it over where some human might need to intervene, like using central fleet of operators that remotely sort of interact with humans or I mean, is all of that type of infrastructure already built into each of the cars?
Elon Musk
Does that make sense? I think there will be sort of a phone home thing where if the car gets stuck, it'll just phone home to Tesla and ask for a solution. Things like being pulled over by police offshore. That's easy for us to program in. That's not a problem.
Elon Musk
It will be possible for some, somebody to take over using the steering wheel at least for some period of time. And then probably down the road we'll just cap the steering wheel so there's no steering control.
Elon Musk
We'll just take the steering Wheel off, put a cap on in the long. Give it like a couple years hardware
Analyst
modification to the car in order for it to enable that.
Elon Musk
Or yeah, we literally just unbolt the steering wheel and put a cap on where the steering wheel handle currently is.
Analyst
But, but that, that is a like future car that you would put out. But what about today's cars where the steering wheel is a mechanism to take over autopilot like so if it's in a robotaxi mode, would someone be able to take it over by just simply moving the steering wheel type?
Elon Musk
Yes, I think there'll be a transition period where people will be able to take over and should be able to take over from the robotaxi. And then once regulators are comfortable with us not having a steering wheel, we'll just delete that. And for cars that are on the, that are in the fleet, you know, obviously with the permission of the owner, if it's owned by somebody else, we would just take the steering wheel off and put a cap where the steering wheel currently attaches.
Analyst
So there might be like two phases to robotaxi. One where the service is provided and you come in as the driver, but could potentially take over. And then in the future there might not be a driver option. Is that how you see it as
Elon Musk
well or like in the future? There will in future. The probability of the steering wheel being taken away in the future is 100% people. Consumers will demand it.
Analyst
But, but initially you would call up.
Elon Musk
This is not. This is, I'm going to clear, does not meet professional prescribing a point of view about the world. This is me predicting what consumers will demand. Consumers will demand in the future that people are not allowed to drive these three ton death machines.
Analyst
I totally agree with that. But in order for a Model 3 today to be part of the Robotaxi network, when you call it, you would then get into the driver's seat essentially because just to be on the same.
Elon Musk
Okay, that makes sense.
Analyst
Thank you.
Elon Musk
Exactly. Just a sort of like, you know, there were amphibians, you know, but then pretty much that things just become like land creatures. There'll be a little bit of an amphibian phase.
Analyst
Hi.
Elon Musk
Sorry. I can see what the. Okay.
Speaker
Yes.
Andrej Karpathy
The strategy we've heard from other players in the robo taxi space is to select a certain municipal area to create geo fenced self driving. That way you're using an HD map to have a more confined area with a bit more safety.
Speaker
A.
Andrej Karpathy
We didn't hear much today around the importance of HD maps to what Extent is an HD map necessary for you? And second, we also didn't hear much about deploying this into specific municipalities where you're working with the municipality to get
Analyst
the buy in from them and you're
Andrej Karpathy
also getting a more defined area. So what's the importance of HD maps and to what extent are you looking at specific municipalities for rollout?
Elon Musk
I think HTMAPs are a mistake. We actually had HTMAPs for a while. Actually can't can that because you either need HTMAPs, in which case if anything changes about the environment, the car will break down, or you don't need HTML in which case why are you wasting your time doing HD maps? So the HD maps thing, like the two main crutches that should not be used and will in retrospect be obviously false and foolish are LIDAR and HD maps.
Elon Musk
Mark my words.
Analyst
Hello.
Elon Musk
If you need a geofenced area, you don't have real self driving.
Andrej Karpathy
Just it sounds like maybe battery supply could be the only bottleneck left towards this vision. And also could you just clarify how you get the battery packs to last a million miles?
Elon Musk
I think cells will be a constraint. That's a subject for a whole separate. That's a whole separate subject.
Elon Musk
And I think we're actually going to want to push our sort of standard range plus battery more than our long range battery because the energy content in the long range pack is 50% higher kilowatt hours.
Elon Musk
So essentially you can make you know, a third more cars if you, if you just. If they're all sort of standard range plus instead of the long range pack. So one's like around 50 kilowatt hours, the other one's around 75 kilowatt hours. So we're actually probably going to bias our sales intentionally towards the smaller battery pack in order to have a higher volume of what basically you want the obvious thing to do is to maximize the number of autonomous units or the number of maximize the output that will substitute result in the biggest autonomous leak down the road.
Elon Musk
So we're doing a number of things in that regard, but it's just not for today's meeting.
Elon Musk
The million mile life is basically just about getting the cycle life of the pack to you know, you need basically on the order. Like let's say you've got a basic math, if you've got a 250 mile range pack, you know you're going to need 4,000 cycles.
Elon Musk
So very achievable. We already do that with our stationary storage. Some of our stationary storage solutions like power pack, we're ready to deploy power pack with 4,000 cycle life capability.
Speaker
Yeah.
Analyst
Can I ask.
Analyst
Sorry, yeah, I wanted to.
Elon Musk
It's like ventriloquism.
Analyst
No, it's obviously significant. Very constructive margin implications to the extent you can drive attach rates much higher of the full self driving option. I'd just be curious if you can level set kind of where you are in terms of those attach rates and how you expect to educate consumers about the Robotax scenario so that attach rates do materially improve improve over time.
Elon Musk
Sorry, it's a bit hard to hear your question.
Analyst
Yeah, just curious where we are today in terms of full self driving attach rates in terms of the financial implications. I think it's hugely beneficial if those attach rates materially increase because of the higher gross margin dollar that flow through. To the extent people do sign up for full fsd, Just curious how you see that ramping
Andrej Karpathy
or what the attach
Analyst
rates are today versus you know, when do you expect. How do you expect to educate consumers and get them aware that they should be attaching FSD to their vehicle purchases?
Elon Musk
We're going to ramp that up massively after today.
Elon Musk
Yeah, I mean the fundamental, really fundamental message that consumers should be taking today is that it's financially insane to buy anything other than a Tesla.
Elon Musk
It'll be like owning a horse in three years. I mean fine if you want to own a horse, but you should go into it with that expectation.
Elon Musk
If you buy a car that does not have the hardware necessary for full self driving, it's like buying a horse and the only car that has the hardware necessary for full self driving is a Tesla.
Elon Musk
Like people should really think about their purchase any other vehicle. It's basically crazy to buy any other car than Tesla.
Elon Musk
We need to make that convey that argument clearly and we will after today.
Analyst
Perfect. Thanks for bringing the future to present very informational time today. I was wondering like you did not talk much about Tesla pickup and let me give a context for that. I could be wrong but the way I'm looking at Tesla Network it was as an early adopter and something as a test bread. I think Tesla's pickup may be the first phase of putting the vehicles in network because the utility of Tesla pickup would be pretty much people who are either loading a lot of stuff or are in the profession of construction or little here and there odd items like picking up stuff from Home Depot.
Analyst
I would say that, you know, maybe it needs to have a two stage process pickup trucks exclusively for Tesla Network as a starting point. Then people like me can buy them later. But what are your thoughts on that?
Elon Musk
Well, today was really just about autonomy. There's, there's a lot that we could talk about such as cell production, pickup truck and future vehicle vehicles. But today was just focus on autonomy. But I agree it's a major thing. I'm very excited for the Tesla pickup truck unveil later this year. It's going to be great.
Stuart Bowers
Colin Lang and UBS Just so we understand the definitions, when you refer to feature complete self driving, it sounds like you're talking level five, no geofence. Is that what's expected by the end of the year?
Elon Musk
Just so we're all.
Stuart Bowers
And then the regulatory process, I mean
Pete Bannon
have you talked to regulators about this? This seems quite an aggressive timeline from
Stuart Bowers
what other people have put out there. I mean are they, you know, what are the hurdles that are needed and what is the timeline to get approval?
Andrej Karpathy
And do you need things like in
Stuart Bowers
California and are they tracking miles that you know, with an operator behind that? Do you need those things? What is that process going to look like?
Elon Musk
Yeah, I mean we talk to regulators around the world all the time as we introduce, you know, additional features like navigate on autopilot.
Elon Musk
This requires like regulatory approval on a per jurisdiction basis.
Elon Musk
So but I think fundamentally regulators in my experience are convinced by data. So if you have a massive amount of data that shows that autonomy is safe, they listen to it. They may take time to digest the information they're processed by. May take a bit of time, but they have always come to the right conclusion from what I've seen.
Andrej Karpathy
Oh, I have a question over here.
Elon Musk
I've got license and pillar. Okay.
Andrej Karpathy
I just wanted to, just to, you know, some of the work we've done trying to better understand the ride hail market. It looks like it's very concentrated in major dense urban centers. So is the way to think about this that the robo taxis would probably deploy more into that area and the additional full self driving for personally owned vehicles would be in the suburban areas?
Elon Musk
I think like probably, yeah, like Tesla owned robo taxis would be in dense urban areas along with customer vehicles. And then as you get to medium and low density areas, it would tend to be more that people own the car and occasionally lend it out.
Elon Musk
Yeah, there are a lot of edge cases in Manhattan and say downtown San Francisco, but those are, you know, and there are various cities around the world that have challenging open environments.
Elon Musk
But we do not expect this to be a significant issue. And when I say future complete, I mean it will work in downtown San Francisco and downtown Manhattan. This Year.
Andrej Karpathy
Hi, I have a neural net architecture question. Do you use different models for say, path planning and perception or different types of AI and sort of how do you split up that problem across the different pieces of autonomy?
Elon Musk
Well, essentially right now, AI or neural nets are used really for object recognition. And we're still basically just using it as still frames, so identifying objects and still frames and tying it together in a perception path planning layer thereafter. But what's happening is steadily is that the neural net is kind of eating into the software base more and more. And so over time, we expect the neural net to do more and more.
Elon Musk
Now, from a computational cost standpoint, there are some things that are very simple for heuristic and very difficult for a neural net. And so it probably makes sense to maintain some level of heuristics in the system because they're just computationally a thousand times easier than a neural net. Like a neural net is like a cruise missile, and if you're trying to swat a fly, just use a fly swatter, not a cruise missile.
Elon Musk
So, but over time, I would expect that it moves really to just training on against video and then video in car steering and pedals out, or basically video in lateral and longitudinal acceleration out almost entirely.
Elon Musk
That's what we're going to use the dojo system for. There's no system that can currently do that.
Andrej Karpathy
Maybe over here.
Stuart Bowers
Just going back to the sensor suite discussion, Elon. One area I'd like to just talk
Elon Musk
about is a lack of side radars.
Stuart Bowers
And in a situation where you have an intersection with a stop sign, where
Andrej Karpathy
there's maybe a 35, 40 mile per
Stuart Bowers
hour cross traffic, are you comfortable with
Andrej Karpathy
the sensor suite, the side cameras being
Analyst
able to handle that?
Stuart Bowers
Just maybe talk a bit about that?
Elon Musk
Yeah, no problem.
Elon Musk
Essentially, the car is going to do kind of what a human would do. You can think of a human as like basically a camera on a slow gimbal. And it's quite remarkable that people are able to drive the car in the way that they are, because you can't look in all directions at once. The car can literally look in all directions at once with multiple cameras. So humans are able to drive just by sort of looking this way, looking that way.
Elon Musk
They're actually stuck in their driver's seat. They can't really get out of the driver's seat. So it's like kind of one camera on a gimbal and is able to drive. A conscientious driver can drive with very high safety. The, the cameras in the cars have a better vantage point than the person. So they're like up in the B pillar or in front of the rear view mirror. They've really got a great vantage point. So if you're turning onto a road that's got a lot of high speed traffic, you can just do what a person does.
Elon Musk
Just turn a little bit. Don't go fully into the road. Let the camera see what's going on. And if things look good and then the rear cameras don't show any oncoming traffic, off you go. And if it looks sketchy, you can just pull back a little bit. Just like a person. The behavior is like remarkably. It starts to become remarkably lifelike. It's like quite eerie, actually. The car just starts behaving like a person over here.
Elon Musk
Here we go then.
Stuart Bowers
Trouble quiz right here.
Elon Musk
Okay.
Stuart Bowers
Given all the value you're creating in your auto business by wrapping all of this technology around yourselves, I guess I'm curious as to why you would still be taking some of your cell capacity and putting it into powerwall and power pack. Wouldn't it make sense to put every single unit you can make into this part of your business?
Elon Musk
We're already stolen almost all the cell lines that were meant to go to powerwall and power pack and use them for model three. I mean, last year, in order to make our Model 3 production and not be self starved, we had to convert all of the 2170 lines at the gigafactory to car sales.
Elon Musk
So our actual output in total gigawatt hours of stationary storage compared to vehicles is an order of magnitude different. And for stationary storage, we can basically use a whole bunch of miscellaneous cells out there. So we can just gather cells from multiple suppliers all around the world, and you don't have a homologation issue or a safety issue like you have with cars. So that's basically our stationary battery business has been just kind of feeding off scraps for quite a while.
Elon Musk
So.
Elon Musk
But like, really think of like the production as being. There are many, many constraints of a massive production system. It's like the degree to which manufacturing a supply chain is underappreciated is amazing. There are a whole series of constraints. And what is the constraint in one week may not be the constraint in another week.
Elon Musk
It's insanely difficult to make a car, especially one which is rapidly evolving. So.
Analyst
Yeah.
Elon Musk
But I'll just take a few more questions and then I think we'll just break four so you can try out the cars.
Pete Bannon
Hi, Elon, Adam, Jonas.
Stuart Bowers
Questions on safety.
Analyst
What.
Stuart Bowers
What data can you share with us today?
Pete Bannon
How safe this technology is, which would
Andrej Karpathy
obviously be important in a regulatory or insurance discussion.
Elon Musk
Well, we publish the accidents per mile every quarter. And what we see right now is that autopilot is about twice as safe as a normal, you know, normal driver on average. And we expect that to increase quite a bit over time.
Elon Musk
Like I said, in the future it will be. Consumers will want to outlaw, and I'm saying they will succeed, nor am I saying I agree with this position, but in the future, consumers will want to outlaw people driving their own cars because it is unsafe. If you think of like elevators, elevators used to be operated on a big lever, like go up and down the floor and there's like a big relay and you had elevator operators, but then periodically they would get tired or drunk or something and then they'd turn the lever at the wrong time and sever somebody in half.
Elon Musk
So now you do not have elevator operators. And it would be quite alarming if you went into an elevator that had a big lever that could just move between floors arbitrarily.
Analyst
So.
Elon Musk
So there's just buttons and in the long term, again, not a value judgment. I'm not saying I want the world to be this way. I'm saying consumers will most likely demand that people are not allowed to drive cars.
Pete Bannon
And Elon, a follow up, can you
Stuart Bowers
share with us how much Tesla's spending
Andrej Karpathy
on Autopilot or autonomous technology by order of magnitude on an annual basis? Thank you.
Elon Musk
It's basically our entire expense structure.
Elon Musk
Question on the, on the economics of the Tesla network. Just so I understand, it looked like. So you get a Model 3 off lease, $25,000 goes on, the balance sheet would be an asset, and then you. It would cash flow $30,000 a year, roughly. Is that the way to think about. Yeah, something like that, yeah.
Andrej Karpathy
And then just in terms of financing
Analyst
of it, there's a question earlier you
Elon Musk
mentioned you would do it. Is it cash flow neutral to the Robotaxi program or cash flow neutral to
Analyst
Tesla as a whole?
Elon Musk
Sorry, the cash flow neutral in terms of.
Andrej Karpathy
He asked a question about financing the
Elon Musk
robo tax, yet it looks to me
Analyst
like they're self financing.
Elon Musk
But you mentioned they would be basically cash flow neutral. Is that what you're referring to? I'm just saying between now and when the Robotaxis are fully deployed throughout the world, the sensible thing for us is to maximize rates and drive the company to cash flow neutral.
Elon Musk
Once the Robotaxi fleet is active, I would expect to be extremely cash flow positive. And so you were talking about production yeah. To produce them all.
Andrej Karpathy
Okay, thanks.
Elon Musk
Maximize the number of autonomous units made.
Analyst
Thank you.
Elon Musk
Okay, just maybe one. One last question here. Hello.
Analyst
If I.
Andrej Karpathy
If I add my Tesla to the Robotaxi network, who.
Elon Musk
Who is liable for an accident?
Analyst
Is it Tesla or is it me?
Pete Bannon
If the vehicle has an accident and harms.
Elon Musk
Probably Tesla. It's probably Tesla.
Speaker
If.
Elon Musk
Yeah.
Elon Musk
I think the right thing to do is just make sure there are very, very few accidents. All right, thanks, everyone. Please, enjoy the drives.
Stuart Bowers
Thank you.
Andrej Karpathy
Thank you very.
Speaker
Much. Sam.
Speaker
Sa.
Speaker
Sam.
Speaker
It.
Speaker
Sam.
Analyst
Sa.
Speaker
It.
Speaker
It.
Speaker
It's.
Speaker
Sa.
Speaker
Sam.
Speaker
It.
Speaker
Ra.