机器翻译,已尽力保留原意与数字
内容摘要
Tesla 2022年 AI 日:首次展示能够行走的 Optimus 人形机器人原型,以及全自动驾驶和 Dojo 超级计算机的最新进展。
Tesla AI Day 2022: the first walking prototype of the Optimus humanoid robot, plus updates on Full Self-Driving and the Dojo supercomputer.
中文实录Transcript
31 个段落
第 1 段
好了,欢迎大家,给每个人一点时间回到观众席,然后好了,很好,欢迎来到 Tesla 2022年 AI 日。我们有一些非常令人兴奋的东西要向你们展示,我想你们会相当印象深刻。我确实想就我们的 Optimus 机器人设定一些预期,因为正如你们所知,去年只是一个穿着机器人服装的人。但我们现在,我们已经取得了长足的进步,而这点我认为我们,你知道,与之相比,它会非常令人印象深刻。我们将讨论用于全自动驾驶的 AI 所取得的进展,以及它们如何更广泛地应用于现实世界的 AI 问题,比如人形机器人,甚至超越这一点。我认为,我们在 Tesla 这里所做的事情有可能为 AGI 作出有意义的贡献。而且我实际上认为,从治理角度看,Tesla 是做这件事的一个不错实体,因为我们是一家上市公司,只有一个股票类别,而这意味着公众控制着 Tesla,我认为这实际上是一件好事。所以如果我疯了,你们可以解雇我。这很重要。也许我没疯。我不知道。所以,是的,我们将大量讨论 AI 自动驾驶方面的进展,以及 Dojo 方面的进展,然后我们会请团队上台,进行一次长时间的问答,这样你们就可以提出尖锐的问题。但无论你们想问什么,存在主义问题、技术问题,不过我们希望尽可能多地留出时间进行问答。所以让我们看看,就这样。那是因为,嘿,大家好,我是米拉娜,从事自动驾驶工作,而它是关于,而且我是莉齐,也是这个项目的机械工程师。好。那么在这之前,我们是不是应该把机器人请上来?今天我们有1个小小的额外提示。
第 2 段
实际上,这是我们第一次在没有任何后备支撑、起重机、机械装置的情况下尝试这个机器人。没有线缆。什么都没有。是的,我想今晚和你们一起做这件事。那是第一次。让我们看看。
第 3 段
准备好了吗?开始吧。我我我觉得这个虫子长了些胸部。顺便说一句,这本质上就是在你们的 Tesla 汽车中运行的那台简单的自动驾驶计算机。这是,这是机器人第一次在没有系绳的情况下运行,是今晚在舞台上。所以这个机器人实际上能做的比我们刚刚展示给你们的多得多,我们只是不想让它脸朝下摔倒。所以我们,我们现在会向你们展示一些视频,里面是机器人做其他许多事情。是的,那些风险更低。是的,我们应该关掉屏幕,各位。是的,是的,我们想多展示一点过去几个月里我们用这个机器人做了什么,还有就是在舞台上四处行走和跳舞。只是卑微的起步,但你们可以看到自动驾驶神经网络正在运行,因为它直接在那个,在那个新平台上为机器人进行了重新训练。那是我的浇水壶,是的,当你,当你看到渲染视图时。那就是机器人。什么是,那就是机器人所看到的世界。所以它,它非常清楚地识别出了物体,比如这就是那个物体。
第 4 段
它应该拿起来,正在把它拿起来。是的。我们使用了与自动驾驶相同的流程来连接数据并训练神经网络,而我们没有把它部署到机器人上。这个例子更进一步展示了上半身。这是我们会想要,比如,在几个月内,在接下来的几个月里,我会说,把它完善到极致的东西。不过这确实也是它正在工作的弗里蒙特工厂里的一个真实工位。而且这并不是我们今天唯一要展示的东西,对吧?是的,当然。所以,你们看到的那个是我们称为 bumble sea 的东西。那算是我们粗糙的开发机器人,使用半现成的执行器。但实际上我们已经比那更进一步了,团队完成了令人难以置信的工作。我们实际上已经有了一个执行器、电池组、控制系统、一切都完全由 Tesla 设计和制造的乐观主义者机器人。嗯,它还没有完全准备好行走,但我想它几周后就会走了。不过我们想向你们展示这个机器人,这个实际上与我们将投入生产的产品相当接近的东西,并向你们展示它能做的所有所有事情,所以让我们把它请上来。好了。是的。所以这里你们看到的是乐观主义者,具有这些,我们预计乐观主义者量产单元1会具备的自由度,也就是能够独立移动所有手指,让拇指拥有2个自由度。所以它拥有可对握的拇指。而且左右手都有,因此它能够操作工具并做有用的事情。
第 5 段
我们的目标是尽快制造出一个有用的人形机器人,而且我们也采用了设计汽车时所用的同一种严谨方法来设计它,也就是说,针对所有制造进行设计,使这种机器人能够以低成本、高可靠性进行大批量生产。所以,那一点极其重要。我的意思是,你们都看过非常令人印象深刻的人形机器人演示,而那,那很好。但它们缺少什么?它们缺少大脑,因为它们不具备自行在世界中行动的智能。而且它们,它们也非常昂贵,嗯,并且产量很低。嗯,而,呃,这个,这就是乐观主义者的设计目标:成为一个能力极强的机器人,但以,以非常大的规模生产,最终可能是数百万台。而且预计它的成本会远低于一辆汽车。所以,呃,我会说,我猜可能不到20,000美元。好了,我认为,很少有人意识到乐观主义者的潜力。和往常一样,Tesla 的演示都是紧赶着完成的。所以,嗯,是的,呃,我是,团队付出了,投入了,而团队已经投入了多得令人难以置信的工作,呃,每天工作,你知道,每周7天,熬着凌晨3点的油,来,来赶上今天的演示。嗯,我对他们所做的工作感到无比自豪,他们确实做得非常出色。我只想请大家为整个乐观主义者团队鼓掌。所以,你知道,现在仍然有很多工作要做,要,呃,完善乐观主义者,并改进它。
第 6 段
显然,这只是乐观主义者第1版,而这其实正是我们举办这场活动的原因,就是要说服像你们这样世界上最有才华的一些人,嗯,加入 Tesla,帮助它成为现实,并让它大规模实现,从而能够帮助数百万人。而且,而且它的潜力真的是令人难以置信,因为你得问,什么,什么是经济?经济是,呃,某种生产实体乘以生产率,呃,人头数乘以人均生产率。当人头数不再受到限制时,到那个时候,甚至都不清楚经济到底意味着什么。经济会变得近乎无限,嗯,所以,什么,什么,你知道,如果它得到实现,在希望是良性的情景下,这意味着一个富足的未来,一个,嗯,没有贫困的未来,人们可以拥有你想要的任何产品和服务。这确实是对我们所知文明的一次,一次根本性变革。显然,我们希望确保这场变革是积极的,而且,嗯,安全的。而,但,但这也是为什么我认为,由 Tesla 这个实体来做这件事非常重要,它只有单一类别的股票、公开交易、由公众持有,嗯,这一点不应被忽视。我认为这至关重要,因为如果公众不喜欢 Tesla 正在做的事,公众可以买入 Tesla 的股票,并投出不同的票。这是一件大事,嗯,比如,非常重要的一点是,我不能只做我想做的事。你知道,有时人们会那么认为,但事实并非如此。嗯,所以,你知道,非常重要的是,让这件事得以发生的那个,那个公司实体,是公众能够适当地施加影响的实体。因此,我认为 Tesla 的结构是,是,是实现这一点的理想结构,嗯。而且就像我说的,你知道,自动驾驶汽车肯定会对世界产生巨大的影响。嗯,我认为它们会将运输的生产率提高至少半个数量级,也许1个数量级,也许更多,嗯。我认为 Optimus 在经济产出方面可能有2个数量级的,呃,潜在提升。比如,比如,这不清楚。甚至根本不清楚实际的极限在哪里,嗯,所以。但我们需要以正确的方式做这件事,我们需要谨慎且安全地做,并确保结果对,呃,文明有益,而且,而且是人类想要的。呃,我不能,这显然也极其重要,所以,嗯。而且,而且我希望你们会考虑,呃,加入 Tesla,实现这些目标。嗯,在 Tesla,我们,我们真的在意在这里做正确的事,或者说立志做正确的事,而且,而且真的不要用良好意图为通往地狱的道路付钱。我认为通往地狱的道路大多是用恶意铺成的,但偶尔,里面也会有一个善意。
第 7 段
所以我们想做正确的事。嗯,所以,你知道,考虑加入我们,帮助它实现,嗯。说到这里,让我们,让我们,呃,我们要进入下一阶段。好了,今天你们已经看到了几个机器人。让我们快速回顾一下时间线。去年,我们发布了 Tesla 机器人概念,但一个概念并不能让我们走得很远。我们知道,我们需要一个真正的开发和集成平台,以尽快获得现实生活中的经验。所以,刚才出来为你们做了那套小动作的机器人,我们在6个月内就把它造了出来,并让它运行起来,此后的几个月里一直在进行软件集成和硬件升级。但与此同时,我们也一直在设计下一代,也就是这边这个。所以这个家伙植根于,某种程度上说,车辆设计流程的基础,你知道。我们正在利用所有那些我们已经拥有的经验。显然,自去年以来发生了很多变化,但仍有几件事保持不变,你们会注意到,我们仍然非常细致地专注于真正的人体形态。我们认为这很重要,原因有几个,但这很有意思。我们花了很多时间思考人体有多么不可思议。我们通常拥有这种惊人的活动范围,以及非常惊人的力量。嗯,一个有意思的练习是,如果你把指尖放在前面的椅子上,你会注意到,比如说,在不移动指尖的情况下,你的肩膀和肘部有巨大的活动范围,你可以让这些关节到处移动。嗯,但机器人,你知道,它的主要功能是完成真正有用的工作。而且它也许不一定马上需要所有那些自由度。所以我们把它精简到了最低限度,某种程度上是28个基本自由度,当然,除此之外还有我们的手。人类在某些事情上也相当高效,而在其他时候则没那么高效。例如,我们可以吃少量食物来维持自己数小时。这很好。呃,但当我们只是有点坐在那里时,无意冒犯,但我们有点低效。我们只是在消耗能量。所以在机器人平台上,我们要做的是把待机功耗降至最低,尽可能降低。这样一来,我们只要拨一下开关,机器人就会立刻变成能够完成有用工作的东西。那么,让我们详细谈谈这最新一代,好吗?
第 8 段
所以在这里的屏幕上,你会看到橙色的是执行器,我们稍后会讲到,蓝色的是电气系统。现在,我们已经有了某种基于人体的研究,也有了第一个开发平台,因此在这个设计中,我们可以同时借鉴研究和执行成果。同样,我们使用的是车辆设计基础。因此,我们要让它从概念经过设计和分析,再到制造和验证。在这个过程中,我们会针对成本和效率等方面进行优化,因为这些是最终将这款产品推向规模化的关键指标。我们要怎么做呢?我们会减少零件数量,并尽可能降低每个元件的功耗。我们会做一些事情,比如减少四肢末端的传感器和布线。你可以想象,手脚中有大量质量,会让它们的移动变得非常困难,也非常耗电。而且,我们会把配电和计算都集中到平台的物理中心。因此,在躯干中部,实际上就是躯干,我们放置了电池包。它的容量为2.3千瓦时,恰好足以支持大约一整天的工作。这个电池包真正独特的地方在于,所有电池电子元件都集成在电池包内的一块PCB上。这意味着,从传感到熔断、充电管理和配电,全部集中在一个地方,全都在一个地方。我们还借鉴了车辆产品和能源产品,把所有这些关键特性融入这块电池。因此,它拥有精简的制造流程、非常高效且简单的冷却方法、电池管理以及安全性。当然,我们可以利用Tesla现有的基础设施和供应链来制造它。接下来谈谈某种意义上的大脑,它不在头部,但离得很近。我们的中央计算机也位于躯干中。大家知道,Tesla生产的每辆车都已经配备完全自动驾驶计算机。我们希望把自动辅助驾驶的硬件和软件都用于人形机器人平台。但由于它在需求和形态因素方面有所不同,我们首先要改变几件事。所以,它仍然会,它将完成人脑所做的一切:处理视觉数据,根据多个传感输入做出split sescan决策,以及进行通信。为了支持通信,它配备了无线连接和音频支持。它还具备硬件级安全功能,这对于保护机器人及其周围的人都很重要。现在,我们已经有了某种核心,接下来需要给这片天空装上一些四肢。嗯,我们也很希望向大家稍微展示一下执行器和功能完整的双手。但在这之前,我想介绍Malcolm,他会稍微讲讲机器人的结构基础。Tesla有能力分析高度复杂的系统。没有什么比碰撞复杂多少。这里可以看到一次来自瓶子3的模拟碰撞,叠加在实际的物理碰撞之上。它的准确程度实际上令人难以置信。为了让大家了解这个模型有多复杂,它包含每一个“不是”、螺栓和垫圈、每一个点焊,并且拥有3500万个自由度,相当惊人。确实可以说,如果没有这样的模型,我们就无法制造出世界上最安全的汽车。那么,我们能否利用汽车方面的能力和方法来影响机器人呢?我们可以制作一个模型,而且由于我们有碰撞软件,这里使用的是同一款软件。我们可以让它摔倒。这样做的目的是确保它摔倒时,理想情况下它不会摔倒,但损伤只是表面的。比如说,我们不希望它弄坏自己的齿轮箱和手臂。
第 9 段
那相当于机器人的肩膀脱臼。修理起来困难且昂贵。所以我们希望它能拍掉自己身上的灰尘,继续完成交给它的工作。我们还可以使用同一个模型,并利用一个先前已求解模型的输入来驱动执行器,让它活起来。因此,这会为我们希望机器人执行的任务生成动作。这些任务包括搬起箱子、转身、下蹲、上楼梯。无论任务集合是什么,我们都可以把它播放给模型。这里展示的只是简单行走。我们可以生成所有部件中的应力,这有助于我们优化这些部件。这些不是在跳舞的机器人,这些实际上是机器人的模态行为,也就是前5个模态。通常人们制造机器人时,会确保第一模态处于较高的个位数范围,接近10赫兹。这样做是为了让行走控制更容易。如果你无法保证自己的脚在哪里,而它还在晃来晃去,走路就会非常困难。如果只制造1个机器人,那没问题。我们想制造数千个,也许数百万个。我们没有用碳纤维、钛来制造的奢侈条件。我们想用塑料制造它们,而这些东西没有那么坚硬。因此,我们无法采用这些高目标。
第 10 段
我称它们为愚蠢的目标。我们必须让它们在较低的目标下工作。那么,这样能很好地工作吗?嗯,如果你仔细想想,抱歉这么说,但我们只不过是塞进了骨头的几袋湿软果冻。我们不是高频的。如果我从腿上开始,我不会以10赫兹振动。我们人类以很大的频率运作。所以我们知道机器人实际上可以做到,只不过这会让控制更加困难。因此,我们从这里获取信息,也就是模态数据和刚度,并将其输入控制系统,让它能够行走。然后,只是稍微改变一下税项,看看膝盖。我们可以从生物学中获得一些启发,看看膝盖的机械优势是什么。事实证明,它实际上表现得与四杆连杆十分相似,而且具有很强的非线性。这其实并不令人意外,因为想想看,当你弯腿蹲下时,膝盖弯曲时承受的扭矩远大于伸直时。因此,你会预期得到一个非线性函数,而事实上,生物学结构就是非线性的。它与之匹配得相当准确。所以这是一种表示形式,四杆连杆显然在物理上并不是四杆连杆,正如我所说,它们的特性相似,但是我弯腰蹲下,这不太科学。让我们更科学一点。我们已经让所有任务通过、通过这张图。这里展示的是设置纠察线、行走、下蹲,也就是我所说的我们在应力方面做过的那些任务。而这就是膝盖处看到的扭矩,横轴是膝盖弯曲程度。这里显示的是膝盖完成所有这些任务的要求。然后画一条曲线,让它在这块区域的顶部冲浪,这表示机器人完成这些任务需要做到这些。
第 11 段
所以,如果我们看一下四杆连杆,它实际上就是绿色曲线。它表示,四杆连杆的非线性实际上使力的特性线性化了,而这真正表达的是,那会降低力。这让执行器具有尽可能低的力,也就是效率最高的状态。我们希望缓慢地消耗能量。蓝色曲线是什么?蓝色曲线实际上表示,如果我们没有四杆连杆,只是在我腿上这里伸出一条杆,杆上装有一个执行器,也就是一个简单的二杆连杆,那么这就是使用简单二杆连杆所能达到的最佳效果。它表明,那会在执行器中产生大得多的力,效率并不高。那么它在实践中是什么样子呢?嗯,正如你将看到的,它非常紧凑地封装在膝盖中,你会看到它在第二个上变透明。你会看到那里的四杆连杆正在作用于执行器。这决定了执行器上的力和位移。现在把时间交给Constantina,由她更详细地向大家介绍这些执行器是如何制造、设计和优化的。谢谢。所以,我是,我想和大家谈谈我们机器人中的设计流程和执行器产品组合。在动力总成设计方面,汽车和机器人之间有很多相似之处。这里最重要的是能量、质量和成本。我们正在把大部分汽车设计经验沿用到机器人上。所以在这个具体案例中,你会看到一辆配有2个驱动单元的汽车。这些驱动单元用于让汽车加速,即0到60英里/小时的时间,或者驾驶一个城市驾驶场地,而机器人有28个执行器,但在执行器层面要完成哪些任务并不明显。
第 12 段
所以,我们有一些更高层级的任务,比如行走、爬楼梯或搬运重物,这些任务需要转化为关节,转化为关节规格,因此我们使用我们的模型,生成关节的扭矩速度轨迹,随后将其输入我们的优化模型,并运行优化流程。这是机器人能够完成的场景之一,也就是转身和行走。所以,当我们得到这条扭矩速度轨迹时,会把它叠加到执行器的效率图上,然后我们就能沿着轨迹生成这项任务随时间变化的功耗和能量、累计能量。所以,这让我们能够确定这个特定执行器的系统成本,并在点云中放入一个简单的点。然后,我们通过在集群中求解,对数十万个执行器执行这项操作。红线表示帕累托前沿,也就是我们寻找最优解时偏好的区域。所以,x 表示我们为这个特定关节选定的首选执行器设计。现在,我们需要对每一个关节都这样做。我们有 28 个关节需要优化,而且我们解析点云,我们针对每项关节规格再次解析点云,这一次红色坐标轴表示为每个关节定制的执行器设计。这里的问题是,我们有太多种独特的执行器设计,即使利用对称性,数量仍然太多。为了制造出可大规模生产的东西,我们需要能够减少独特执行器设计的数量。因此,我们会开展一种称为共通性研究的工作,再次解析点云,这一次寻找能够同时满足不止一个关节的关节性能要求的执行器。所以,最终的产品组合包含 6 种执行器,它们以彩色图的形式显示在中间的图中,嗯,这些执行器也可以在这张幻灯片中看到。我们有 3 种旋转执行器和 3 种线性执行器,它们都具有很高的单位质量输出力或扭矩。尤其是旋转执行器,在高速侧集成了机械离合器、角接触球轴承,在高速侧,以及在低速侧有交叉滚子轴承,而年传动是一种应变波年。嗯,这里有 3 个集成式传感器和定制的永磁电机。线性执行器,抱歉,线性执行器采用行星滚柱,并使用倒置行星螺杆作为齿轮系,从而实现效率、紧凑性和耐用性。所以,为了展示我们的线性执行器的力量能力,我们搭建了一项实验,以便在其极限条件下进行测试。我就让各位欣赏这段视频。我们的执行器能够举起一架重 0.5 吨、长 9 英尺的音乐会三角钢琴。而且,这是一项要求,不是什么锦上添花的东西,因为我们的肌肉在采用直接驱动时也能做到同样的事;在直接驱动时,我们的股四头肌也能做到同样的事,只不过膝盖是一个增速联动系统,它把力转化为脚后跟末端执行器处的速度,目的是赋予人体敏捷性。所以,这是人体令人惊叹的主要方面之一。我的部分到这里就结束了,我想请我的同事迈克上台,他将为各位介绍手部设计。非常感谢。感谢各位来看我们。我们刚刚看到了人类和人形执行器可以有多么强大。不过,人类也极其灵巧。人手能够以每秒 300 度的速度移动,拥有数万个触觉传感器,并且能够抓握和操纵我们日常生活中的几乎所有物体。我们的机器人手部设计受到了生物学的启发。我们有 5 根手指、一根可对握的拇指。我们的手指由既柔韧又坚固的金属肌腱驱动。我们既能够完成大开口的力量抓握,也针对小型、纤薄和易损物体的精密夹持进行了优化。那么,为什么要采用类似人手的机器人手呢?嗯,主要原因是,我们的工厂以及周围的世界都是按人体工学设计的。所以,这意味着它能确保工厂里的物体可以被抓握,同时也能确保我们以前可能从未见过的新物体可以被人手抓握,也可以被我们的机器人手抓握。反过来的情况相当有意思,因为这意味着这些物体是按照我们的手来设计的,而不必为了适应一种新物体去改变我们的手。关于我们的手,一些基本数据是,它有 6 个执行器和 11 个自由度。它有一个手内控制器,用于驱动手指并接收传感器反馈。传感器反馈对于进一步了解我们正在抓握的物体非常重要,对本体感觉也很重要,而本体感觉就是我们识别自己的手处于空间中什么位置的能力。我们的手有一个重要方面,就是它具有适应性。这种适应性本质上涉及一些复杂机构,使手能够适应正在抓握的物体。另一个重要部分是,我们采用了不可反向驱动的手指驱动装置。这种离合机构使我们无须启动手部电机,就能握住并搬运物体。各位刚刚听到了我们如何着手,我们如何着手设计 Tesla 机器人的硬件。现在,我把时间交给米兰和我们的自主系统团队,让他们为这个机器人注入生命。谢谢你,迈克尔。好的,所以,我们之前在视频中展示的所有那些很酷的东西,只用了短短几个月就得以实现,这要归功于过去几年我们在 Autopilot 上完成的出色工作。其中大多数组件都相当容易地移植到了机器人的环境中。仔细想想,我们只是从轮式机器人转向了腿式机器人。所以,有些组件相当相似,另一些则需要投入更多工作。例如,我们的计算机视觉神经网络直接从 Autopilot 移植到了机器人的情境中。这与稍后 Autopilot 团队将更详细介绍的占用网络完全相同,现在它正在这段视频中的机器人上运行。真正改变的唯一一件事,就是我们必须重新采集训练数据。我们还在尝试设法改进这些占用网络,利用在你们的辐射场上完成的工作,对机器人的环境进行非常出色的体积渲染,例如这里是机器人可能需要与之互动的一些机械设备。另一个值得思考的有趣问题是,在室内环境中,主要是在那种 GPS 信号感知的情况下,怎样让机器人导航到目的地?比如说,找到离它最近的充电站。因此,我们一直在训练更多神经网络,以识别机器人摄像头视频流中的高频特征、关键点,并随着机器人在其环境中导航,跨帧追踪它们随时间的变化。我们正利用这些点,更准确地估计机器人行走时在其环境中的姿态和轨迹。我们在仿真方面也做了相当多的工作,而这实际上就是 Autopilot 模拟器。我们已经把机器人的运动代码集成进了它,这是一段运动控制代码在 Autopilot 模拟器中运行的视频,展示了机器人工作随时间的演变。
第 13 段
所以,正如各位所见,我们在 4 月份起步相当缓慢,随着我们在过去几个月中解锁更多关节,以及像手臂平衡这样更深入、更先进的技术,开始加速。因此,运动具体来说是一个非常不同的组件,因为我们正在从汽车环境转向机器人的环境。所以,我认为这值得更深入地讲一讲,现在我想请我的同事们开始介绍这部分。谢谢你,米兰。大家好,我是费利克斯。我是这个项目的一名机器人工程师,接下来我要谈谈行走。行走看起来很容易,对吧?人们每天都在走。你甚至无须去想它。但行走的某些方面从工程到技术来看颇具挑战性。而且我认为,这正是让我更容易思考它的事情之一。但行走的某些方面从工程角度来看颇具挑战性。
第 14 段
例如,身体自我意识,这意味着对自身有良好的表征。你的四肢有多长?你的四肢质量是多少?你的脚有多大?这一切都很重要。还要有一扇节能的门。你可以想象,行走有不同的风格,而且它们的效率全都相同。最重要的是保持平衡。不要摔倒。当然,还要把所有肢体的动作协调起来。现在,人类会自然而然地完成这一切。但作为工程师或机器人专家,我们必须思考这些问题。接下来,我将向各位展示我们如何在运动规划和控制栈中解决这些问题。我们从运动规划开始。还有我们对机器人的表征,这意味着一个包含机器人运动学、动力学和接触属性的模型。利用该模型和为机器人设定的期望路径,我们的运动规划器会为整个系统生成参考轨迹。这意味着相对于我们模型的假设可行的轨迹。规划器目前分 3 个阶段工作。
第 15 段
它从规划落脚点开始,最终形成整个运动照片系统。让我们更深入地了解一下这是如何运作的。在这段视频中,我们看到系统在一个规划时域内,沿着期望路径规划落脚点。我们从这里开始,然后添加连接这些落脚点的足部轨迹,使用脚尖离地和脚跟着地,就像人类,就像人类所做的那样;这为我们提供了最大的右侧和更小的膝部弯曲,从而使系统具有很高的效率。最后一个阶段是找出一条质心轨迹,它能为我们提供整个系统在动力学上可行的运动,以保持平衡。众所周知,计划固然很好,但我们也必须在现实中实现它们。让我们说说如何,看看我们能怎样做到这一点。谢谢你,Felix。大家好,我叫 Anand,我要和大家谈谈控制。那么,让我们把 Felix 刚才谈到的运动计划放到现实世界中的真实机器人上。让我们看看会发生什么。它走了几步,然后摔倒了。嗯,这有点令人失望。但我们在这里缺少几个关键部分,有了它们,机器人就能行走。正如 Felix 提到的,运动规划器使用的是自身的理想化版本,以及周围现实的一个版本。这并不完全正确。它还通过轨迹和扳手来表达自己的意图,也就是它为了对世界进行移动而想施加的力和扭矩的扳手。现实远比任何类似模型复杂得多。此外,机器人也没有被简化。它存在振动和模态、柔顺性、传感器噪声,等等等等。那么,当你把机器人放进现实世界时,这会对现实世界产生什么影响?嗯,意料之外的力会引起未建模的动力学,而行星基本上并不了解这些,这会导致失稳,尤其是对于像双足移动这种动态稳定的系统。那么我们能对此做些什么?
第 16 段
嗯,我们测量现实。我们使用传感器以及我们对世界的理解来进行状态估计。在这里,你可以看到姿态和骨盆位姿,它实质上相当于人类的前庭系统,同时还会跟踪机器人在办公环境中行走时的质心轨迹。现在,我们已经拥有闭环所需的全部组成部分。因此,我们使用更完善的机器人模型。我们使用通过状态估计获得的对现实的理解,并将我们想要的情况与我们预期现实正在对我们做的事情进行比较,以便对机器人的行为进行修正。这里的机器人显然不喜欢被戳,但它在保持直立方面表现得令人钦佩。这里最后一点是,仅仅会走路的机器人还不够。我们需要它使用双手和手臂,才能发挥作用。我们来谈谈操作。大家好,我叫 Eric,是 Tesla 机器人的机器人工程师。我想谈谈我们如何让机器人在现实世界中操作物品。我们希望它在操作物体时看起来尽可能自然,同时也希望迅速实现这一点。所以,我们把这个过程分成了2个步骤。首先是生成一个自然动作参考库,或者可以称之为演示,然后我们在线调整这些动作参考,使其适应当前的现实世界情境。假设我们有一个人类拿起物体的演示。我们可以获取该演示的动作捕捉数据,这里将其可视化为一系列表示双手、手肘和躯干位置的关键帧。我们可以使用逆运动学将其映射到机器人上。如果收集很多这样的数据,我们现在就有了一个可供使用的动作库。但单次演示无法泛化到现实世界中的各种变化。例如,这只适用于位于一个非常特定位置的箱子。因此,我们还通过一个轨迹优化程序来处理这些参考轨迹,该程序会求解手应该在哪里,以及机器人在需要调整动作以适应现实世界时应该如何保持平衡。比如,如果箱子在这个位置,那么我们的优化器就会改为生成这条轨迹。接下来 Milan 将谈谈,呃,Optimus 接下来会怎样,呃,Tesla 谎言。谢谢。好的,希望到现在为止,你们已经很清楚我们过去几个月一直在做什么了。嗯,我们开始拥有了一些可用的东西,但它离真正有用还很远。前方仍有一条漫长而令人兴奋的道路。嗯,我认为未来几周内的第一件事,是让 Optimus 至少与 bumble 分开,看看你们之前看到的另一个机器人原型,而且可能还不止于此。我们也将开始专注于我们其中一家工厂里的真实使用场景,并真正尝试、尝试把这件事彻底搞定,而且我把在现实世界中部署这款产品所需的全部要素都耗尽了。我之前提到过,你知道,室内导航,嗯,优雅的管理,甚至是维修,以及扩大这款产品规模所需的所有组件。但是,嗯,我不知道你们怎么想,但看过我们今晚展示的内容后,我相当确定我们能在未来几个月或几年内完成这件事,嗯,让这款产品成为现实,并改变整个经济。嗯,所以我要感谢整个 Optimus 团队过去几个月的辛勤工作。我认为这相当了不起。所有这些工作仅用了6或8个月就完成了。
第 17 段
非常感谢。大家好。嗨,我是 Ashok。我和 Milan 一起领导自动辅助驾驶团队。天哪,要超越刚才的 optinist 环节实在太难了。不过他无论如何都会试试。过去几年生产的每一辆 Tesla,我们认为都配备了让汽车实现自动驾驶的硬件。我们一直在开发软件,为其增加越来越高级别的自动驾驶能力。去年这个时候,大约有2000辆汽车在运行我们的 FSD Beta 软件。自那以后,我们大幅提升了软件的稳健性和能力,截至今天,我们已经将它推送给160,000名客户。这并非毫无代价,它凝聚了工程团队过去1年的血汗。嗯,例如,仅在过去1年里,我们就训练了75,000个神经网络模型。大约每8分钟就有一个模型。那就是,你知道,由团队产出的,然后我们在大型集群上评估它们,之后我们发布了其中281个模型,它们确实改善了汽车的表现。而且,这种创新贯穿整个技术栈。规划软件、基础设施、工具,甚至招聘,一切都在迈向下一个层级。FSD Beta 软件驾驶汽车的能力相当强。它应该能够从一个停车场行驶到另一个停车场,处理城市街道驾驶,在交通信号灯和停车标志前停车,在十字路口与物体进行协商、转弯等等。所有这些都来自,呃,经过我们神经网络处理的摄像头数据流,而这些神经网络就在汽车本身运行。数据不会传回服务器或任何其他地方。它在车上运行并生成所有输出,呃,用于在车上形成世界模型,而规划软件据此驾驶汽车。今天,我们将深入介绍构成这个系统的许多组件。占用网络充当系统的基础几何层。这是一个多摄像头视频神经网络,它根据图像预测机器人周围世界的完整物理占用情况。因此,任何实际存在的东西,比如树木、墙壁、建筑物、汽车、球,无论什么,只要它实际存在,网络就会预测它们,同时预测它们未来的运动。在这个基础几何层之上,我们还有更多语义层。为了在道路上行驶,我们当然需要车道。但道路有许多不同的车道,它们以各种方式连接。因此,对典型的计算机视觉技术而言,预测车道集合及其连接关系实际上是一个非常困难的问题。所以,我们一路深入到语言技术,然后从其他领域而不只是计算机视觉领域引入最先进的技术,使这项任务成为可能。对于车辆,我们需要它们完整的运动学状态,以便针对它们进行控制。所有这些都直接来自神经网络。视频流,即原始视频流,进入网络,经过大量处理,然后输出完整的运动学状态,也就是位置、速度、加速度、加加速度,所有这些。它们直接从网络输出,只需极少的后处理。这对我来说真的很令人着迷,因为这究竟需要多少?甚至可能吗?我们生活在怎样的世界里,才让这种魔法成为可能,让这些网络能够预测这些位置的四阶导数,而人们曾以为我们甚至无法检测这些物体。我的看法是,这并非毫无代价。它需要海量数据。
第 18 段
所以我们必须是复杂的自动标注系统,透过原始传感器数据显现出来 在服务器上运行大量离线计算。这花了很多时间。这花了很多时间。这花了很多时间 在服务器上运行大量离线计算。这可能要花几个小时,运行昂贵的神经网络 将信息提炼成标签,用来训练我们的车载神经网络 除此之外,我们还使用仿真系统以合成方式创建图像,而且由于这是仿真 我们轻而易举地就拥有所有标签 所有这些都会经过一条运转良好的数据引擎流水线,在那里我们首先用一些数据训练一个基线模型 将它部署到车上,看看有哪些失败,而一旦我们知道这些失败 我们就留意车队中它失败的案例 提供正确的标签,并把数据加入训练集 这个过程会系统性地修复问题,而我们会对车上运行的每项任务都这样做 是的,为了训练这些新的巨型神经网络 今年我们将训练基础设施扩充了大约 40% 到 50% 所以目前我们在美国的多个训练集群中拥有大约 14,000 块 GPU 我们还改进了 AI 编译器,它现在支持这些神经网络所需的新操作 并将它们映射到我们底层最合适的硬件资源上 而我们现在的推理引擎能够将单个神经网络的执行分布到两个独立的片上系统上 本质上就是同一台完全自动驾驶计算机内互联的两台独立计算机 为了实现这一点,我们必须严格控制这个新系统的端到端延迟 所以我们在整个 FSD 平台上部署了更先进的调度代码 所有这些在车内运行的神经网络 共同生成向量空间,也就是机器人或汽车周围世界的模型 然后规划系统在此基础上运行,得出能够避免碰撞或平顺的轨迹 利用基于模型的优化与神经网络的组合朝目的地推进 神经网络帮助对其进行优化,使其真正快速 今天我们非常兴奋地展示所有这些领域的进展 我们的工程负责人已经准备好上台解释这些不同的模块,而这些模块不仅为汽车提供动力 同样的组件也运行在 Milan 先前展示的 Optimus 机器人上 接下来我欢迎 Paril 开始讲解规划部分 大家好,我是 Paril Jain 今天我们就用这个十字路口场景 今天我们就用这个十字路口场景直接深入了解我们如何在 Autopilot 中进行规划和决策 所以我们正从一条支路接近这个十字路口,而且必须让所有横向行驶的车辆先行 对,随着,因为它们即将进入十字路口 十字路口另一侧的行人决定在没有人行横道的地方过马路 现在我们需要让这名行人先行 让右侧驶来的车辆先行,同时还要理解行人与十字路口另一侧车辆之间的关系 所以有很多这类对象内依赖关系 我们需要一眼就迅速解决 而人类非常擅长这一点 我们观察一个场景,理解所有可能的交互,评估最有希望的那些 然后通常最终选择一个合理的方案 所以我们来看一看 Autopilot 系统评估过的几种交互方案 我们本可以用非常激进的纵向横向运动曲线抢在这名行人前面通过 那么显然,我们这样是在对行人耍混蛋,而且会吓到这名行人和他可爱的宠物 我们本可以缓慢向前移动 为行人之间的空隙留短,或者终止从右侧驶来的车辆 同样,我们这样是在对从右侧驶来的车辆耍混蛋 但如果这是唯一可用的安全交互方案,你不应该直接否定它 最后是我们最终选择的交互方式 一开始保持低速,找到合理的空隙,然后在所有参与者通过后完成机动动作 现在,对所有这些交互进行评估并非易事 尤其是当你关心为其他参与者的高阶导数建模时 例如,当你在从右侧驶来的车辆前面强行插入时,那辆车需要多大的纵向加加速度? 单纯依靠边际预测进行碰撞检查,只能让你走到一定程度,因为你会错过许多有效的交互 这基本上归结为针对自车和所有其他参与者的轨迹,求解一个多参与者联合轨迹规划问题 现在,无论你如何优化,这个优化问题的运行速度都会有一个极限 即使经过大量增量近似,它仍会接近接近 10 毫秒的量级 现在,对于一个典型的拥挤无保护升降场景 假设你有超过 20 个对象 每个对象都有多个不同的未来模式,相关交互组合的数量就会急剧膨胀 规划器需要每 50 毫秒做出一次决策。
第 19 段
那么,我们如何实时解决这个问题?我们依赖一个我们称为交互搜索的框架,它基本上是针对一系列机动轨迹开展的瘫痪式研究。这里的状态空间对应于自车的运动学状态、其他主体的运动学状态、它们名义上的未来多种多模态预测,以及场景中的所有静态实体。行动空间则是事情开始变得有趣的地方。我们使用一组候选机动轨迹,在一系列交互决策上进行分支,同时也为更长时域的机动设置增量目标。让我们非常快速地走一遍这项研究,了解它是如何运作的。我们从一组视觉测量结果开始,即车道、占用情况、移动物体。这些会被表示为过去的吸引物以及潜在特征。我们用它来创建一组候选目标。车道同样来自车道网络,或者来自与根据人类示范得出的概率掩码相对应的非结构化区域。一旦获得一批这样的候选目标,我们就结合经典优化方法创建3条轨迹。以及我们的网络规划器,它同样使用来自客户车队的数据进行训练。现在,一旦获得一批这样的3条轨迹,我们就用它们开始针对交互进行分支。我们找出最关键的交互。在我们的例子中,这会是相对于行人的交互。我们是在它前面强行通过,还是让行。显然,左边的选项是一个高惩罚选项,它很可能不会被优先考虑。因此,我们进一步在右边的选项上分支,而我们正是在那里引入越来越复杂的交互。通过越来越多的约束,逐步构建这个优化问题。树搜索继续推进,在更多交互上分支,在更多目标上分支。这里的许多刺儿在于对树搜索中每一个节点的评估。在每个节点内部,我们最初先使用经典优化方法创建轨迹。像我所描述的那些约束会被逐步加入。每个行动需要接近1到5毫秒。现在,尽管这是个相当不错的数字,但当你想评估超过100%的交互时,它就无法扩展。因此,我们最终构建了可以在规划器循环中运行的轻量级可查询网络。这些网络使用来自车队的人类示范,以及放宽时间限制的离线求解器进行训练。借此,我们得以将每个行动的运行时间降至接近100微秒。现在,仅仅这样做还不够,因为你仍然需要遍历这个庞大的树搜索。而且你需要高效地剪枝搜索空间。因此,你需要对这些轨迹中的每一条进行新的评分。其中一些相当标准,你会进行一系列碰撞检查,也会进行一系列舒适度分析。对于给定的肥料,需要怎样的加加速度和访问权限。客户车队数据在这里再次发挥重要作用。我们运行两组同样轻量级的可查询网络,二者实际上相互增强。其中一组使用FSD测试版车队的人工接管数据进行训练。它会对给定的肥料在接下来几秒内导致人工接管的可能性给出评分。第2组则完全基于人工示范,也就是人类驾驶数据,针对你选定的行动与人类驾驶轨迹有多接近给出评分。这种评分帮助我们剪枝搜索空间,继续在交互上进一步分支,并将算力集中于最有希望的结果。这种架构很酷的一点在于,它让我们能够在数据驱动方法之间形成一种很酷的融合,在这种方法中,你不必依赖大量人工设计的代价。但同时也通过基于物理的检查将其立足于现实。我刚才描述的很多内容都是针对我们能够在场景中观察到的主体。但同一个框架也扩展到我们拥有的所有其他系统。我们使用8个摄像头的视频流生成世界的3D占用情况。这里的蓝色掩码对应于我们所说的可见区域。它基本上会在你看到的场景中第1处遮挡位置被阻断。我们使用这个可见性掩码来生成场景的可见性。我们使用8个摄像头的视频流生成世界的3D占用情况。这里的蓝色掩码对应于我们所说的可见区域。在你看到的场景中第1处遮挡位置,我们使用这个可见性掩码生成我们称为幽灵物体的东西,你可以在左上角看到。现在,如果你正确地建模这些幽灵物体的生成区域和状态转移。如果你根据它们存在的可能性来调整控制响应,就可以提取出一些非常不错的、类似人类的行为。现在我把话交给Phil,由他进一步介绍我们如何生成这些占用网络。大家好,我叫Phil,我将分享我们在过去1年里构建的占用网络的细节。这个网络是我们用来对车辆周围的3D物理工作进行建模的解决方案。而且它目前没有显示在面向客户的可视化界面中。你们将在这里看到的是我们内部实验室工具输出的原始网络结果。占用网络以我们全部8个摄像头的视频流作为输入。直接在矢量空间中生成一个统一的体积占用结果。对于车辆周围的每一个3D位置,它都会预测该位置被占用或未被占用的概率。由于它具有视频接触,因此能够预测瞬间被遮挡的障碍物。对于每个位置,它还会生成一组语义,例如路缘、汽车、行人和道路碎片,这里以颜色编码表示。它还会预测用于运动的占用流。由于该模型是一个通用网络,它不会明确区分静态物体和动态物体。它能够生成并建模随机运动,例如这里这个蜂拥而动的训练器。这个网络目前运行在所有配备FSD计算机的Tesla上。而且效率极高,借助我们的神经线路加速器,大约每10毫秒运行一次。那么它是如何工作的?让我们看看架构。首先,我们使用相机标定对每个摄像头图像进行校正。我们这里展示的图像就是提供给网络的图像。它实际上不是典型的8位RGB图像。从顶部的第1张图可以看到,我们提供给网络的是12位原始照片账户图像。因为它多了4位信息,所以动态范围提升了16倍,同时也降低了延迟。因为我们不再需要在循环中运行ISP。我们使用一组reglets和bif-fps作为骨干网络,提取图像空间特征。接下来,我们构建一组3D位置查询,并将图像空间特征作为键和值送入注意力模块。注意力模块的输出是高维空间特征。这些空间特征利用车辆里程计在时间上进行对齐,以推导运动。接下来,这些时空特征经过一组反卷积,生成最终的占用和占用流输出。它们以固定大小的体素网格形式构成,而这对于规划和控制来说可能不够精确。为了获得更高的分辨率,我们还会生成逐体素特征图,并将其与3D空间点查询一起输入MLP,从而获得任意位置上的位置和语义。在更好地了解模型之后,让我们看看另一个例子。这里有一辆铰接式公交车停在道路右侧,在这里以一个L形体素突出显示。当我们靠近时,公交车开始移动。汽车的前部首先变成蓝色,这表明模型预测。公交车的前部具有一条很长的零占用流。随着公交车继续移动,整辆公交车都变成蓝色,而且你还可以看到网络预测出了公交车精确的曲率。对于传统目标检测网络而言,这是一个非常复杂的问题,因为你必须判断我是要用一个长方体,还是可能用两个来适配这种曲率。但对于占用网络而言,因为我们关心的只是可见空间中的占用情况,所以我们能够精确地建模这种曲率。除了体素网格之外,占用网络还会生成一个胡言乱语表面。这个胡言乱语表面同时包含3D几何形状和语义。它们对控制非常有用,尤其是在多丘陵和多弯道的道路上。表面和体素网格并不是彼此独立预测的。实际上,体素网格会隐式地与表面对齐。这里,我们正处于一个山丘任务中,你可以看到表面的3D几何形状得到了很好的预测。规划器可以利用这些信息决定,也许我们需要为这个山丘任务进一步减速。而且你还可以看到,体素网格始终与表面对齐。除了体素和表面,我们也对神经辐射场或NERF近期取得的突破感到非常兴奋。我们正在研究将一些最后的NERF特征纳入占用网络训练,同时也研究将我们的网络输出用作NERF的输入状态。事实上,Ashok对此非常兴奋。
第 20 段
有一段时间以来,这一直是他的个人周末项目。关于这些 NERF,因为我认为学术界正在利用海量大型语言数据集构建这些语言基础模型。但我认为,就视觉而言,NERF 将为计算机视觉提供基础模型,因为它们以几何为基础。而几何为我们提供了一种很好的方法来监督这些网络,并冻结掉定义本体的要求。而这种监督基本上是免费的,因为你只需要以可微方式渲染这些图像。所以我认为,未来这种占用网络的理念,也就是输入图像,然后由网络生成一致的场景体积表示。随后可以将其以可微方式渲染成任何曾被观察到的图像。我个人认为这是计算机视觉的未来,我们现在正在对此开展一些初步工作。但我认为未来,无论是在 Tesla 还是学术界,我们都会看到,这种对体积占用进行单次预测的组合将成为未来。这是我个人的押注。谢谢 Ashok。这里展示的是使用我们的免费数据进行 3D 重建的一个早期结果示例。我们的首要目标不是在图像空间中实现完美的 RGB 复现,而是在 3D 空间中准确表示用于驾驶的世界。我们希望针对全球所有天气和光照条件下的全部免费数据做到这一点。显然,这是一个非常具有挑战性的问题,我们希望你们能来帮助。最后,占用网络使用大型自动标注数据集进行训练,全程无需人工参与。接下来,我把时间交给 Tim,由他介绍训练这个网络需要些什么。谢谢 Phil。好的,大家好。我们来谈谈一些训练基础设施。我们已经看过几个视频,不,是 4 个或 5 个,我想,而且更加关心、更加担心这上面的更多片段。我们刚刚一直在看 Phil 介绍的占用网络。仅 Phil 的那些视频,训练该网络就需要 14 亿帧。也就是你们刚才看到的内容;如果你有 100,000 个 GPU,需要 1 小时。但如果你只有 1 个 GPU,就需要 100,000 小时。所以这不是一个人能够等着训练任务跑完的合理时长,对吧?我们希望以比这更快的速度交付。因此,这意味着你需要采用并行方式。为此,你需要更多算力。这意味着你需要一台超级计算机。因此,我们在内部建造了 3 台超级计算机,由 14,000 个 GPU 组成。其中,我们使用 10,000 个 GPU 进行训练,使用约 4,000 个 GPU 进行自动标注。所有这些视频都存储在一个容量为 30 PB 的分布式托管视频缓存中。你不应该认为我们的数据集是固定的。比方说,就像你看待图像网络或类似东西那样,你知道,它大概有 100 万帧。你应该把它看成一种流动性很强的东西。所以,我们每天都有 50 万个这样的视频流入和流出这个集群。这些集群中的每一个。我们每秒追踪 400,000 个这种 Python 视频实例化。所以调用量非常大。为了管控这个分布式视频缓存的保留策略,我们需要捕获这些信息。因此,支撑这一切的是大量基础设施,全部由我们在内部构建和管理。所以你不能只是买来 14,000 个 GPU,再买 30 PB 的闪存 NVMe。然后把它们组装起来,就开始训练。实际上,这需要大量工作,我会稍微介绍一些。通常,你真正想做的是拿到你的加速器。它可以是 GPU,也可以是 Dojo,我们稍后会谈到。因为那是最昂贵的组件,所以你希望瓶颈出现在那里。因此,这意味着系统的每一个部分都需要拥有超过这款加速器的性能。所以这真的非常复杂。这意味着你的存储系统需要具备足够的容量和带宽,将所有数据输送到节点中。这些节点需要具备适当的 CPU 和内存能力,为你的机器学习框架提供数据。然后,这个机器学习框架需要把数据交给 GPU,之后你才能开始训练。但接着,你还需要跨数百或数千个 GPU,以可靠且步调一致的方式完成这项工作。而且还要足够快,所以你还需要互连系统。极其复杂。稍后我们会进一步讨论 Dojo。首先,我想带你们了解一下我们对集群进行的一些优化。我们正在接收大量视频,而视频与比方说图像或文本训练非常不同,我认为后两者已经非常成熟。从字面意义上说,视频要多复杂一个维度。因此,我们需要从存储层一直到加速器进行端到端优化。优化其中的每一个部分。因为我们使用直接来自车队的光子计数视频进行训练。我们直接使用这些视频训练,完全不对它们进行后处理。具体做法是,我们精确定位到为批次选定的帧。把它们以及它们依赖的帧一起加载进来,也就是你的眼帧或关键帧。我们把它们打包,移入共享内存,再移入来自 GPU 的双条,然后使用仅经过加速的硬件解码器实际解码视频。因此,我们以原生方式在 GPU 上完成这项工作,而且这一切都封装在一个非常好用的 PyTorch 扩展中。这样做使占用网络的训练速度提高了 30% 以上。而且基本上释放了整颗 CPU,让它可以做任何其他事情。当然,你不能只用视频进行训练,还需要某种真值。而这其实也是一个有意思的问题。存储真值的目标是,你要确保以最少的文件系统操作获取所需的真值。并且仅加载所需内容的最小体量,从而优化集群整体的总吞吐量。因为你应该把计算集群视为一个内部具有固定约束和阈值的大型设备。为此,我们推出了一种原生自有格式,叫作 small。我们使用它来存储真值、特征缓存和所有推理输出。所以其中有很多张量。这里只画一张示意图,假设这是你想存储的表。那么,这就是它部署到磁盘上之后的样子。你要做的是,把任何想要建立索引的内容,例如视频时间戳,全部放入标头,这样在首次读取标头时,你就能确切知道该去磁盘上的哪个位置。然后,如果你有任何张量,就尝试转置维度,把另一个维度放到最后,作为连续维度。然后还要尝试不同类型的压缩。接着查看哪一种最优,再把那一种存储下来。如果你要做特征缓存,这其实是一个非常重要的技巧。机器学习网络输出无法辨识。稍微旋转一下维度,存储效率最多可提高 20%。然后,在存储时,我们还会按照大小对各列排序。这样,所有小列和小值都会放在一起。所以,当你定位单个值时,很可能会在一次读取中顺带读到更多稍后会用到的值。这样你就不需要再执行一次文件系统操作。我还可以继续说下去,而且我刚才确实说了很多;我只谈到了我们内部的 2 个项目。实际上,这是为了优化内部现有算力而开展的一项规模庞大、持续不断的工作的一部分。通过积累并汇总所有这些优化,我们现在训练占用网络的速度提升到了原来的 2 倍,仅仅是因为效率提高到了 2 倍。现在,如果我们再加入更多算力并采用并行方式,就能在几小时而不是几天内完成训练。接下来,我想把时间交给算力的最大用户 John。大家好,我叫 John Emmons,负责 Autopilot 视觉团队。今天我要向大家介绍 2 个主题,第一个是我们如何预测车道。第二个是我们如何预测道路上其他参与者的未来行为。在 Autopilot 的早期阶段,我们把车道检测问题建模为一项图像空间即时分割任务。不过,我们的网络超级简单,实际上,它只能根据几种不同类型的几何形态来预测车道。具体而言,它会分割本车道,也可以分割相邻车道,另外还针对分岔和汇入设置了一些特殊情况。这种对问题的简化建模适用于高速公路等高度结构化的道路。但如今,我们正试图构建一个能够完成复杂得多的操控动作的系统。具体来说,我们希望在交叉路口左转和右转,而那里的道路拓扑可能要复杂、多样得多。当我们尝试在这里应用这种对问题的简化建模时,它就彻底失效了。先退一步来看,我们在这里试图做的是预测一组稀疏的车道实例及其连通关系。我们希望让神经网络基本上预测出这张图,其中节点是车道段。而边则编码这些车道之间的连通关系。所以,我们有一个车道检测神经网络,它由 3 个组件构成。在第一个组件中,我们有一组卷积层、注意力层和其他神经网络层。它们对车辆上 8 个摄像头的视频流进行编码,并生成丰富的视觉表示。然后,我们利用粗略的道路级地图数据来增强这种视觉表示。我们通过一组额外的神经网络层对这些数据进行编码,并将其称为车道引导模块。这张地图不是高清地图,但它提供了许多有用的提示,包括交叉路口内部的车道拓扑、不同道路上的车道数量,以及一组能够帮助我们的其他属性。这里的前 2 个组件会生成一个以某种方式编码世界的稠密张量。但我们真正想做的是,把这个稠密张量转换成一组稀疏的车道及其连通关系。我们像处理图像描述任务一样处理这个问题,其中输入是这个稠密张量,输出文本则以我们在 Tesla 开发的一种特殊语言进行预测,用来编码车道及其连通关系。在这种车道语言中,单词和词元是 3D 空间中的车道位置。在词元的排序中,词元里经过加密的修饰符会编码这些车道之间的连接关系。通过把这项任务建模为语言问题,我们可以利用语言领域最新的自回归架构和技术,来处理这个问题的多重二元性。在 Autopilot,我们不只是在解决计算机视觉问题,还在更广泛地应用语言建模和机器学习领域的最先进技术。现在,我会进一步介绍这个语言组件的一些细节。我在屏幕上展示的是一幅卫星图像,它以某种方式表示车辆周围的局部区域。这组鼻子和边就是我们所说的车道图,也是我们最终希望这个神经网络输出的内容。我们从空白状态开始。我们要在这个绿点处作出第一个预测。这个绿点的位置被编码为一个路线网格中的索引,该网格将 3D 世界离散化。现在,我们不会直接预测这个索引,因为这样做的计算成本太高。网格点实在太多,而预测它们之上的分类分布会同时对训练时和测试时产生影响。所以,我们的做法是先对世界进行粗略离散化,然后预测这个可能位置上的热力图,然后我们锁定最可能的位置。以此为条件,我们接着细化预测并得到精确的点。现在我们知道了这个词元的位置,但我们不知道它是否紧密。不过在这个案例中,它是一条新车道的起点。所以我们将它预测为一个起始词元。因为它是起始词元,所以在我们的语言中没有额外属性。然后,我们取得第一次前向传递的预测结果,并使用学习得到的位置嵌入对其进行编码。这会产生一组张量,我们把它们组合到一起。这实际上就是我们的车道语言中的第一个词。我们把它添加到这里句子的第一个位置。然后,我们以类似方式预测下一个车道点,继续这一过程。现在,这个车道点不是一条新车道的起点,它实际上是前一条车道的延续。所以它是延续词元类型。仅仅知道这条车道与此前预测的车道相连还不够。我们还想编码它的精确几何形状,我们通过回归一组样条系数来做到这一点。然后,我们取得这条车道,再次对它编码,并将它作为句子中的下一个词添加进去。我们继续预测这些延续车道,直到到达预测网格的末端。然后,我们转向另一个车道段。所以你可以看到那里的青色点。现在,它在拓扑上并不与那个粉色点相连。它实际上是从那里的绿色点分叉出来的。所以它具有分叉类型。而分叉词元实际上会反向指向它们的分叉所源自的先前词元。所以你可以在这里看到,分叉点预测器实际上是索引 0。因此,它实际上是在反向引用一个已经预测出来的词元,就像你在语言中所做的那样。我们一遍又一遍地继续这个过程,直到枚举完车道图中的所有词元。然后,网络预测句末词元。是的,我只是想指出,我们这么做的原因并不只是因为我们想构建某种复杂的东西。虽然这里的神经网络几乎感觉像一台图灵完备机器。原因是我们尝试过简单的方法,例如尝试只对道路沿线的车道之类的东西进行分割。但问题是,当存在不确定性时,比如说你无法清楚地看到道路。那里可能有 2 条车道,也可能有 3 条车道,而你无法判断。简单的基于分割的方法会把它们全都画出来。这有点像 2.5 条车道的情况。而当预测是这样的时候,后处理算法会滑稽地失败。是的,问题还不止于此。我的意思是,你需要预测交叉路口内部的这些连接车道。而用 Ashok 提到的方法,这根本不可能做到。这就是为什么我们不得不升级为这种方法。是的,当它像这样重叠时,分割就会完全失控。但即使你非常努力地尝试把它们放在不同的层上,这仍然只是一个非常困难的问题。而语言恰好提供了一个非常好的框架,可以从后验分布中获得一个样本,而不是试图在后处理中完成所有这些工作。但这实际上并不只适用于 Autopilot,对吧?John,这可以用于乐观主义者。是的,我想它们不会被称为车道。但你可以想象,大概在这里这个阶段,你可能会有某种路径,以某种方式编码人们可能行走的位置。是的,基本上,如果你身处工厂或家庭环境中,你只需要求机器人:好的,请规划前往厨房的路线,或者请规划前往工厂中某个位置的路线。然后,我们预测一组会穿过通道的路径,引导机器人。并说,好的,这就是你前往厨房的方式。它确实为我们提供了一个很好的框架,用来对这些不同路径建模。这些路径能为下游规划器简化导航问题。好的,所以我们最终从这个车道检测网络得到的,是一组车道及其连接关系,它们直接来自网络。这里没有把这些密集预测稀疏化为稀疏预测的额外步骤。这就是网络直接输出的、未经筛选的结果。好的,我稍微谈了谈车道。接下来我会简要介绍我们如何对未来路径和物体上的其他语义进行建模与预测。所以我只会非常快速地讲 2 个例子。这里右侧的视频中,有一辆车实际上正在闯红灯,并在我们前方转弯。为了处理这样的情况,我们会为所有物体预测一组短时间范围的未来轨迹。我们可以利用这些轨迹预判这里的危险情况。并采取避免碰撞所需的任何破坏和转向操作。右侧的视频中,我们前方有 2 辆车。左侧车道上的那辆停着,显然正在装货、卸货。我不知道为什么司机决定把车停在那里。但重要的是,我们的神经网络预测它处于停止状态。也就是那里的红色。正如你注意到的,另一条车道上的车辆同样静止不动。但那辆车显然只是在等待红灯变绿。所以,尽管这 2 个物体都静止不动,速度为 0。这里真正重要的是语义。这样我们才不会被困在那辆停得很别扭的车后面。在尝试构建实时系统时,预测所有这些智能体属性会带来一些实际问题。我们需要最大限度提高对象部分堆栈的帧率。这样 Autopilot 才能对不断变化的环境迅速作出反应。这里每 1 毫秒都真的很重要。为了尽量降低推理延迟,我们的神经网络被分为 2 个阶段。在第一阶段,我们识别智能体存在于 3D 空间中的哪些位置。在第二阶段,我们接着取出那些 3D 位置上的张量。将其与车辆上的额外数据附加在一起。然后完成其余处理。这个具体化步骤让神经网络能够把算力集中在最重要的区域。这使我们只需付出一小部分延迟成本,就能获得卓越的性能。所以,把所有这些结合起来。Autopilot 视觉堆栈预测的不只是世界的几何结构和运动学。它还会预测一组丰富的语义,从而实现安全且像人类一样的驾驶。现在我要把时间交给 Sri,他会告诉我们如何在 FSD 计算机上运行所有这些很酷的神经网络。谢谢。大家好,我是 Sri。今天,我将让大家大致了解在车内运行这些 FSD 网络需要做些什么。以及我们如何优化推理延迟?今天,我只会重点介绍 John 刚才谈到的 FSD 车道网络。所以,当我们开始这项工作时,我们想知道能否在 trip 引擎上原生运行这个 FSD 车道网络。trip 引擎是我们在 FSD 计算机中构建的内部神经网络加速器。构建这个硬件时,我们让它保持简单,并确保它能以快得离谱的速度做好一件事:密集点积。但这个架构是自回归且迭代式的。它会在内循环中处理多个注意力-注意力块。在每一步直接生成稀疏点。所以,这里的挑战是,我们怎样才能在一个密集点积引擎上完成这种稀疏点预测和稀疏计算。让我们看看我们是怎样在 trip 上实现这一点的。所以,网络会预测点最可能的空间位置的热力图。为了在 trip 上实现这一点,我们实际上在 SRAM 中构建了一个查找表。并且我们设计了这个嵌入的维度,使我们只用矩阵乘法就能完成所有这些事情。不仅如此,我们还想把这个嵌入存入词元缓存。这样我们就不必在每次迭代中重新计算它,而是可以为未来的点预测复用它。同样,我们在这里用了一些技巧,只在点积引擎上完成所有这些运算。实际上很酷的是,我们的团队找到了创造性的方法,将所有这些运算映射到 trip 引擎上。而这些方式在这个硬件设计之初甚至都没有被设想过。但为了让它正常工作,我们要做的并不只有这些。实际上,我们实现了大量运算和功能,来使这个模型能够被编译。以提高 intate 准确率并优化性能。所有这些都帮助我们以不到 10 毫秒的延迟运行这个拥有 7500 万个参数的模型。功耗仅为 8 瓦。但这并不是车内运行的唯一架构。我们还需要在车内运行许多其他架构、模块和网络。为了让大家感受一下规模,所有网络合计约有 10 亿个参数。会产生大约 1000 个神经网络信号。所以我们需要确保对它们进行联合优化,从而最大限度提高算力利用率和吞吐量,并最大限度降低延迟。因此,我们专门为神经网络构建了一个编译器,它与传统编译器具有相似的结构。如你所见,它接收一个包含 15 万个节点和 37.5 万条连接的庞大神经网络图。取得这个东西,将其划分为相互独立的子图。然后针对推理设备原生编译每一个子图。之后,我们还有一个与传统链接器结构相似的神经网络链接器。我们会在其中执行这种链接时优化。在那里,我们解决一个受算力、内存和内存带宽约束的离线优化问题。从而生成一个在车内执行的优化调度方案。在运行时方面,我们设计了一个混合调度系统,它基本上会在 1 个 SOC 上进行异构调度。并在 2 个 SOC 之间进行分布式调度,以模型并行方式运行这些网络。为了达到 100 TOPS 的算力利用率,我们需要对软件的所有层级进行优化。从调整网络架构和编译器,一直到实现一个低延迟、高带宽的 RDMA 链路。这个链路横跨 2 个 SOC,事实上还要进一步深入,理解并优化 SOC 中加速器的缓存一致和非一致数据路径。为了确保获得最高帧率,需要在每一个层级进行大量优化,因为这里每 1 毫秒都很重要。而这只是车内所运行神经网络的可视化。这本质上就是我们的数字大脑。如你所见,这些运算无非就是矩阵乘法、卷积,仅举几个在车内运行的真实运算为例。要训练这个拥有 10 亿个参数的网络,你需要大量标注数据。所以 Egan 将谈谈我们如何通过自动标注流水线实现这一点。谢谢你,Sri。大家好,我是 Egan Zhang,我负责 Autopilot 的几何视觉工作。那么,是的,让我们谈谈自动标注。我们有多种自动标注框架来支持各种类型的网络。但今天我想重点介绍这里这个很棒的车道网络。所以,为了成功训练这个网络并使其泛化到各个地方,我们认为我们从可能 100 万个甚至更多的交叉路口走过了数千万次行程。那要怎么做到呢?获得足够数量的行程当然是可以实现的,因为正如 Tim 之前解释的,我们已经拥有大约每天 50 万次行程的缓存率。然而,将所有这些数据转换为训练形式是一个极具挑战性的技术问题。为了解决这个挑战,我们尝试了各种人工标注和自动标注方式。从第 1 列到第 2 列,从第 2 列到第 3 列,每一次进步都为我们带来了近 100 倍的吞吐量提升。但我们仍然运行着一台更出色的自动标注机器,它可以为我们提供良好的质量、多样性和可扩展性。为了满足所有这些要求,尽管数量巨大,这里所需的工程工作方面,我们开发了一台由多行程重建驱动的新型自动标注机器。因此,对于10,000次行程的标注,它只需在集群上运行12小时,就能取代500万小时的人工标注。那么我们是怎么解决的?有3大步骤。第一步是通过多摄像头、视觉、惯性或几何进行高精度轨迹和结构恢复。因此在这里,包括地面在内的所有特征都由神经网络从视频中推断出来,然后在向量空间中进行跟踪和重建。因此,这条车内轨迹的典型行程率大约是每米1.3厘米和每米0.45毫升,考虑到其紧凑的计算需求,这相当不错。随后,恢复出的表面和道路细节也被用作后续人工验证工作的强力指导。这也在每辆FSD车辆中启用,因此我们会随行程数据一起获得预处理后的轨迹和结构。第二步是多行程重建,这是这台机器庞大而核心的部分。视频展示了之前显示的行程如何被重建并与其他行程对齐,基本上是来自不同车辆的其他行程,而不是同一辆车。因此,这是通过多个内部步骤完成的,比如粗略对齐、成对匹配、联合优化,然后进一步进行表面细化。最后,由人工分析员介入并最终确定标签。因此,每个繁重的步骤都已经在集群上完全并行化,所以整个过程通常只需要几个小时。最后一步实际上是自动标注新行程。因此在这里,我们使用相同的多行程对齐引擎,但只在预先构建的重建结果和每个新行程之间使用。所以这比把所有片段完全放在一起重建要简单得多。正因如此,自动标注每次行程只需30分钟,而不是数小时的人工标注。这也是这台机器可扩展性的关键。只要我们有可用的算力和行程数据,这台机器就能轻松扩展。这个场景中新自动标注了大约50次行程,其中一些显示在这里,因此有来自不同车辆的53次行程。这就是我们如何捕获世界的时空切片,并将其转化为网络监督。我想指出的一点是,Jagan刚才谈到了我们如何自动标注车道。我们几乎为所做的每项任务都配备了自动标签,包括我们的规划器。其中许多是完全自动的,没有人参与。例如,对于物体,所有运动学信息、形状、未来状态,一切都完全来自自动标注。我们的占用情况也是如此,而我们实际上已经围绕这件事构建了一台机器。是的,所以如果你能往回翻一张幻灯片。再往回一张,上面写着在集群上并行化。这听起来相当直接,但实际上并非如此。也许分享一下这种东西是如何形成的会很有意思。前些时候,我们完全没有任何自动标注,然后有人编写了一个脚本。它开始起作用,开始运行得更好,直到你的处理量变得相当高。我们显然需要一个解决方案。所以我们团队里还有另外2名工程师,他们当时就说,你知道,那是个有意思的,你知道,东西。我们需要做的是构建一整张图,本质上由需要一个接一个运行的Python函数组成。首先拉取片段,然后做一些清理,然后进行一些网络推理,再进行另一次网络推理。直到最终得到这个。但你需要大规模完成这件事,所以我告诉他们,我们大概需要以,你知道,每天100,000个片段为目标。或者说100,000个条目,这似乎不错。于是工程师们说,嗯,我们可以用一点Postgres,再下点苦功,我们能做到。与此同时,过了一段时间,我们现在每天都会执行2,000万个这样的函数。再说一次,我们拉取大约50万个片段,并以流式方式在这些片段上运行大量函数,每一个都是如此。这就是不仅运行训练,而且运行自动标注所需的后端基础设施。是的,它确实就像一座生产标签的工厂,有生产线、产量、质量、库存。所有这些相同的概念都应用到了这座标签工厂,就像它们应用于,你知道,我们的汽车工厂一样。没错。好的,谢谢Tim和Ashok。那么,是的,为了结束这一部分,我想再分享几个对网络来说肯定更具挑战性且有趣的例子。甚至对人类来说可能也是如此。从上面开始,有一些例子,比如光线不足的情况、雾夜、环岛,以及被停放车辆严重遮挡的情况。甚至还有雨夜,摄像头镜头上带着雨滴。这些都很有挑战性,但一旦其他片段完整重建了它们的原始场景,所有这些都可以被自动标注。这样我们的汽车就能更好地驶过这些具有挑战性的场景。那么现在,我把麦克风交给David,进一步了解Sim如何在这些标签之上创造新世界。谢谢。谢谢你,Yegan。我叫David,我要谈谈仿真。仿真在提供难以获取和/或难以标注的数据方面发挥着关键作用。然而,3D场景的制作速度是出了名的慢。以我身后正在播放的仿真场景为例。这是旧金山市场街的一个复杂十字路口。艺术家需要2周才能完成。对我们来说,这慢得令人痛苦。不过,我要谈的是如何利用Yegan的自动化真值标签以及一些全新的工具,让我们仅用5分钟就能以程序化方式生成这个场景以及许多类似场景。这比以前快了惊人的1,000倍。那么,让我们深入了解这样的场景是如何创建的。我们首先把自动化真值标签输送到Houdini软件内的仿真世界创建工具中。从道路边界标签开始,我们可以生成实体道路网格,并使用车道图标签对其重新拓扑。这有助于确定重要的道路细节,比如横向道路坡度和细致的材质混合。接下来,我们可以使用线数据,让几何体沿其表面扫掠并将其投射到道路上,从而创建车道标线贴花。接下来,利用中央分隔带边缘,我们可以生成交通岛几何体,并用随机化植被填充它。这会极大改变场景的可见性。现在,外部世界可以通过一系列随机化启发式规则生成。模块化建筑生成器会制造视觉遮挡,而随机放置的消防栓等物体可以改变曲线的颜色,树木则可以让树叶落在下方,遮住线条或边缘。接下来,我们可以引入地图数据,以确定交通交通灯或停车标志等事物的位置。我们可以沿其法线进行追踪,以收集车道数量等重要信息,甚至能在标志本身上获得准确的街道名称。接下来,利用车道图,我们可以确定车道连通性,在道路上生成方向性道路标记及其配套道路标志。最后,利用车道图本身,我们可以确定车道邻接关系以及其他有用指标,从而在我们的模拟器中生成随机化的交通排列。再说一次,这一切都是自动完成的,过程中没有艺术家参与,并且会在几分钟内完成。现在,这使我们能够做一些相当酷的事情。由于一切都基于数据和启发式规则,我们可以开始对参数进行模糊测试,为单一真值创建视觉变体。它既可以是物体摆放和随机材质替换这样细微的变化,也可以是全新生物群系或城市、郊区、乡村等环境位置这样更为剧烈的变化。这使我们能够针对需要更多真值的特定真值,创建无限且有针对性的排列。而这一切只需点击一个按钮即可完成。我们甚至还可以更进一步,修改真值本身。假设John希望他的网络更多关注方向性道路标记,以便更好地检测前方即将出现的受限左转车道。我们可以开始在模拟器中以程序化方式修改车道图,帮助创建穿过这个十字路口的全新通行流线,从而帮助网络把注意力集中到道路标记上,做出更准确的预测。这很好地展示了这套工具如何让我们创建永远无法从现实世界采集的新数据。这款工具真正的力量在于它的架构,以及我们如何并行运行所有任务以实现无限扩展。你们刚才看到了图块创建器工具的运行,它把真值标签转换为对应内容。接下来,我们可以使用图块提取器工具,把这些数据划分为面积约150米见方的地理哈希图块。然后,我们将这些数据分别保存到几何体文件和实例文件中。这样我们就有了一个易于加载的干净数据源,并使我们未来不依赖特定渲染引擎。然后,我们可以使用图块加载器工具,通过地理哈希ID调取任意数量的缓存图块。目前,我们处理的大约是这些5x5图块,或者通常是3x3图块,以车队热点或有意思的车道图位置为中心。图块加载器还会把这些图块集转换为U资产,供虚幻引擎使用,并根据你们在第一张幻灯片中看到的内容生成成品。这确实让我们为规模和扩展做好了准备。正如你们在我们身后的地图上看到的,我们可以轻松生成旧金山市的大部分街道。而这并没有花费数年甚至数月的工作,而是由1个人在2周内完成。我们可以继续利用工具内部的PDG网络管理并扩充所有这些数据。这让我们可以投入算力,在一夜之间重新生成所有这些图块集。这确保了所有环境在质量和特征方面保持一致,这对训练极其重要,因为新的本体和信号会不断发布。现在再回到原点,因为我们根据真值数据生成了所有这些图块集,它们包含现实世界中所有奇怪的复杂细节。我们可以将其与程序化的视觉和交通多样性相结合,创建无限且有针对性的数据,供网络学习。SIM部分到此结束,我把话筒交给Kate,请她谈谈我们如何利用所有这些数据改进Autopilot。谢谢。谢谢David,大家好,我叫Kate Park,我来这里是要谈谈数据引擎。它是我们通过数据改进神经网络的流程。我们将向你们展示如何通过数据以确定性方式解决人工接管问题。并带你们了解这个特定片段的整个生命周期。在这个场景中,Autopilot正在接近一个转弯处,并错误地预测那辆横向车辆因交通状况而停下,因此认为它是一辆我们应该为之减速的车辆。实际上,车里没有人,它只是以一种别扭的方式停在那里。我们构建了这套工具,用来识别错误预测、修正标签,并将这个片段归入一个评估集。这个特定片段恰好是我们诊断为转弯处具有挑战性的停放车辆的126个片段之一。得益于这套基础设施,我们无需针对这一特定挑战案例投入任何定制工程资源,就能整理出这个评估集。要真正解决这个挑战案例,需要挖掘数千个类似示例。而这是Tesla可以轻松做到的事情。我们只需使用数据获取基础设施,请求数据,并使用之前展示的工具修正标签。通过精准锁定当前模型的错误预测,我们只会把最有价值的示例添加到训练集中。我们精准修复了13,900个片段,而由于这些都是当前模型难以处理的示例,我们甚至不需要改变模型架构,使用这些新的高价值数据进行一次简单的权重更新就是足以解决这个挑战性案例。所以你可以看到,我们不再把那辆横穿的车辆预测为已停下(如橙色所示),而是预测为已停放(如红色所示)。在学术界,我们经常看到人们保持数据不变,但在 Tesla,情况完全相反。我们一次又一次地看到,数据是解决这些干预问题的最佳杠杆之一,即使不是最具决定性的杠杆。我们刚刚向你们展示了一个挑战性案例的数据引擎闭环,也就是这些在转弯处停放的车辆。但即使只针对车辆运动这一个信号,也存在许多挑战性案例。我们把这个数据引擎闭环应用于我们诊断出的每一个挑战性案例,无论是公交车、弯曲道路、停止的车辆,还是停车场。我们并不只是添加一次数据,而是一次又一次地这样做,以完善语义。事实上,今年我们更新了5次车辆运动信号,并且每次权重更新都使用新数据进行训练。我们不断提高车辆运动判断的准确率。这个数据引擎框架适用于我们的所有信号,无论它们是3D数据还是多摄像头视频,无论数据是人工标注、自动标注还是模拟生成的,无论是离线模型还是在线模型。Tesla 能够大规模做到这一点,是因为车队优势、我们的 NG 团队构建的基础设施,以及为我们的网络提供数据的标注资源。为了使用所有这些数据进行训练,我们需要极其庞大的算力,所以接下来我会把时间交给皮特和加内什,请他们介绍 Dojo 超级计算平台。谢谢。谢谢你,凯蒂。谢谢大家,谢谢你们坚持到现在,我们快讲完了。我叫皮特·班农,负责 Tesla 的定制芯片和低电压团队。我叫加内什·伦克,负责 Dojo 项目。谢谢。我经常被问到,为什么一家汽车公司要建造一台用于训练的超级计算机?
第 21 段
而这个问题从根本上误解了 Tesla 的本质。从核心来看,Tesla 是一家硬核科技公司。在整个公司里,人们都在科学和工程领域努力工作,以推进我们在制造汽车、能源解决方案、机器人以及任何其他能够改善全世界人类生活状况的事物方面所掌握的基础认识和方法。能参与其中是一件超级令人兴奋的事,而能负责其中半导体团队里非常小的一部分,是一种荣幸。今晚我们会谈一点 Dojo,并向大家介绍过去1年里我们取得的进展。但在此之前,我想先简单介绍一下我们几年前开始进行的初始设计。刚开始时,我们的目标是大幅改善自动驾驶团队的训练延迟。他们如今训练的一些最大型神经网络需要运行超过1个月,这妨碍了他们快速探索替代方案并对其进行评估。因此,如果我们能以在成本和能耗上具有竞争力的方式实现30倍加速,那会非常好。为此,我们希望打造一款拥有大量算术单元的芯片,并且能够以非常高的效率利用这些单元。我们花了很多时间研究能否利用 DRAM 和各种封装构想做到这一点,但全都失败了。最终,尽管这感觉像是一种反常之举,我们还是决定放弃将 DRAM 作为这个系统的主要存储介质,转而专注于嵌入芯片的 SRAM。遗憾的是,SRAM 提供的容量不大,但带宽极高、延迟极低,这使我们能够让算术单元实现高利用率。那些选择,那个特定的选择,又引出了一大堆其他选择。例如,如果你想拥有虚拟内存,就需要页表,而页表要占用大量空间;我们没有空间,所以没有虚拟内存。因此我们也没有中断,这个加速器是一块裸键、原始的硬件,它被呈现给编译器,而编译器负责以确定性的方式调度发生的一切。所以系统中不需要中断,甚至也不希望有中断。我们还选择采用模型并行作为训练方法,这不是典型情况;如今大多数机器使用数据并行,而这会消耗额外的内存容量,显然我们没有这些容量。因此,所有这些选择促使我们打造出一台与如今现有设备截然不同的机器。我们还有一大堆其他目标,其中最重要的目标之一就是没有限制。因此,我们希望打造一种计算结构,在大多数情况下都能不受限制地扩展,我的意思是,显然这里那里还是存在物理限制。但基本上,如果你的模型对这台计算机来说太大,你只需要去买一台更大的计算机,这就是我们所追求的。如今机器的封装方式使 GPU、CPU、DRAM 容量和网络容量等要素之间存在相当固定的比例。我们确实希望把这一切解耦,这样随着模型演进,我们就可以改变这些不同要素之间的比例,让系统更加灵活,以满足自动驾驶团队的需求。而事实的确如此,“没有限制”的理念始终是我们的指路星,我们所有的选择都围绕着它展开。甚至到了这样的程度:我们不希望传统数据中心基础设施限制我们高速执行这些程序的能力。这就是为什么我们对数据中心进行了垂直整合,对整个数据中心进行了垂直整合。通过对数据中心进行垂直整合,我们能够发掘新的效率水平,能够优化整个数据中心堆栈中的供电、冷却以及系统管理,而不是逐个盒子进行处理,再把那些盒子集成到数据中心里。为此,我们也希望尽早进行集成,从而找出软件工作负载的规模极限。因此,我们很早就把 Dojo 环境集成到了自动驾驶软件中,并从中吸取了很多经验。今天,Bill Chang 将介绍我们的硬件更新,以及我们一路上遇到的一些挑战。Rajiv Kurian 将让大家一窥我们的编译器技术,并介绍我们取得的一些很酷的成果。很好。谢谢 Pete,谢谢 Ganesh。今晚我会先从我们系统的高层愿景讲起,这将有助于为我们正在解决的挑战和问题奠定背景,然后也会讲到软件将如何利用它来获得性能。我们对 Dojo 的愿景是打造一个单一、统一而且非常大型的加速器。软件将看到一个无缝的计算平面,其中包含可全局寻址的超高速内存,而且所有部分都通过统一的高带宽和低延迟连接在一起。要实现这一点,我们需要利用密度来获得性能。我们利用技术获得这种密度,从而打破从芯片一直到横向扩展系统的各层级。数十年来,硅技术一直在这样做。芯片一直遵循摩尔定律进行密度集成,从而获得性能扩展。实现这一愿景的关键一步是我们的训练单元块。或许我们可以用极高带宽集成25个裸片,但只需把它们连接在一起,就能将其扩展至任意数量的额外单元块。去年,我们展示了首个能够正常工作的训练单元块,当时已经有工作负载在上面运行。此后,这里的团队一直在努力而勤勉地推进其规模化部署。我们取得了惊人的进展,并一路达成了许多里程碑。当然,我们也遇到了许多意料之外的挑战。但正是在这里,我们的快速失败理念让我们得以突破自己的边界。通过提高密度来获得性能带来了全新的挑战。其中一个领域是供电。我们需要在这里为计算裸片供电,而这会直接影响我们的顶线计算性能。但我们需要以前所未有的密度做到这一点。我们需要能够匹配裸片间距,同时达到接近每平方毫米1安培的功率密度。由于集成程度极高,这需要采用多层垂直供电解决方案。又因为这里存在复杂的异质材料堆叠,我们必须谨慎管理材料之间的过渡,尤其是 CTE。那么在这种情况下,热膨胀系数为什么重要?CTE 是一种基本材料属性,如果不加以谨慎管理,那个堆叠结构实际上会把自己撕裂。我们一开始与供应商合作开发这种供电解决方案。但后来意识到,我们实际上必须在内部自行开发。为了平衡进度和风险,我们构建了快速迭代版本,既支持系统启动和软件开发,也用于找出能够满足最终生产目标的最优设计和堆叠结构。最终,我们成功将 CTE 降低了50%以上,并使性能达到初始版本的3倍。不用说,要在密度条件下最大化性能的同时找到这种最优材料堆叠结构,是极其困难的。我们一路上的确遇到了一些意料之外的挑战。这里有一个例子:我们突破了集成的边界,结果导致元件故障。最初是在我们扩展到更大、更长的工作负载时,一个单元块上的某个单一位置会间歇性失效。开始时这些还是可恢复的故障,但随着我们把功率推得越来越高,它们会变成永久性故障。要理解这种故障,你必须理解我们为何以及如何构建供电模块。解决每个层级的密度问题,是实际实现系统性能的基石。因为我们的 XY 平面用于高带宽通信,所以其他所有东西都必须垂直堆叠。这意味着除裸片之外的所有其他元件都必须集成到供电模块中。这包括时钟、电源以及系统控制器。在这个案例中,故障是由于振荡器的时钟输出丢失所致。经过广泛调试,我们发现根本原因是附近的电容器产生压电效应,导致模块振动。会鸣叫的电容器并不是什么新现象,事实上在电源设计中非常常见。但通常时钟芯片会放在电路板上非常安静的区域,往往不会受到电源电路的影响。不过,因为我们需要实现这种程度的集成,这些振荡器必须放置在非常近的位置。由于我们的开关频率以及随后产生的振动共振,它在 MEMS 振荡器上引发了平面外振动,导致振荡器开裂。这个问题的解决方案需要采取多管齐下的方法。我们可以使用软端子电容器来减少振动。我们可以更新 MEMS 部件,使其在平面外方向上具有更低的 Q 因子。我们还可以更新开关频率,让共振进一步远离这些敏感频段。除了系统层面的密度之外,我们在基础设施层面也取得了许多进展。我们知道,为了支持前所未有的功率和冷却密度,必须阅读检查数据中心基础设施的每一个方面。我们引入了完全定制设计的 CDU,以支持 Dojo 的高密度冷却需求。令人惊叹的是,相比购买现成设备再进行改造,我们只花了其成本的一小部分就做到了这一点。由于我们的 Dojo 机柜集成了足以匹配整整一排标准 IT 机架的供电和冷却能力,我们需要把机柜和基础设施结合起来谨慎设计。我们已经对这个机柜进行了数次迭代,以优化设计。今年早些时候,我们开始对供电和冷却基础设施进行低测试,并成功将其推到了超过2兆瓦,随后变电站跳闸,我们还接到了市政府打来的电话。去年,我们只介绍了系统的几个组件:定制 D1 裸片和训练单元块,但我们曾预告,退出舱是我们的最终目标。接下来我们会逐一介绍构建这个退出舱所需的系统其余部分。系统托盘是实现单一加速器愿景的关键部分。它让我们能够无缝连接各个单元块,不仅可以在机柜内部连接,也可以在不同机柜之间连接。我们可以在整个加速器中以非常紧密的间距连接这些单元块。这就是我们实现统一通信的方式。这是一种层压母线,使我们能够集成极高功率、机械和热支撑,并实现极高密度的集成。它高75毫米,可承载6个单元块,重量为135千克。这相当于3到4个满载的高性能机架。接下来,我们需要向训练单元块输送数据。为此,我们开发了 Dojo 接口处理器。它为系统提供高带宽 DRAM,用于暂存训练数据。它还通过 TTP 向训练单元块提供完整的内存带宽,TTP 是我们用来在整个加速器中进行通信的定制协议。它还拥有高速以太网,帮助我们通过标准以太网扩展这种定制协议。我们为此提供原生硬件支持,几乎不产生软件开销。最后,我们可以通过标准 Gen4 PCIe 接口连接它。我们为每个托盘配备20张这样的卡,从而获得640 GB 的高带宽 DRAM。这为训练单元块提供了解耦的内存层。这些卡是高带宽输入路径,可同时通过 PCIe 和以太网输入。它们还提供一条高倍率 X-Z 连接路径,可以在我们的大型 Dojo 加速器中实现捷径。现在我们实际上将主机直接集成在我们的系统托盘下方。这些主机负责摄取处理,并通过 PCIe 连接到我们的接口处理器。这些主机可以为基于视频的训练提供硬件视频解码器支持。而我们的用户应用程序会部署在这些主机上,因此我们可以为它们提供标准的 X86 Linux 环境。现在,我们可以把其中 2 套组件装入一个机柜,并为其配备冗余电源,将三相 480 伏交流电直接转换为 52 伏直流电。现在,通过在每个层级都专注于密度,我们可以实现单一加速器的愿景。首先从我们定制 D1 裸片上的统一节点开始,我们可以在完全一体化的训练瓦片中将它们连接起来。最后再跨越机柜边界无缝连接它们,构成我们的 Dojo 加速器。综合起来,我们可以在 Exapod 中容纳 2 个完整的加速器,合计提供 1 exa-flop 的机器学习计算能力。总的来说,这种程度的技术和集成在计算史上只实现过寥寥数次。接下来,我们将看到软件如何利用它来提升性能。谢谢 Bill,我叫 Rajiv,接下来我要谈一些数字。我们的软件栈始于 PyTorch 扩展,这体现了我们开箱即用地运行标准 PyTorch 模型的承诺。我们还会进一步介绍 JIT 编译器,以及为硬件输入数据的摄取流水线。抽象地说,性能等于 TOPS 乘以利用率,再乘以加速器占用率。我们已经看到硬件如何提供峰值性能,而编译器的工作是在代码于硬件上运行时挖掘硬件的利用率。摄取流水线的工作则是确保数据能够以足够高的吞吐量输入,使硬件永远不会缺少数据。那么,我们来谈谈为什么通信受限模型难以扩展。但在此之前,先看看为什么类似 ResNet 50 的模型更容易扩展。你从单个加速器开始,运行前向和反向传播,然后运行优化器。接着,为了扩大规模,你在多个加速器上运行它的多个副本。虽然反向传播产生的梯度确实需要进行归约,并且这会引入一些通信,但这可以与反向传播以流水线方式执行。这种设置的扩展效果相当好,几乎呈线性。对于激活值大得多的模型,我们一想运行前向传播就会遇到问题。单个加速器所能容纳的批量大小通常小于批归一化表面。为了绕过这个问题,研究人员通常会在多个加速器上以同步批归一化模式运行这种设置。这会把受延迟限制的通信引入前向传播的关键路径,而我们本来就已经存在通信瓶颈。虽然有办法绕过这一问题,但通常涉及繁琐的手动工作,而这更适合交给编译器。归根结底,无法回避的事实是,如果你的状态无法装入单个加速器,你就可能受到通信限制。即使我们的机器学习工程师付出了巨大努力,我们仍然看到这类模型无法线性扩展。doger 系统正是为了让这类模型能以高利用率运行而构建的。高密度集成不仅是为了加速模型中受计算限制的部分,也是为了加速受延迟限制的部分,例如批归一化,以及受带宽限制的部分,例如梯度全归约或参数全收集。可以从 doger 网格中划分出一个切片来运行任何模型。用户唯一需要做的,就是确保该切片足够大,能够容纳其特定模型的批归一化表面。之后,该分区会将自身呈现为一个大型加速器,使用户不必担心执行的内部细节。而维护这种抽象正是编译器的工作。细粒度同步原语和统一的低延迟,使跨集成边界加速所有形式的并行处理变得容易。张量通常以分片形式存储在 SRAM 中,并在某一层执行前即时复制。我们依靠 doger 的高带宽来隐藏这种复制时间。张量复制和其他数据传输会与计算重叠进行,而编译器也可以在有利可图时重新计算各层。我们预计大多数模型都能开箱即用。例如,我们采用了最近发布的 Stable Diffusion 模型,并在几分钟内让它运行在 Dojo 上。编译器开箱即用便能以模型并行方式,将其映射到 25 个 Dojo 裸片上。这里是一些由运行在 Dojo 上的 Stable Diffusion 生成的火星上的 Cybertruck 图片。看起来,它距离达到 Tesla 设计工作室团队的水平还有一段路要走。我们已经讨论过通信瓶颈会如何妨碍扩展性。也许,对编译器及底层硬件的一项资产测试,就是执行跨裸片批归一化层。如前所述,这可能成为串行瓶颈。批归一化的通信阶段始于各节点计算其局部均值和标准差。然后协调对这些值进行归约,再把这些值广播回去,之后它们恢复并行工作。那么,在 25 个 Dojo 裸片上,理想的批归一化会是什么样?假设之前较少的激活值已经被拆分到各个裸片上。我们会期望每个裸片上的 350 个节点相互协调,生成裸片局部的均值和标准差值。理想情况下,这些值会进一步归约,最终值落在瓦片中部附近的某处。随后,我们希望看到这个值从中心向外辐射式广播。让我们看看编译器实际上如何跨 25 个裸片执行一次真实的批归一化操作。通信树提取自编译器,而计时来自一个真实硬件的。我们即将看到 25 个裸片上的 8,750 个节点相互协调,先归约批归一化的均值和标准差值,然后再广播。先进行裸片局部归约,然后朝瓦片中部进行全局归约。接着,归约后的值从中部向外辐射式广播,并由硬件的广播设施加速。这项操作在 25 个 Dojo 裸片上仅需 5 微秒。同一项操作在 24 个 GPU 上需要 150 微秒。这相较 GPU 是数量级上的提升。虽然我们是在批归一化的语境下讨论一种已经使用的操作,但需要再次强调,同样的优势适用于所有其他通信原语。而这些原语对于大规模训练至关重要。那么,完整模型的性能如何?虽然我们认为 ResNet 50 并不能很好地代表 Tesla 的真实工作负载,但它是一项标准基准测试,所以我们先从这里开始。我们已经能够在逐裸片比较中达到 100。不过,或许能够体现 Dojo 能力的一点是,我们只用每个裸片 8 的批量就能达到这个数字。但 Dojo 真正要解决的是更庞大、更复杂的模型。因此,当我们着手处理真实工作负载时,我们研究了当前 GPU 集群的使用模式。有 2 类模型尤为突出:自动标注网络,即一类用于生成真实标签的离线模型;以及你们听说过的占用网络。自动标注网络是具有高算术强度的大型模型,而占用网络可能受摄取限制。我们选择这些模型,是因为它们合计占据了当前 GPU 集群使用量的很大一部分,而且会以不同方式考验系统。那么,我们在这 2 种网络上的表现如何?接下来要看到的结果,在 GPU 和 Dojo 上都是通过多裸片系统测得的,但已归一化为每个裸片的数据。在我们的自动标注网络上,采用旧一代 VRM 运行的现有硬件已经能够超过 A100 的性能。在采用新款 VRM 的量产硬件上,这意味着吞吐量可达到 A100 的 2 倍。我们的模型表明,通过一些关键的编译器优化,我们可以达到 A100 性能的 3 倍以上。在占用网络上,我们看到了更大的飞跃。量产硬件上接近 3 倍,而且还有进一步提升的空间。那么,这对 Tesla 意味着什么?按照目前的编译器性能水平,我们只需一个 Dojo 瓦片,就能取代 1、2、3、4、5 和 6 个 GPU 机箱的机器学习计算能力。而且这个 Dojo 瓦片的成本低于其中 1 个 GPU 机箱。它真正意味着,以前训练耗时超过 1 个月的网络,现在不到 1 周就能完成。可惜,当我们进行测量时,结果并不太好。
第 22 段
在 PyTorch 层面,我们一开始并没有看到预期的性能,而这张时间线图展示了我们的问题。那些极小、极细的绿色条,是在加速器上运行的编译代码。这一行大部分都是空白,硬件只是在等待数据。凭借我们的密集型机器学习计算,Dojo 主机的机器学习算力实际上是 GPU 主机的 10 倍。数据加载器运行在这一台主机上,根本无法跟上所有这些机器学习硬件。因此,为了解决数据加载器的可扩展性问题,我们知道必须突破这台单一主机的限制。Tesla 传输协议可以在主机、瓦片和输入处理器之间无缝传输数据。因此,我们扩展了 Tesla 传输协议,使其能够通过以太网工作。然后,我们打造了 Dojo 网络接口卡 D-NIC,以便通过以太网利用 TTP。这使任何配备 D-NIC 卡的主机都能对其他 TTP 端点执行 DMA2,并从这些端点执行 DMA。所以我们先从 Dojo 网格开始,然后增加了一个配备 D-NIC 卡的数据加载主机层。我们通过一台以太网交换机将这些主机连接到网格。现在,这个数据加载层中的每台主机都能通过硬件加速 DMA 访问 Dojo 网格中的所有 TTP 端点。这些优化实施后,我们的利用率从 4% 上升到了 97%。因此,数据加载部分大幅缩短,机器学习硬件得以持续忙碌。实际上,我们预计这个数字很快会达到 100%。这些改动实施后,我们从 PyTorch 层看到了全部预期的加速效果,并恢复了正常运转。所以,我们从一种突破传统集成边界的硬件设计起步,以实现我们打造单一巨型加速器的愿景。我们已经看到编译器层和输入层如何构建在该硬件之上。在这些复杂的现实世界网络上验证了我们的性能后,我们就知道首次大规模部署会以什么为目标:具有高算术强度的自动标注网络。如今,它占用了 4,000 个 GPU,分布在 72 个 GPU 机架中。凭借我们的密集型计算机和高性能,我们预计只用 4 个 Dojo 机柜就能提供同等吞吐量。这 4 个 Dojo 机柜将成为我们计划在 2023 年第一季度建造的首个 ExaPOD 的一部分。这将使 Tesla 的自动标注能力提高到两倍以上。首个 ExaPOD 是我们计划在帕洛阿尔托、就在这堵墙对面建造的总计 7 个 ExaPOD 之一。我们还展示了其中一个 ExaPOD 的机柜,供大家观看。6 个瓦片密集地装在一个托盘上,54 千万亿次浮点运算/秒的算力,640 吉字节高带宽内存,电源和主机被击败。大量的算力。我们正在构建所有集群组件的新版本,并持续改进软件,以达到新的规模极限。我们相信,下一代硬件还能带来另外 10 倍的提升。为了实现我们的宏伟目标,我们需要最优秀的软件和硬件工程师。所以请来和我们聊聊,或者访问 tesla.com。好了,希望这些细节已经足够了。现在我们可以进入提问环节。各位,我想团队可以上台了。我们真的想展示 Tesla 在人工智能、计算硬件、机器人执行器方面的深度和广度,并努力真正改变外界对公司的看法,因为很多人认为我们就只是一家汽车公司,或者说我们造很酷的汽车,随便怎么说。但大多数人完全不知道,Tesla 可以说是现实世界人工智能硬件和软件领域的领导者,而且我们正在构建可以说是自 Kray-1 超级计算机以来最激进的计算机架构。我认为,如果你有兴趣开发世界上一些最先进、并将真正对世界产生积极影响的技术,Tesla 就是该来的地方。那么,来吧,开始提问。我想前面有一支麦克风,后面也有一支麦克风。直接把麦克风扔给大家。大家都扑向麦克风。是的,你好,非常感谢。我在这里深受震撼。我对 Optimus 印象非常深刻,但我想知道为什么没有驱动手部。为什么你们为手部选择了肌腱驱动方案?
第 23 段
因为肌腱不是很耐用。还有,为什么采用弹簧加载?很好,太棒了,是的,这是个很好的问题。你知道,说到任何类型的驱动方案,都要在不同选择之间作出权衡,比如是采用肌腱驱动系统,还是采用某种连杆式系统。把麦克风靠近嘴一点。再近一点,能听见我吗?很好。是的,我们采用肌腱式系统的主要原因是,你知道,首先我们其实研究过一些合成肌腱,但我们发现,金属船用缆绳要坚固得多。这些缆绳的一个优势是非常有利于减少零部件。我们确实想制造很多这样的手,所以存在一大堆零件、一大堆小连杆,会在大量制造某种东西时成为问题。你知道,从某种意义上说,肌腱优于连杆的一大原因是,它可以实现消除回程间隙。消除回程间隙实质上可以让你的手指不出现任何空隙,也不会出现断断续续的动作。至于弹簧加载,它主要让我们能够实现主动张开。因此,我们不必使用 2 个执行器来驱动手指闭合再张开,而是可以用肌腱驱动手指闭合,然后由弹簧被动伸展。这也是我们自己的手上能看到的机制,对吧?我们能够主动屈曲,也能够伸展。是的。我的意思是,我们对 Optimus 的目标,是尽快让它成为一款最大程度有用的机器人。解决人形机器人的各种问题有很多种方法,而且我们采用的所有技术方案可能并不全都走在正确的方向上。我应该说,我们愿意随着时间推移不断演进你们在这里看到的技术方案,它们并非一成不变。但我们必须选定某种方案,而且我们希望选择一种能让我们尽快生产机器人,并像我说的那样,让它尽快发挥作用的方案。我们正努力遵循一条目标:以最快路径打造出能够大批量制造的实用机器人。我们会在 Tesla 内部、在我们的工厂里测试机器人,然后看看它究竟有多大用处。因为你必须,你得在现实中形成闭环,以确认机器人事实上是有用的。是的,所以我们会直接用它来制造东西。我们相信,凭借目前设计的手部可以做到这一点。但我确信还会有手部第 2 版、第 3 版,而且随着时间推移,我们可能会相当大幅度地改变架构。你好,Optimus 机器人确实令人印象深刻,你们做得很棒,双足机器人真的很难。但我注意到,你们的计划中可能遗漏了一点,那就是承认人类精神的效用。我想知道,Optimus 将来是否会拥有个性,并且能在叠我们的衣服时听我们的笑话而发笑。是的,当然。我认为我们希望推出真正有趣的 Optimus 版本。这样 Optimus 既可以具备实用性并执行任务,也可以有点像朋友和伙伴,陪你一起消磨时间。我相信人们会为这款机器人想出各种富有创意的用途。你知道,问题在于,一旦解决了核心智能和执行器,你实际上就可以,嗯,给机器人穿上各种服装吧。我的意思是,你可以改变机器人的外观,可以用许多不同方式给机器人加上外皮。我相信人们会找到非常有趣的方式来,嗯,打造 Optimus 的不同版本。感谢精彩的演示。我想知道 Optimus 中是否有与干预等效的机制。在人类对正在发生的事情持不同意见的时刻进行标注,似乎非常重要;而在人形机器人中,这或许也是一种值得获取的信息来源。是的,我认为我们会有远程操作机器人并在它做坏事时进行干预的方法,尤其是在我们训练机器人并让它逐步运转起来时。希望我们可以,你知道,以一种方式设计它,让我们能够在机器人将要撞到什么东西时把它停下来。我们只要像这样抓住它,它就会停下,不会像,你知道,压碎你的手之类的。而这些全都是干预数据。是的,我们也可以从模拟系统中学到很多东西。在那里,我们可以检查碰撞,并监督系统将那些行为视为不良行为。是的,我的意思是,所以 Optimus,随着时间推移,我们希望它成为,你知道,一个仿生人,就是你在科幻电影中见过的那种仿生人,比如《星际迷航:下一代》里的 Data。但显然,我们可以对机器人进行编程,使它不那么像机器人而更加友好。而且,你知道,它显然可以学习模仿人类,感觉非常自然。所以,随着人工智能整体上的进步,我们可以把这些能力加入机器人。而且,你知道,它显然应该能够执行简单指令,甚至凭直觉判断你想要什么。你可以给它一个高层级指令,然后它可以将其分解成一系列动作,并执行这些动作。你好,是的,想到你们认为凭借 Optimus 可以让经济产出实现多个数量级的提升,确实令人兴奋。真的非常令人兴奋。Tesla 创立时,其使命是加速可再生能源或可持续交通的到来。那么,随着 Optimus 的出现,你是否仍然认为这会继续作为 Tesla 的使命宣言,还是会把它更新成,比如,你知道,加速我也不知道,无限富足或无限经济的到来?是的,严格来说,Optimus 并不严格地与加速可持续能源直接一致。如果它完成任务的效率比人更高,那么我想,它确实有助于可持续能源。但我认为,随着 Optimus 的出现,这项使命实际上确实有所拓宽,变成,你知道,我也不知道,让未来变得精彩。所以,你知道,我想,当你看到 Optimus,而且我了解你,但我很期待看到 Optimus 会变成什么样。而且,你知道,这就像,你知道,如果可以的话,我的意思是,你可以问任何一项特定技术:你想看看它在 1 年、2 年、3 年、4 年、5 年、10 年后会是什么样吗?我要说,Optimus 肯定是你绝对想看看会发展成什么样的技术。相比之下,你知道,其他很多技术已经,你知道,算是进入平台期了。关于在这里说名字,但是,你知道,所以,我认为 Optimus 在大约 5 年、10 年后会变得不可思议,令人震撼。我非常有兴趣看到它成为现实,希望你们也是。我这里有个简短的问题,Justin,我想知道,你们是否计划扩展机器人的对话能力?
第 24 段
我对此的第2个完整问题是,最终目标是什么?Optimus 的最终目标是什么?是的,Optimus 肯定会具备对话能力,所以,你将能够和它说话并进行对话,而且感觉会相当自然。所以,从最终目标的角度看,我不知道,我认为它会不断演进,我不确定它最终会走向何方,但肯定会到达某个有趣的地方。而且,你知道,我们始终必须小心,你知道,不要走上终结者的道路。那是一个,你知道,我本来想,也许我们应该以一段视频开场,比如终结者一开始就这样,你知道,碾碎头骨。但那可能会,你知道,人们可能不会太当真。所以,你知道,我们确实希望 Optimus 是安全的,因此我们正在设计安全防护措施,让你可以在本地停止机器人。而且,你知道,基本上会有一个无法通过互联网更新的本地控制 ROM。坦率地说,我认为这相当重要,是必不可少的。所以,比如一个无法被更改的本地停止按钮或遥控器之类的东西。不过,我的意思是,它肯定会很有趣,不会无聊。好的,是的,我看到你们今天有一款非常有吸引力的产品,包括 Dojo 及其应用。所以,我想知道 Dojo 平台的未来是什么?所以,你知道,比如像 AWS 那样提供基础设施和服务,还是会像 NVIDIA 那样销售芯片?所以,基本上说,未来是什么?因为我看你们使用 7nm,所以开发成本很容易就超过1000万美元。你们如何从商业角度开展这项业务?Dojo 是一台非常大的计算机,实际上会消耗大量电力,并且需要大量冷却。所以,我认为让 Dojo 以类似 Amazon Web Services 的方式运营,可能比尝试把它卖给别人更合理。所以,运营 Dojo 最有效的方式,就是让它成为一项你可以使用的服务。它可以在线使用,你可以在那里以快得多的速度和更低的成本训练模型。而随着世界向软件 2.0 过渡,这也在宾果卡上。我认识的某个人必须知道要喝5杯龙舌兰酒。所以,让我们看看,软件 2.0 将使用大量神经网络训练。所以,随着时间推移,出现越来越多的神经网络相关事物,人们会想要使用速度最快、成本最低的神经网络训练系统,这算是合乎情理的。所以,我认为这个方向有很多机会。嗨,我叫 Ali Jahanian。感谢举办这次活动,它非常鼓舞人心。我的问题是,我想知道你对能够理解我们的情感和艺术,并能为我们的创造力作出贡献的人形机器人有何愿景?嗯,我认为你已经看到,机器人至少能够生成非常有趣的艺术作品,比如 Dali 和 Dali 2。而且我认为,我们将开始看到 AI 甚至能够生成具有连贯性的电影,比如有趣的电影,还能讲笑话。所以,除了 Tesla 之外,许多公司的 AI 进步速度之快都相当惊人。我们正迈向一个非常有趣的未来。是的,那么,你们有人想对此发表评论吗?
第 25 段
是的,我想 Optimus 机器人可以创作实体艺术,而不只是数字艺术。你可以通过文字或语音要求一些舞蹈动作,然后将来你就可以创作出这些东西。所以,会有很多实体艺术,而不只是数字艺术。哦,是的,计算机绝对可以创作实体艺术,是的,100%。是的,比如跳舞、踢足球,或者任何你……我的意思是,随着时间推移,它肯定需要变得更加敏捷。非常感谢这次演示。现在,关于 Tesla Autopilot 的幻灯片,我注意到你们使用的模型在很大程度上受到语言模型的启发。我想知道这背后的历史,以及它带来了多大程度的改进。我觉得用语言模型来处理车道转换是一个非常有趣、令人好奇的选择。所以,我们转向语言建模的原因大致有2个方面。所以,第1个……大声点,靠近些。好的,明白了。是的,所以语言模型通过2种方式帮助我们。第1种方式是,它让我们能够预测原本无法预测的车道。正如 Ashok 之前提到的,基本上,当我们以某种密集 3D 的方式预测车道时,你只能对某些类型的车道进行建模,但我们希望获得交叉路口内部那些纵横交错的连接。如果不把它做成图预测,就根本无法做到。如果你试图用密集分割来实现,它就是行不通。此外,车道预测是一个多模态问题。有时你就是没有足够的视觉信息,无法准确知道交叉路口另一侧是什么样子。所以你需要一种能够泛化并生成连贯预测的方法。你不希望同时预测2条车道和3条车道。你希望确定其中一种,而像这样的语言模型所提供的通用模型能做到这一点。嗨。嗨,我叫 Giovanni。是的,感谢这次演示。真的很棒。我有一个问题想问 FSD 团队。对于神经网络,你们如何测试……如何对它进行单元测试、软件单元测试?你们是有一大批,还是我不知道,几千个,或者……
第 26 段
是的,有一些案例,神经网络训练完成后,你们必须让它通过这些案例,之后才能将它作为产品发布,对吧?是的,基本上你们对此采用什么软件单元测试策略?是的,很高兴你问这个问题。我们定义了一系列测试,首先是针对软件本身的单元测试。但之后对于神经网络模型,我们定义了 VAP 集,你可以定义……我们发现,如果你只有一个大型测试集,那是不够的。对于不同的故障模式,我们需要复杂的 VAP 集。然后我们对它们进行整理,并在产品存在期间不断扩充它们。所以多年以来,我们积累了数十万个过去曾失败的示例。我们对这些示例进行了整理,因此对于任何新模型,我们都会根据这些失败案例的全部历史记录进行测试。然后继续向这个测试集添加内容。除此之外,我们还有影子模式,会把这些模型以静默方式部署到车上。然后我们会收到关于它们在哪里失败或成功的数据。而且还有一个广泛的 QA 项目。出现回归时很难发布。在产品到达客户手中之前,大约有9层过滤机制。但我们有非常好的基础设施,能让这一切高效完成。我是其中一名 QA 测试人员,所以我对车辆进行 QA……是的,QA 测试人员。是的,所以我经常待在车里,就是不断排队测试任何不会彻底崩溃的最新 alpha 版本。是的,发现了很多错误。嗨,很棒的活动。
第 27 段
我有一个关于自动驾驶基础模型的问题。我们都看到,大模型确实能够……当你用数据和模型参数从 GP3 扩展到 POM 时,它现在实际上可以进行推理。你是否认为,有必要用数据和规模来扩展基础模型,然后至少可以得到一个可能解决所有问题的教师模型,再将其蒸馏为学生模型?你是这样看待基础模型与自动驾驶的相关性吗?这与我们的自动标注模型非常相似。所以,我们不只有在车上运行的模型。我们还会训练完全离线、规模极大、无法在车上实时运行的模型。所以,我们只在服务器上离线运行这些模型,生成非常好的标签,随后用来训练在线网络。所以,这是这些师生模型的一种蒸馏形式。就基础模型而言,我们正在构建一些非常、非常大的数据集,规模达到数 PB。而且我们看到,当拥有这些大型数据集时,其中一些任务的表现非常好。比如我提到的运动学,输入视频,输出所有物体的全部运动学数据,一直到4阶导数。人们曾认为我们无法用摄像头进行检测。检测、深度、速度、加速度。想象一下,要让这些高阶导数准确,它们必须有多精确。而这一切都来自这类大型数据集和大型模型。所以,我们正以自己的方式,在几何、运动学以及诸如此类的领域看到基础模型的对应形式。John,你有什么要补充的吗?是的,我会简短说一下。基本上,每当我们在更大的数据集上训练时,都会看到模型性能大幅提升。而且基本上,每当我们通过其他一些辅助任务的预训练步骤来初始化网络时,我们基本上都会看到改进。使用大型数据集的自监督或监督方法都有很大帮助。嗨,所以在一开始,埃隆说 Tesla 可能有兴趣构建通用人工智能系统。鉴于这类技术可能带来的变革性影响,专门投资于技术性 AGI 安全专业能力似乎是审慎之举。我知道 Tesla 开展了大量关于技术性窄域 AI 安全的研究。我想知道 Tesla 是否打算尝试专门建立技术性通用人工智能安全方面的专业能力。嗯,我的意思是,如果我们开始看起来将对通用人工智能作出重大贡献,那么我们肯定会投资于安全,我非常信奉 AI 安全。我认为政府层面应该设立某种 AI 监管机构,就像任何影响公共安全的事物都有监管机构一样。所以,我们有针对飞机和汽车以及某种食品和药品的监管机构,因为它们会影响公共安全,而 AI 也会影响公共安全。所以我认为,而这其实是政府目前还不理解的事情,我认为应该有一个裁判,努力确保 AGI 的公共安全。你可以想一想,创造 AGI 所必需的要素有哪些?比如,可获取的数据集极其重要。而如果你有大量汽车和人形机器人,处理来自现实世界的数 PB 视频数据和音频数据,就像人类一样,那可能是最大的数据集,很可能就是最大的数据集。因为除此之外,你显然还可以逐步扫描互联网。但互联网无法真正做到的,是在现实世界中拥有数百万或数亿台摄像头。正如我所说,还包括音频和其他传感器。所以,我认为我们可能会拥有最多的数据,也可能拥有最多的训练算力。因此,我们可能会对 AGI 作出贡献。嘿,我注意到 Semi 就在后面,但我们还没有过多谈到它。我只是想知道,对于 Semi 卡车,从感知角度来看,你们正在考虑哪些变化?我想,它的要求显然与普通汽车有很大不同。而如果你认为事实并非如此,为什么会是这样?
第 28 段
不,我想,基本上你可以驾驶汽车。想想看,是什么在驾驶任何车辆?是一个带有眼睛的生物神经网络,本质上就是带有摄像头。你的主要传感器是什么?慢速万向架上的2个摄像头,一个非常慢的万向架。那就是你的头。所以,如果一个在慢速万向架上装有2个摄像头的生物神经网络能够驾驶半挂卡车,那么,如果你有大约8个摄像头,拥有连续的360度视野,以更高的帧率运行,而且反应速度快得多,那么我认为很明显,你应该能够把半挂卡车或任何车辆开得比人类好得多。嗨,我叫 Akshay,谢谢你们举办这场活动。假设 Optimus 会用于不同的使用场景,并且会针对这些使用场景以不同速度演进,是否有可能以某种方式独立开发和部署不同的软件及硬件组件,并将它们部署到 Optimus 中,从而加快 Optimus 整体功能的开发?好吧,我们没理解。很遗憾,我们的神经网络没有理解这个问题。下一个问题。嗨,我想把话题切换到 Autopilot。那么你们计划什么时候向美国和加拿大以外的国家推出 FSD 测试版?另外,我的下一个问题是,你认为当前 Autopilot 技术栈中最大的瓶颈、技术或障碍是什么?你们打算如何解决它,让 Autopilot 在安全保障和人类信心等性能指标方面显著优于人类?我想你还提到,对于 FSD V11,你们将把高速公路和城市道路合并为一个统一的技术栈,并进行一些重大的架构改进,能否稍微详细讲讲?谢谢。嗯,这是一大堆问题。我们希望能够……我认为从技术角度看,到今年年底,应该可以在全球推出 FSD 测试版。但在许多国家,我们需要获得监管批准,所以在其他国家,我们在一定程度上受制于监管审批。但我认为,从技术角度看,到今年年底,它将可以面向全球进行测试。我们预计下个月会发布一个相当大的改进,它将始终尤其擅长评估快速移动的横向交通流的速度,还有许多其他方面。那么,有人想详细说说吗?我来说吧,过去量产版 Autopilot 与完全自动驾驶测试版之间有许多差异,但随着时间推移,这些差异变得越来越小。我想就在几个月前,我们现在已经在所有车辆的 FSD 和量产版 Autopilot 中使用相同的纯视觉目标检测技术栈。仍然存在少数差异,主要差异是我们目前预测车道的方式。正如我在演讲中提到的,我们升级了车道建模,使它能够处理这些更复杂的几何结构。量产版 Autopilot 仍然使用一种更简单的车道模型,但我们正在扩展当前的 FSD 测试版模型,使其也能在各种高速公路场景中工作。我驾驶的 FSD 测试版实际上已经采用了集成技术栈。因此,它在城市街道和高速公路上都使用 FSD 技术栈,对我来说运行得相当好。但我们需要在各种天气条件下验证它,比如暴雨、降雪、沙尘,并确保它在广泛的环境中都比量产技术栈表现得更好。不过我们已经非常接近了。我想是……我不知道,也许……肯定会在今年年底之前,也许是11月。是的,在我们个人的驾驶中,FSD 技术栈在高速公路上的表现已经远远优于我们现有的量产技术栈。我们还预计在今年年底之前把停车场技术栈纳入 FSD 技术栈。因此,基本上,到今年年底,这将使我们达到这样的程度:你坐进停车场里的汽车,然后一直开到停车场尽头的一个停车位。就需要优化的根本指标而言,就是每两次必要干预之间能行驶多少英里。也就是大幅增加汽车能够完全自主行驶多少英里,之后才需要进行一次关乎安全的干预。所以,是的,这就是我们每周衡量的根本指标,而我们正在这个指标上取得巨大的改进。嗨,谢谢,非常感谢你们的演示,非常鼓舞人心。我叫 Daisy,其实我有一个非技术问题想问你。我很好奇,如果你回到20多岁时,你希望自己当时知道哪些事情?你会给年轻时的自己什么建议?嗯,我正试着想出一些有用的话。是的,一个联合 Tesla 会是一件事。是的,我想,就是尽量让自己接触尽可能多的聪明人。我不怎么读很多书。你知道,不过我确实那么做过。所以,我认为不必总是过于紧绷,以及更享受当下,也有一些可取之处,我会这样对20多岁的自己说。偶尔停下来闻闻玫瑰花香,可能会是个好主意。你知道,就像我们在 Quageline 环礁研发 Falcon 1 火箭时,我们在一座美丽的小岛上研发火箭,而在整个那段时间里,我一次都没有在海滩上喝过饮料。我就想,我本来应该在海滩上喝杯饮料,那完全没问题。非常感谢。我想,你已经通过 Optimus 让所有机器人领域的人都兴奋起来了。这感觉非常像10年前的自动驾驶。但事实证明,自动驾驶比10年前看起来的要困难。那么,我们现在知道哪些10年前不知道的事情,可以让例如人形机器人上的 AGI 更快到来?嗯,我是说,在我看来,AGI 正在非常迅速地进步。几乎每周都会有某项重大消息公布。而且,是的,我是说,到这个时候,AI 似乎几乎能在任何基于规则的游戏中获胜。它能够创作极其令人印象深刻的艺术作品,进行非常复杂的对话,你知道,还能写文章。而这些能力还在不断提升。有这么多更有才华的人在研究 AI,硬件也在变得更好。无论我们在 Tesla 做什么,AI 都处于一条超级……一条强劲的指数级改进曲线上。显然,我们会在一定程度上受益于 AI 的这条指数级改进曲线。比如,Tesla 也必须非常擅长执行器、电机、变速箱、控制器、电力电子、蓄电池和传感器。而且,你知道,确实,我会说,4轮机器人与有手臂和腿的机器人之间最大的区别,就是把执行器做好。这是执行器和传感器的问题。显然,还有如何控制这些执行器和传感器。但它是……是的,执行器和传感器,以及如何控制执行器。我不知道,我们必须具备创造一款有吸引力的机器人所必需的……各种要素。而我们正在做这件事,所以……
第 29 段
嗨,伊兰。你确实正在把人类带到下一个层次。确切地说,Tesla,而你正在把人类带到下一个层次。所以,你说擎天柱,Optimus 将用于下一座 Tesla 工厂。我的问题是,一座新的 Tesla 工厂会完全由 Optimus 程序运行吗?普通公众什么时候能订购一个人形机器人?是的,我想会……你知道,我们将让 Optimus 从工厂中非常简单的任务开始。你知道,比如可能只是装载一个零件,就像你在视频中看到的那样。你知道,把一个零件从一个地方搬到另一个地方,或者把一个零件装入我们某个更传统的机器人工作单元中,你知道,它会把车身焊接起来。所以我们会从……你知道,只是尝试解决:怎样才能让它具备任何用处?然后逐步扩大它能派上用场的情形数量。我认为 Optimus 能派上用场的情形数量将呈指数级增长。真的、真的非常快。至于人们什么时候能够订购一个,我不知道,我认为不会太遥远。嗯,我想你的意思是,人们什么时候能够收到一个?所以,我不知道,我会说可能在3年内,而且不会超过5年。在3至5年内,你可能就能收到一个 Optimus。我觉得,实现 AGI 进步的最佳方式,是让世界各地尽可能多的聪明人参与进来。鉴于 Tesla 与机器人公司相比所具备的规模和资源,也鉴于当前人形机器人研究的状况,类似 Tesla 这样的公司以某种方式将部分仿真硬件部件开源,是否合理?我认为 Tesla 仍然可以成为占主导地位的平台方,它可以成为类似 Android OS,或者面向整个人形机器人研究领域的类似 iOS 的东西。与其把 Optimus 仅仅留给 Tesla 的研究人员或工厂本身,是否可以将它开放出来,让全世界探索人形机器人研究?
第 30 段
我认为,我们必须谨慎防范 Optimus 可能被用于不好的方式,因为这是可以做的事情之一。所以我认为,我们会提供一种 Optimus,你可以向 Optimus 下达指令,但这些指令受一些你无法违背的机器人法则约束。因此,不伤害他人,而且我认为 Optimus 可能还会有相当多与安全相关的事项。我们可能再回答几个问题,然后感谢大家前来。2个问题,1个深入的问题和1个宽泛的问题。关于 Optimus 的深入问题:当前以及理想的控制器带宽是多少?然后是更宽泛的问题:这充分展示了这家公司的深度和广度。Tesla 究竟有什么独特之处,使这一切成为可能?有人想回答带宽问题吗?所以,技术带宽……靠近你的嘴,大声点。关于带宽问题,你必须理解或弄清楚你希望它执行什么任务。如果你对该任务进行频率变换,你希望自己的肢体做什么?你的带宽就是从那里得出的。它不是一个你可以直接明确说出的数字,你需要理解自己的使用场景,而带宽就是从那里得出的。宽泛的问题是什么?
第 31 段
关于广度和深度这个问题,我可以回答广度和深度。关于带宽问题,我认为我们最终可能只会提高带宽。这会转化为机器人的实际灵巧度和反应时间。可以肯定地说,它不是1赫兹,也许你不需要一路提高到100赫兹。但也许是10、25,我不知道。随着时间推移,我认为带宽会大幅提高。或者把它转化为灵巧度和延迟。随着时间推移,你会希望将其降至最低。最大限度降低延迟,最大限度提高灵巧度。就广度和深度而言,我想我们现在是一家相当大的公司了。所以我们拥有许多不同领域的专业能力,而这些能力是我们为了制造电动汽车、继而制造自动驾驶电动汽车而必然需要培养的。Tesla基本上就像一整套初创公司。而到目前为止,它们几乎全都相当成功。所以我们一定是做对了某些事情。而我认为,经营公司时,我的核心职责之一就是营造一个能让杰出工程师蓬勃发展的环境。我认为在许多公司,我不知道,也许是大多数公司,如果某人是一位真正有才华、有干劲的工程师,他们实际上无法……他们的才华在许多公司会受到压制。而在有些公司,工程人才受到压制的方式也许并没有明显的坏处,而是那里实在太舒适了,你拿到的钱又实在太多了,你实际上需要产出的成果又实在太少了,以至于它就像一个甜蜜陷阱。所以硅谷有几个甜蜜陷阱一样的地方。它们对工程师而言看起来不一定像是糟糕的地方。但你必须问,一位优秀工程师进去了,最后做出了什么?那些工程人才的产出似乎非常低,尽管他们看起来过得很开心。这就是为什么我说硅谷有几家甜蜜陷阱公司。Tesla不是甜蜜陷阱,我们要求很高,就像是:你会完成一大堆该死的事情,而且会非常酷。这不会轻松。但如果你是一位才华超群的工程师,我认为你的才华在这里会得到比其他任何地方更充分的运用。你知道,SpaceX也是这样。嗨,Ilan,我有2个问题。两个都是问自动辅助驾驶团队的。事情是这样的,过去几年我一直在关注你们的进展。所以今天,你们对车道检测之类的东西做了改动。就像你说的,你们以前做的是即时语义分割。现在你们构建了用于构建车道的迁移模型。那么你们目前还面临哪些其他常见挑战,也就是你们未来要解决的挑战?作为一名充满好奇心的工程师,我想了解这些,这样我们研究人员也能着手研究这些问题。开始研究这些问题。第2个问题是,我对数据引擎非常好奇。你们讲过一个案例,比如汽车停了下来。那么,你们如何从现有数据中找到与之非常相似的案例?如果能多讲一点数据引擎,那就太好了。我用占用网络作为例子来回答第1个问题。你们在演示中看到的东西,1年前还不存在。所以我们只花了1年的时间。实际上,我们交付了超过12个占用网络。而要用1个基础模型真正表示周围各处的整个物理世界,并且始终以此为条件,实际上非常非常具有挑战性。所以就在1年多以前,我们还算是在一个2D世界中驾驶。如果有一堵墙,如果有一条曲线,我们差不多会用同样的静态边缘来表示,显然,你知道,这并不理想,对吧?驾驶时,曲线和墙壁之间存在很大差别,你会作出不同的选择,对吧?所以在意识到必须转向3D之后,我们基本上必须重新思考整个问题,并考虑如何处理它。所以这算是我们在过去1年中遇到、并且已经攻克的挑战的1个例子。是的,回答一下我们实际上如何获取那些棘手的停车案例样本。对此有几种做法,但这里有2个例子。第1,我们可以针对信号内部的不一致设置触发器。比方说,停车状态位在“已停车”和“行驶中”之间闪烁。我们会触发并把它传回来。第2,我们可以更多地利用影子模式逻辑。所以,如果客户忽略了那辆车,但我们认为应该为它停车,我们也会把那些数据取回来。所以这些就是各种不同的触发逻辑,使我们能够取回那些数据活动。嗨。感谢你们精彩的演示,非常感谢。有很多公司都在关注AGI问题。而这个问题如此困难的原因之一,是这个问题本身就很难定义。几家公司有几种不同的定义,它们关注不同的事情。那么Tesla是什么,Tesla如何定义AGI问题,你们具体关注什么?嗯,我们实际上并没有专门关注AGI。我只是说,AGI似乎很可能会成为我们正在做的事情的一种涌现特性。因为我们正在创造老式自动驾驶汽车和自主运行的人形机器人,它们实际上伴随着一股真正巨大的数据流,这些数据不断涌入并得到处理。这远远是规模最大的现实世界数据,而且是那种仅靠搜索互联网无法获得的数据。因为你必须置身于现实世界,与人互动、与道路互动。而且,你知道,地球是一个很大的地方,现实既混乱又复杂。所以我认为,这有点像是,它似乎很可能会成为一种涌现特性。如果你拥有数千万或数亿辆自动驾驶车辆,也许还有数量相当的人形机器人。也许人形机器人方面的数量还会更多。嗯,那就是规模最大的数据,而如果这些视频正在得到处理,汽车显然就很可能会远比人类驾驶员优秀。而人形机器人可能会变得越来越难以与人类区分。这样一来,就像我说的,你就会拥有这种AGI的涌现特性。并且可以说,人类作为一个整体,在某种程度上也是一种超级智能。尤其是随着我们提高人类之间的数据传输速率。回想早期,那件事似乎发生在很久以前,互联网就像是人类获得了一个神经系统。现在突然之间,人类中的任何一个个体都可以通过连接互联网了解人类的全部知识。几乎所有知识,或者肯定是其中很大一部分。而以前我们通过渗透来交换信息。比如为了传输数据,你必须写一封信。必须有人亲自把信带给另一个人。中间还会经历一大堆事情,然后就是……是啊,我是说,仔细想想,那慢得不可思议。而且,即使你身在美国国会图书馆,你仍然无法获取全世界的所有信息。当然你也无法搜索这些信息,而且显然只有极少数人在美国国会图书馆。所以我的意思是,互联网是伟大的某种平等要素之一,就信息和知识获取而言,它一直是历史上最大的平等化力量。我认为任何研究历史的人都会同意这一点。因为你知道,回到1000年前,书籍寥寥无几。而且书会贵得不可思议,却只有少数人识字。甚至只有更少的人拥有书。现在看看,你可以即时访问任何书籍。基本上可以免费学习任何东西。这非常不可思议。所以,你知道,最近有人问我最愿意身处哪个历史时期。我的回答就是现在。这是历史上最有趣的时代,而我读过很多历史。所以让我们尽最大努力让它延续下去。回到前面的1个问题,我想问的是,随着时间推移,在Tesla自动辅助驾驶方面发生的事情,是神经网络逐渐吸收了越来越多的软件。当然,在极限情况下,你可以直接获取汽车看到的视频,并将其与方向盘和踏板的操控输入进行比较,这些输入非常简单。原则上,你可以在中间什么都没有的情况下进行训练。因为人类使用生物神经网络时就是这样做的。你可以根据视频进行训练,而训练视频的是方向盘和踏板的动作,中间没有其他软件。我们还没有达到那一步,但正在逐渐朝那个方向发展。好了,最后1个问题。你们进行得怎么样?我想前排这里有1个问题。你好,他们就在那里。我们回答2个问题,好吧。他们在这里。感谢如此精彩的演示。我们最后回答你的问题。好的,酷。FSD被如此多人使用,你们如何从性能统计数据的角度评估公司的风险承受能力?你是否认为,需要提高透明度,或由第三方加强监管,来确定什么程度算足够好,并为许多英里行驶里程中的性能设定阈值?Tesla的首要设计要求是安全。而且这一要求适用于所有方面。所以就汽车的机械安全性而言,在政府测试过的所有汽车中,我们的受伤概率最低。仅就被动机械安全而言,基本上就是碰撞结构、安全气囊之类的东西。我们的主动安全评级也是最高的。而且我认为,主动安全会达到好得离谱的程度。就是好得远超人类,荒谬地好。然后关于自动辅助驾驶,我们确实大体上公布了行驶里程统计数据,包括没有任何自动驾驶能力的汽车、没有任何自动驾驶能力的Tesla汽车、硬件1、硬件2、硬件3,以及参与FSD测试版的汽车。我们看到整个过程中都在稳步改善。有时会出现这样一种二分问题:你是否应该等到汽车比人安全3倍之后,才部署任何技术?但我认为,那实际上在道德上是错误的。当你相信增加自动驾驶能力能够减少伤害和死亡时,我认为你在道德上有义务部署它。尽管很多人会起诉你并责怪你。因为那些被你挽救生命的人不知道自己的生命得救了。而那些偶尔死亡或受伤的人肯定知道,或者他们的州知道,自动辅助驾驶出了问题。这就是为什么你必须看总行驶里程中的数字。发生了多少起事故,多少起是严重事故,有多少人死亡。而我们在道路上行驶的汽车已经远超300万辆。所以每天的行驶里程非常多。它不会完美。但重要的是,它非常明确地比不部署更安全。是的。所以,我想,最后1个问题。我想,是的,谢谢。这里是最后1个问题。好的,嗨。那么,我并不从事硬件工作。所以也许硬件团队和你们可以指点我一下。为什么Optimus的设计必须保持对称?因为我们人类有惯用手,对吧?我们使用某一组肌肉的频率高于其他肌肉。随着时间推移会产生磨损,对吧?所以,也许随着时间推移,你会开始更多地看到一些关节故障,或者一些执行器故障。我明白这还处于极其早期的阶段。另外,我们人类已经把如此多的幻想和虚构建立在超人能力之上。比如我们所有人都不想直接走到那边。我们想伸长手臂,而且就像,我们有所有这些,你知道,很多幻想、奇幻的设计。所以,考虑到在电池和计算强度方面正在发生的其他一切,也许你们可以利用所有这些方面,想出一些,嗯,我不知道,就你们正在制造的机器人而言更有意思的东西。我希望你们能够探索那些方向。是的,我觉得如果能像,你知道,让神探加杰特成为现实,那会很酷。那会非常棒。所以,是的,我的意思是,现在我们只是想让基础的人形机器人良好运作。我们的目标是通过这条通往实用型人形机器人的道路。我认为这会让我们脚踏实地,字面意义上的。并确保我们正在做有用的事情。比如最难做到的事情之一就是有用。真正做到有用,然后让曲线下面积所代表的效用很高。比如你平均为每个人提供了多少帮助,然后你帮助了多少人?总效用。试图真正向大量的人交付人们喜欢的有用产品,难得简直不可思议。令人难以置信。你知道,所以我可以说,天哪,一家已经交付产品的公司和一家还没有交付产品的公司之间,差别大得要命。这完全是天壤之别。而且即使你交付了产品,你能否让产出的价值高于投入的成本?这同样难得不可思议,尤其是硬件。所以,不过我认为随着时间推移,做一些有创意的事情会很酷。比如有8条手臂之类的。并且有不同版本。也许,你知道,会有一些硬件,比如有些公司能够给一个Optimus添加东西。比如也许我们,你知道,增加一个电源端口之类的。或者把它们连接起来,你可以给你的Optimus添加附件。就像你可以把它们添加到你的手机上一样。随着时间推移,可以做很多很酷的事情。而且也许可以形成一个由小公司或者大公司组成的生态系统,为Optimus制造附加组件。那么,说到这里,我想感谢团队的辛勤工作。你们太棒了。也感谢大家前来。对于所有在线上观看的人,感谢收看。我认为这会成为那种很棒的视频,你可以,比如,快进到你觉得最有意思的部分。但我们试图为你提供海量的细节,确确实实是为了让你可以在方便的时候观看视频。你可以专注于自己觉得有意思的部分,并跳过其他部分。所以,谢谢大家,我们会这样做,争取每年都做一次。我们甚至可能会做每月播客。所以,不过我认为,算是带着你们一路同行会很棒。并且向你们展示正在发生哪些很酷的事情。是的,谢谢。好了,谢谢。谢谢。
Paragraph 1
All right, welcome everybody give everyone a moment to Get back in the audience and All right great welcome to Tesla AI day 2022 We've got some really exciting things to show you I think you'll be pretty impressed I do want to set some expectations with respect to our Optimus robot as As you know last year was just a person in a robot suit But we've now we've come a long way and that's I think we you know compared to that it's gonna be very impressive and We're gonna talk about The advancements in AI for full self-driving as well as how they apply to more generally to real-world AI problems Like a humanoid robot and and even going beyond that I think there's some potential that what we're doing here at Tesla could make a meaningful contribution to AGI and And I think actually Tesla is a good Antity to do it from a governance standpoint because we're a publicly traded company with one class of stock and That means that the public controls Tesla and I think that's actually a good thing So if I go crazy you can fire me. This is important Maybe I'm not crazy. I don't know so Yeah, so we're gonna talk a lot about our progress in AI autopilot as well as progress in with dojo and Then we're gonna bring the team out and to do a long Q&A so you can ask tough questions But whatever you'd like existential questions technical questions, but we want to have As much time for Q&A as possible. So let's see you with that That's because Hey guys, I'm Milana work on autopilot and it is about and I'm Lizzie Mechanical engineer on the project as well. Okay So should we should we bring up the bot before we do that? We have one One little bonus tip for the day.
Paragraph 2
This is actually the first time we try this robot without any backup support Cranes mechanical mechanisms. No cables. Nothing. Yeah I want to do it with you guys tonight. That is the first time. Let's see.
Paragraph 3
You ready? Let's go I I I think the bug got some boobs This is essentially the simple self-driving computer that runs in your Tesla cars by the way This is the this is literally the first time the robot has operated without a tether was on stage tonight So the robot can actually do a lot more than we just showed you we just didn't want it to fall on its face So we'll we'll show you some videos now of the robot doing a bunch of other things Yeah, which are less risky. Yeah, we should close the screen guys. Yeah Yeah, we wanted to show a little bit more what we've done over the past few months with the bot and just walking around and dancing on stage Just humble beginnings, but you can see the autopilot neural networks running as it's just retrained for the bot directly on that on that new platform That's my watering can yeah when you when you see a rendered view. That's that's the robot. What's the that's the world the robot sees So it's it's very clearly identifying objects like this is the object.
Paragraph 4
It should pick up picking it up Yeah We use the same process as we did for autopilot to connect data and train neural networks that we didn't deploy on the robot That's an example that illustrates the upper body a little bit more Something that will like try to nail down in a few months over the next few months, I would say To perfection, but this is really an actual station in the Fremont factory as well that it's working at And that's not the only thing we have to show today, right? Yeah, absolutely. So That what you saw was what we call bumble sea. That's our sort of rough development robot using semi off-the-shelf actuators But we actually have gone a step further than that already the team's done an incredible job And we actually have an optimist bot with fully tesla designed and built actuators Um battery pack control system everything. Um, it it wasn't quite ready to walk But I think it will walk in a few weeks But we wanted to show you the robot The something that's actually fairly close to what we'll go into production And and show you all all the things that can do so let's bring it up All right Yeah So here you're seeing optimists with these the With the degrees of freedom that we expect to have in optimists production unit one Which is the ability to move all the fingers independently move the To have the thumb have two degrees of freedom. So it has opposable thumbs And both left and right hand so it's able to operate tools and do useful things.
Paragraph 5
Our goal is to make a useful humanoid robot as quickly as possible and We've also designed it using the same discipline that we use in designing the car, which is to say to design it for All manufacturing such that it's possible to make the robot at in high volume at low cost with high reliability So that that's incredibly important. I mean, you've all seen very impressive humanoid robot demonstrations And that that's great. But what are they missing? They're missing a brain that they don't have the the intelligence to navigate the world by themselves And they're they're also very expensive Um and made in low volume. Um, whereas, uh, this this is the optimist's design to be an extremely capable robot But made in in very high volume probably ultimately millions of units And it is expected to cost much less than a car So, uh, I would say probably less than 20,000 dollars would be my guess Okay The the potential for optimist is I think appreciated by very few people As usual Tesla demos are coming in hot So Um, yeah, uh, I'm the team's put on put in and the team has put in an incredible amount of work Uh working days, you know seven days a week Running the 3am oil To to get to the demonstration today. Um, super proud of what they've done is they've really done a great job I'd just like to give a hand to the whole optimist team So, you know that now there's still a lot of work to be done to, uh refine optimists and Improve it.
Paragraph 6
Obviously, this is just optimist version one And that's really why we're holding this event Which is to convince some of the most talented people in the world like you guys um to Join tesla and help make it a reality and bring it to fruition at scale Such that it can help millions of people And the and the potential likes it is is really boggles the mind because you have to say like what what is an economy an economy is uh sort of productive entities times the productivity, uh capita times Productivity per capita at the point at which there is not a limitation on capita The it's not clear what an economy even means at that point. It an economy becomes quasi infinite um so What what you know taken to fruition in the hopefully benign scenario the This means a future of abundance a future where um There is no poverty where people you can have whatever you want In terms of products and services It really is a a fundamental transformation of civilization as we know it Obviously we want to make sure that transformation is a positive one and um safe And but but that's also why I think Tesla as an entity doing this being a single class of stock publicly traded owned by the public um is very important Um and should not be overlooked. I think this is essential because then if the public doesn't like what tesla's doing The public can buy shares in tesla and vote differently This is a big deal Um Like it's very important that that I can't just do what I want You know, sometimes people think that that but it's not true. Um, so You know that it's it's very important that the the corporate entity that has that that makes this happen Is something that the public can properly influence And so I think the tesla structure is is is ideal for that Um And like said that you know self-driving cars will certainly have a Tremendous impact on the world. Um, I think they will improve the productivity of transport by at least A half order of magnitude perhaps an order of magnitude perhaps more um Optimus I think has Maybe a two order of magnitude Uh potential improvement in uh economic output Like like it's it's not clear. It's not clear what the limit actually even is um so But we we need to do this in the right way we need to do it carefully and safely And ensure that the outcome is one that is beneficial to Uh civilization and and one that humanity wants Uh, I can't this is also extremely important obviously so, um And and I hope you will consider uh joining tesla to achieve those goals Um It tesla we're we really care about doing the right thing here or aspire to do the right thing and and really not Pay the road to hell with with good intentions And I think the road is road to hell is mostly paved with bad intentions, but every now and again There's a good intention in there.
Paragraph 7
So we want to do the right thing. Um, so, you know consider joining us and helping make it happen um With that let's let's uh, we want to the next phase All right, so you've seen a couple robots today. Let's do a quick timeline recap So last year we unveiled the tesla bot concept, but a concept doesn't get us very far We knew we needed a real development and integration platform to get real life learnings as quickly as possible So that robot that came out and did the little routine for you guys We had that within six months built working on software integration hardware upgrades over the months since then But in parallel we've also been designing the next generation this one over here So this guy is rooted in the the foundation of sort of the vehicle design process, you know We're leveraging all of those learnings that we already have Obviously there's a lot that's changed since last year, but there's a few things that are still the same you'll notice We still have this really detailed focus on the true human form We think that matters for a few reasons, but it's fun. We spend a lot of time thinking about how amazing the human body is We have this incredible range of motion typically really amazing strength Um a fun exercise is if you put your fingertip on the chair in front of you you'll notice that there's a huge Range of motion that you have in your shoulder and your elbow for example without moving your fingertip you can move those joints all over the place Um, but the robot, you know, its main function is to do real useful work And it maybe doesn't necessarily need all of those degrees of freedom right away So we've stripped it down to a minimum sort of 28 fundamental degrees of freedom and then of course our hands in addition to that Humans are also pretty efficient at some things and not so efficient in other times So for example, we can eat a small amount of food to sustain ourselves for several hours. That's great Uh, but when we're just kind of sitting around no offense, but we're kind of inefficient. We're just sort of burning energy So on the robot platform, what we're going to do is we're going to minimize that idle power consumption drop it as low as possible And that way we can just flip a switch and immediately the robot turns into something that does useful work So let's talk about this latest generation in some detail, shall we?
Paragraph 8
So on the screen here, you'll see in orange are actuators, which we'll get to in a little bit and in blue are electrical system So now that we have our sort of human based research and we have our first development platform We have both research and execution to draw from for this design Again, we're using that vehicle design foundation. So we're taking it from concept through design and analysis and then build and validation Along the way, we're going to optimize for things like cost and efficiency because those are critical metrics to take this product to scale eventually How are we going to do that? Well, we're going to reduce our part count and our power consumption of every element possible We're going to do things like reduce the sensing in the wiring at our extremities You can imagine a lot of mass in your hands and feet is going to be quite difficult and power consumptive to move around And we're going to centralize both our power distribution and our compute to the physical center of the platform So in the middle of our torso, actually it is the torso. We have our battery pack This is sized at 2.3 kilowatt hours, which is perfect for about a full day's worth of work What's really unique about this battery pack is it has all of the battery electronics integrated into a single pcb within the pack So that means everything from sensing to fusing Charge management and power distribution is all on one all in one place We're also leveraging both our vehicle products and our energy products To roll all of those key features into this battery. So that's streamlined manufacturing Really efficient and simple cooling methods battery management and also safety And of course we can leverage tesla's existing infrastructure and supply chain to make it So going on to sort of our brain it's not in the head, but it's pretty close Also in our torso, we have our central computer So as you know tesla already ships full self-driving computers in every vehicle we produce We want to leverage both the autopilot hardware and the software for the humanoid platform But because it's different in requirements and informed factor, we're going to change a few things first So we still are gonna it's going to do everything that a human brain does Processing vision data making split sescan decisions based on multiple sensory inputs and also communications So to support communications, it's equipped with wireless connectivity as well as audio support And then it also has hardware level security features which are important to protect both the robot and the people around the robot So now that we have our sort of core we're going to need some limbs on the sky Um, and we'd love to show you a little bit about our actuators and our fully functional hands as well But the first before we do that, I'd like to introduce Malcolm who's going to speak a little bit about our structural foundation for the robot Tesla have the capabilities to analyze highly complex systems Don't get much more complex than a crash You can see here a simulated crash from bottle three Superimposed on top of the actual physical crash It's actually incredible how um, how accurate it is Just to give you an idea of the complexity of this model It includes every not bolt and washer every spot weld and it has 35 million degrees of freedom quite amazing And it's true to say that if we didn't have models like this, we wouldn't be able to make the safest cars in the world So can we utilize our capabilities and our methods from the automotive side to influence a robot? Well, we can make a model and since we had crash software we're using the same software here We can make it fall down The purpose of this is to make sure that if it falls down ideally it doesn't but it's superficial damage We don't want it to for example break its gearbox and its arms.
Paragraph 9
That's equivalent of a dislocated shoulder of a robot Difficult and expensive to fix So we wanted to dust itself off get on with the job. It's being given We could also take the same model and we can drive the actuators using the inputs from a previously solved model Bringing it to life So this is producing the motions for the tasks we want the robot to do these tasks are picking up boxes turning squatting walking upstairs Whatever the set of tasks are we can play to the model. This is showing just simple walking We can create the stresses in all the components that helps us optimize the components These are not dancing robots these are actually the modal behavior the first five modes of the robot And typically when people make robots they make sure the first mode is up around the top single figures up towards 10 hertz Who is to do this is to make the controls of walking easier. It's very difficult to walk if you can't guarantee where your foot is wobbling around That's okay if you make one robot. We want to make thousands maybe millions We haven't got the luxury of making from carbon fiber titanium. We want to make them plastic things are not quite as stiff So we can't have these high targets.
Paragraph 10
I call them dumb targets We've got to make them work at lower targets So is that it's that good to work? Well, if you think about it, sorry about this, but we're just bags of soggy jelly and bones thrown in We're not high frequency. If I start on my leg, I don't vibrate at 10 hertz We people operate at a lot of frequency. So we know the robot actually can it just makes controls harder So we take the information from this the modal Data and the stiffness and feed it into the control system that allows it to walk And Just changing tax lightly looking at the knee We could take some inspiration from biology and we can look to see what the mechanical advantage of the knee is It turns out it actually represent quite similar to four-bar link and that's quite non-linear That's not surprising really because if you think when you bend your leg down The torque on your knee is much more when it's bent than it is when it's straight So you'd expect a non-linear function and in fact the biology is non-linear. This matches it quite accurately So that's a representation the four-bar link is obviously not physically four-bar link as I said the characteristics are similar, but Me bending down that's not very scientific. Let's be a bit more scientific We've played all the tasks through the through this graph And this is showing picketing is walking squatting the tasks I said we did on the stress And that's the the torque Seen at the knee against the knee bend on the horizontal axis This is showing the requirement for the need to do all these tasks And then put a curve through it surfing over the top of the piece and that's saying this is what's required to make the robot Do these tasks?
Paragraph 11
So if we look at the four-bar link that's actually the green curve And it's saying that the non-linearity of the four-bar link is actually linearized The characteristic of the force what that really says is that's lower the force That's what makes the actuator have the lowest possible force, which is the most efficient. We want to burn energy up slowly What's the blue curve with the blue curve is actually if we didn't have a four-bar link We just had an arm sticking out of my leg here with a with an actuator on it a simple two-bar link That's the best we could do with a simple two-bar link and it shows that that would create a much more force in the actuator Which would not be efficient So what does it look like in practice? well As you'll see but it's very tightly packaged in the knee you'll see it go transparent on the second You'll see the four-bar link there is operating on the actuator. This is determined the force and the displacements on the actuator And now pass you over to Constantina to tell you a lot more detail about how these actuators are made and designed optimized. Thank you So I am I would like to talk to you about The design process and the actuator portfolio In our robot So there are many similarities between a car and the robot when it comes to powertrain design The most important thing that matters here is energy mass and cost We are carrying over most of our designing experience from the car to the robot So in the particular case you see a car with two drive units And the drive units are used in order to accelerate the car zero to 60 miles per hour time or drive a city drive site while The robot that has 28 actuators and It's not obvious. What are the tasks at actuator level?
Paragraph 12
So we have tasks that are higher level like walking or climbing stairs or carrying a heavy object Which need to be translated into joint Into joint specs therefore we use our model That generates The torque speed Trajectories for our joints which subsequently is going to be fed in our optimization model And to run through the optimization process This is one of the scenarios that the robot is capable of doing which is turning and walking So when we have this torque speed trajectory we lay it over an efficiency map of an actuator And we are able along the trajectory to generate The power consumption and the energy cumulative energy for the task versus time So this allows us to define the system cost for the particular actuator and put a simple point into the cloud Then we do this for hundreds of thousands of actuators by solving in our cluster And the red line denotes the Pareto front, which is the preferred area where we will look for optimal So the x denotes the preferred actuator design we have picked for this particular joint So now we need to do this for every joint. We have 28 joints to optimize and we parse our cloud We parse our cloud again for every joint spec and the red axis this time denotes the bespoke actuator designs for every joint The problem here Is that we have too many unique actuator designs and even if we take advantage of the symmetry Still there are too many in order to make something Mass manufacturable we need to be able to reduce the amount of unique actuator designs Therefore we run something called commonality study, which we parse our cloud again Looking this time for actuators that simultaneously meet the joint performance requirements for more than one joint at the same time So the resulting portfolio is six actuators and they show in a color map at the middle figure um And the actuators can be also viewed in this Slide we have three rotary and three linear actuators all of which have a great Output force or torque per mass The rotary actuator in particular has a mechanical clutch integrated On the high speed side angular contact ball bearing and on the high speed side And on the low speed side a cross roller bearing and the year train is a strain wave year Um, there are three integrated sensors here and the bespoke permanent magnet machine The linear actuator I'm sorry The linear actuator has planetary rollers and an inverted planetary Screw as a gear train which allows efficiency and compaction and durability So in order to demonstrate the force capability of our linear actuators, we have set up an experiment in order to test it under its limits And I will let you enjoy the video So our actuator is able to lift A half ton nine foot concert grand piano And This is a requirement it's not something nice to have Because our muscles can do the same when they are direct driven when they are directly driven our quadriceps muscles Can do the same thing it's just that the knee is an upgearing Linked system that converts the force into velocity at the end effector of our heels for purposes of giving To the human body agility So this is one of the main things that are amazing about the human body And I'm concluding my part at this point and I would like to welcome my colleague Mike who's going to talk to you about Hand design. Thank you very much Thanks for seeing us So we just saw how powerful a human and a humanoid actuator can be However, humans are also incredibly dexterous The human hand has the ability to move at 300 degrees per second There's tens of thousands of tactile sensors And it has the ability to grasp and manipulate almost every object in our daily lives For our robotic hand design, we were inspired by biology We have five fingers an opposable thumb Our fingers are driven by metallic tendons that are both flexible and strong We have the ability to complete wide aperture power grasps while also being optimized for precision gripping of small thin and delicate objects So why a human like robotic hand? Well, the main reason is that our factories in the world around us is designed to be ergonomic So what that means is that it ensures that objects in our factory are graspable But it also ensures that new objects that we may have never seen before can be grasped by the human hand And by our robotic hand as well The converse there is is pretty interesting because it's saying that these objects are designed to our hand instead of having to make changes To our hand to accompany a new object Some basic stats about our hand is that it has six actuators and 11 degrees of freedom It has an in-hand controller which drives the fingers and receives sensor feedback Sensor feedback is really important to learn a little bit more about the objects that we're grasping And also for proprioception and that's the ability for us to recognize where our hand is in space One of the important aspects of our hand is that it's adaptive This adaptability is involved essentially as complex mechanisms that allow the hand to adapt the objects that's being grasped Another important part is that we have a non-back drivable finger drive This clutching mechanism allows us to hold and transport objects without having to turn on the hand motors You just heard how we went about going we went about designing the tesla bot hardware Now I'll hand it off to Milan and our autonomy team to bring this robot to life Thanks Michael All right So all those cool things we've shown earlier in the video Were possible just in a matter of a few months. Thanks to the amazing work that we've done autopilot over the past few years Most of those components poured it quite easily over to the bot's environment If you think about it, we're just moving from a robot on wheels to a robot on legs So some of the components are pretty similar and some other require more heavy lifting So for example our computer vision neural networks Were ported directly from autopilot to the bot's situation It's exactly the same occupancy network that we'll talk into a little bit more details later with the autopilot team that is now running on the bot here in this video The only thing that changed really is the training data that we had to recollect We're also trying to find ways to improve those occupancy networks Using work made on your radiance fields to get really great volumetric Rendering of the bot's environments for example here some machinery that the bot might have to interact with Another interesting problem to think about is in indoor environments, mostly with that sense of gps signal How do you get the bot to navigate to its destination? Say for instance to find its nearest charging station So we've been training More neural networks to identify high-frequency features key points within the bot's camera streams And track them across frames over time as the bot navigates with its environment And we're using those points to get a better estimate of the bot's pose and trajectory within its environment as it's walking We also did quite some work on the simulation side and this is literally the autopilot simulator To which we've integrated the robots locomotion code and this is a video of the Motion control code running in the autopilot simulator Showing the evolution of the robot's work over time.
Paragraph 13
So as you can see we started quite slowly in April and start accelerating as we unlock more joints And deeper more advanced techniques like arms balancing over the past few months And so locomotion is specifically one component that's very different as we're moving from the car to the bot's environment And so I think it warrants a little bit more depth and I'd like my colleagues to start talking about this now Thank you Milan. Hi, everyone. I'm Felix. I'm a robotics engineer on the project and I'm going to talk about walking Walking seems easy, right? People do it every day. You don't even have to think about it But there are some aspects of walking which are challenging from engineering to technology And I think that's one of the things that makes it so much easier for me to think about it But there are some aspects of walking which are challenging from engineering perspective.
Paragraph 14
For example Physical self-awareness that means having a good representation of yourself What is the length of your limbs? What is the mass of your limbs? What is the size of your feet? All that matters Also having an energy efficient gate. You can imagine there's different styles of walking and all of them are equally efficient Most important keep balance. Don't fall And of course also coordinate the motion of all of your limbs together So now humans do all of this naturally But as engineers or roboticists we have to think about these problems And the following I'm going to show you how we address them in our locomotion planning and control stack So we start with locomotion planning And our representation of the bot that means a model of the robots kinematics dynamics and the contact properties And using that model and the desired path for the bots our locomotion planner generates reference trajectories for the entire system This means feasible trajectories with respect to the assumptions of our model The planner currently works in three stages.
Paragraph 15
It starts planning footsteps and ends with the entire motion photo system And let's dive a little bit deeper in how this works So in this video we see footsteps being planned over a planning horizon following the desired path And we start from this and add then Foot trajectories that connect these footsteps using toe-off and heel strike just as the humans Just as humans do and this gives us the largest right and less knee bend for high efficiency of the system The last stage is then finding a center of mass trajectory Which gives us a dynamically feasible motion of the entire system to keep balance As we all know plans are good, but we also have to realize them in reality. Let's say how see how we can do this Thank you Felix. Hello everyone. My name is Anand and I'm going to talk to you about controls So let's take the motion plan that Felix just talked about and put it in the real world on a real robot Let's see what happens It takes a couple steps and falls down Well, that's a little disappointing But we are missing a few key pieces here, which will make it walk Now as Felix mentioned the motion planner is using an idealized version of itself and a version of reality around it This is not exactly correct It also expresses its intention Through trajectories and wrenches wrenches of forces and torques that it wants to exert on the world to locomotive Reality is way more complex than any similar model. Also the robot is not simplified It's got vibrations and modes, compliance, sensor noise and on and on and on So what does that do to the real world when you put the bot in the real world? Well, the unexpected forces cause unmodeled dynamics, which essentially the planet doesn't know about and that causes destabilization Especially for a system that is dynamically stable like bipedal locomotion So what can we do about it?
Paragraph 16
Well, we measure reality We use sensors and our understanding of the world to do state estimation And here you can see the attitude and pelvis pose, which is essentially the vestibular system in a human Along with the center of mass trajectory being tracked when the robot is walking in the office environment Now we have all the pieces we need in order to close the loop So we use our better bot model We use the understanding of reality that we've gained through state estimation And we compare what we want versus what we expect the reality expect that reality is doing to us in order to Add corrections to the behavior of the robot Here the robot certainly doesn't appreciate being poked, but it has an admirable job of staying upright The final point here is a robot that walks is not enough We need it to use its hands and arms to be useful. Let's talk about manipulation Hi everyone, my name is Eric robotics engineer on tesla bot And I want to talk about how we've made the robot manipulate things in the real world We wanted to manipulate objects while looking as natural as possible and also get there quickly So what we've done is we've broken this process down into two steps First is generating a library of natural motion references Or we could call them demonstrations and then we've adapted these motion references online to the current real world situation So let's say we have a human demonstration of picking up an object We can get a motion capture of that demonstration, which is visualized right here as A bunch of keyframes representing the location of the hands the elbows the torso We can map that to the robot using inverse kinematics And if we collect a lot of these now we have a library that we can work with But a single demonstration is not generalizable to the variation in the real world For instance, this would only work for a box in a very particular Location So what we've also done is run these Reference trajectories through a trajectory optimization program which solves for where the hand should be how the robot should balance during When it needs to adapt the motion to the real world. So for instance, if the box is In this location, then our optimizer will create this trajectory instead Next Milan's going to talk about uh, what's next for the optimist uh, tesla lie. Thanks Right, so hopefully by now you guys got a good idea of what we've been up to over the past few months Um, we started having something that's usable, but it's far from being useful. There's still a long and exciting road ahead of us um, I think the first thing within the next few weeks is to Get optimists at least apart with bumble see the other bug prototype you saw earlier and probably beyond We are also going to start focusing on the real use case at one of our factories and really going to try to try to Nail this down and I run out all the elements needed to deploy this product in the real world I was mentioning earlier, you know indoor navigation Um graceful for management or even servicing all components needed to scale this product up But um, I don't know about you, but after seeing what we've shown tonight I'm pretty sure we can get this done within the next few months or years Um, and and make this product a reality and change the entire economy Um, so I would like to thank the entire optimist team for their hard work over the past few months I think it's pretty amazing. All of this was done in barely six or eight months.
Paragraph 17
Thank you very much Hey everyone Hi, I'm Ashok. I lead the autopilot team alongside Milan Oh god, it's going to be so hard to top that optinist section He'll try nonetheless anyway Every tesla that has been built over the last several years We think has the hardware to make the car drive itself We have been working on the software to add higher and higher levels of autonomy This time around last year. We are roughly 2000 cars driving our fsd beta software Since then we have significantly improved the software's robustness and capability That we have now shipped it to 160,000 customers as of today This did not come for free it came from the sweat and blood of the engineering team over the last one year Um, for example, we trained 75,000 neural network models just last one year That's roughly a model every eight minutes That's you know coming out of the team and then we evaluate them on our large clusters and then we ship 281 of those models That actually improved the performance of the car And this space of innovation is happening throughout the stack The the planning software the infrastructure the tools even hiring everything is progressing to the next level The fsd beta software is quite capable of driving the car It should be able to navigate from parking lot to parking lot handling city street driving stopping for traffic lights and stop signs Negotiating with objects at intersections making turns and so on All of this comes from the Uh camera streams that go through our neural networks that run on the car itself It's not coming back to the server or anything It runs on the car and produces all the outputs uh to form the world model or on the car and the planning software drives the car based on that Today we'll go into a lot of the components that make up the system The occupancy network acts as the base geometry layer of the system This is a multi-camera video neural network That from the images predicts the full physical occupancy of the world around the robot So anything that's physically present trees walls buildings Cars balls, whatever you it predicts if it's physically present it predicts them along with their future motion On top of this base level of geometry We have more semantic layers in order to navigate the roadways. We need the lanes, of course But then the roadways have lots of different lanes and they connect in all kinds of ways So it's actually a really difficult problem for typical computer vision techniques to predict the set of lanes and their Connectivities So we reached all the way into language technologies and then pull the state of the art from other Domains are not just computer vision to make this task possible For vehicles, we need their full kinematics state to control for them All of this directly comes from neural networks video streams raw video streams come into the networks Goes through a lot of processing and then outputs the full kinematics state that positions velocities acceleration jerk all of that Directly comes out of networks with minimal post processing. That's really fascinating to me because how how much does it take? Even possible what world do we live in that this magic is possible that these networks predicts fourth derivatives of these positions and people thought We couldn't even detect these objects My opinion is that it did not come for free It it required tons of data.
Paragraph 18
So we had to be sophisticated auto labeling systems that shone through raw sensor data Run a ton of offline compute on the servers. It took a lot of time. It took a lot of time. It took a lot of time Run a ton of offline compute on the servers. It can take a few hours run expensive neural networks Distill the information into labels that train our in-car neural networks On top of this we also use our simulation system to synthetically create images and since it's a simulation We trivially have all the labels All of this goes through a well oiled data engine pipeline where we first train a baseline model with some data Ship it to the car see what the failures are and once we know the failures We mind the fleet for the cases where it fails Provide the correct labels and add the data to the training set This process systematically fixes the issues and we do this for every task that runs in the car Yeah, and to train these new massive neural networks This year we expanded our training infrastructure by roughly 40 to 50 percent So that sits us at about 14,000 GPUs today across multiple training clusters in the United States We also worked on our AI compiler which now supports new operations needed by those neural networks And map them to the the best of our underlying hardware resources And our inference engine today is capable of distributing the execution of a single neural network across two independent system on chips Essentially two independent computers interconnected within the same full self-driving computer And to make this possible we have to keep a tight control on the end-to-end latency of this new system So we deployed more advanced scheduling code across the full FSD platform All of these neural networks running in the car Together produce the vector space, which is again the model of the world around the robot or the car And then the planning system operates on top of this coming up with trajectories that avoid collisions or smooth Make progress towards the destination using a combination of model-based optimization Plus neural network that helps optimize it to be really fast Today we are really excited to present progress on all of these areas We have the engineering leads standing by to come in and explain these various blocks and these power not just the car But the same components also run on the Optimus robot that Milan showed earlier With that I welcome Paril to start talking about the planning section Hi all, I'm Paril Jain Let's use this intersection scenario today Let's use this intersection scenario to dive straight into how we do the planning and decision making in autopilot So we are approaching this intersection from a side street and we have to yield to all the crossing vehicles Right with as they are about to enter the intersection The pedestrian on the other side of the intersection decides to cross the road without a crosswalk Now we need to yield to this pedestrian Yield to the vehicles from the right and also understand the relation between the pedestrian and the vehicle on the other side of the intersection So a lot of these intra object dependencies That we need to resolve in a quick glance And humans are really good at this We look at a scene understand all the possible interactions evaluate the most promising ones And generally end up choosing a reasonable one So let's look at a few of these interactions that autopilot system evaluated We could have gone in front of this pedestrian with a very aggressive longitudinal lateral profile Now obviously we are being a jerk to the pedestrian and we would spook the pedestrian and his cute pet We could have moved forward slowly Short for a gap between the pedestrian or end the vehicle from the right Again, we are being a jerk to the vehicle coming from the right But you should not outright reject this interaction in case this is only safe interaction available Lastly the interaction we ended up choosing Stay slow initially find the reasonable gap and then finish the maneuver after all the agents pass Now evaluation of all of these interactions is not trivial Especially when you care about modeling the higher order derivatives for other agents For example, what is the longitudinal jerk required by the vehicle coming from the right when you assert in front of it? Relying purely on collision checks with marginal predictions will only get you so far because you will miss out on a lot of valid interactions This basically boils down to solving a multi-agent joint trajectory planning problem over the trajectories of ego and all the other agents Now how much ever you optimize there's going to be a limit to how fast you can run this optimization problem It will be close to close to order of 10 milliseconds even after a lot of incremental approximations Now for a typical crowded unprotected lift Say you have more than 20 objects Each object having multiple different future modes the number of relevant interaction combinations will blow up The planner needs to make a decision every 50 milliseconds.
Paragraph 19
So how do we solve this in real time? We rely on a framework what we call as interaction search, which is basically a paralyzed research over a bunch of maneuver trajectories The state space here corresponds to the kinematic state of ego, the kinematic state of other agents, their nominal future multiple multi-modal predictions and all the static entities in the scene The action space is where things get interesting We use a set of maneuver trajectory candidates to branch over a bunch of interaction decisions and also incremental goals for a longer horizon maneuver Let's walk through this research very quickly to get a sense of how it works We start with a set of vision measurements namely lanes occupancy moving objects These get represented as past attractions as well as latent features We use this to create a set of goal candidates Lanes again from the lanes network or unstructured regions which correspond to a probability mask derived from human demonstrations Once we have a bunch of these goal candidates, we create three trajectories using a combination of classical optimization approaches As well as our network planner again trained on data from the customer fleet Now once we get a bunch of these three trajectories We use them to start branching on the interactions We find the most critical interaction In our case, this would be the interaction with respect to the pedestrian Whether we assert in front of it or yield to it Obviously the option on the left is a high penalty option, it likely won't get prioritized So we branch further onto the option on the right and that's where we bring in more and more complex interactions Building this optimization problem incrementally with more and more constraints And the tree search keeps flowing, branching on more interactions, branching on more goals Now a lot of pricks here lie in evaluation of each of this node of the tree search Inside each node, initially we started with creating trajectories using classical optimization approaches Where the constraints like I described would be added incrementally And this would take close to 1 to 5 milliseconds per action Now even though this is fairly good number, when you want to evaluate more than 100% interactions, this does not scale So we ended up building lightweight queryable networks that you can run in the loop of the planner These networks are trained on human demonstrations from the fleet as well as offline solvers with relaxed time limits With this, we were able to bring the run time down to close to 100 microseconds per action Now doing this alone is not enough because you still have this massive tree search that you need to go through And you need to efficiently prune the search space So you need to do a new scoring on each of these trajectories Few of these are fairly standard, you do a bunch of collision checks, you do a bunch of comfort analysis What is the jerk and access required for a given manure The customer fleet data plays an important role here again We run two sets of again lightweight queryable networks, both really augmenting each other One of them trained from interventions from the FSD beta fleet Which gives a score on how likely is a given manure to result in interventions over the next few seconds And second, which is purely on human demonstrations, human driven data, giving a score on how close is your given selected action to a human driven trajectory The scoring helps us prune the search space, keep branching further on the interactions and focus the compute on the most promising outcomes The cool part about this architecture is that it allows us to create a cool blend between data driven approaches where you don't have to rely on a lot of hand engineered costs But also ground it in reality with physics based checks Now a lot of what I described was with respect to the agents, we could observe in the scene But the same framework extends to all of the other systems that we have We use the video feed from 8 cameras to generate the 3D occupancy of the world The blue mask here corresponds to the visibility region, we call it It basically gets blocked at the first occlusion you see in the scene We consume this visibility mask to generate the visibility of the scene We use the video feed from 8 cameras to generate the 3D occupancy of the world The blue mask here corresponds to the visibility region, we call it In the first occlusion you see in the scene, we consume this visibility mask to generate what we call as ghost objects which you can see on the top left Now if you model the spawn regions and the state transitions of this ghost objects correctly If you tune your control response as a function of their existence likelihood, you can extract some really nice human-like behaviors Now I'll pass it on to Phil to describe more on how we generate these occupancy networks Hey guys, my name is Phil, I will share the details of the occupancy network we built over the past year This network is our solution to model the physical work in 3D around our cars And it is currently not shown in our customer-facing visualization What you will see here is the raw network output from our internal lab tool The occupancy network takes video streams of all our 8 cameras as input Produces a single unified volumetric occupancy in vector space directly For every 3D location around our car, it predicts the probability of that location being occupied or not Since it has video contacts, it is capable of predicting obstacles that are occluded instantaneously For each location, it also produces a set of semantics such as curb, car, pedestrian, and road debris as color-coded here Occupancy flow is also predicted for motion Since the model is a generalized network, it does not tell static and dynamic objects explicitly It is able to produce and model the random motion such as a swarming trainer here This network is currently running in all Teslas with FSD computers And it is incredibly efficient, runs about every 10 milliseconds with our neural-line accelerator So how does this work? Let's take a look at architecture First, we rectify each camera image with a camera calibration And the images we're showing here are given to the network It's actually not the typical 8-bit RGB image As you can see from the first image on top, we're giving the 12-bit raw photo-account image to the network Since it has 4 bits more information, it has 16 times better dynamic range as well as reduced latency Since we don't have to run ISP in the loop anymore We use a set of reglets and bif-fps as a backbone to extract image space features Next, we construct a set of 3D position queries along with the image space features as keys and values fit into an attention module The output of the attention module is high-dimensional spatial features These spatial features are aligned temporally using vehicle odometry to derive motion Next, these spatial temporal features go through a set of deconvolutions to produce the final occupancy and occupancy flow output They're formed as fixed-size voxel grids, which might not be precise enough for planning on control In order to get a higher resolution, we also produce per voxel feature maps which we feed into MLP with 3D spatial point queries to get position and semantics at any arbitrary location After knowing the model better, let's take a look at another example Here we have an articulated bus parked on the right side of the road, highlighted as an L-shaped voxel here As we approach, the bus starts to move. The front of the car turns blue first, indicating the model predicts The front of the bus has a long zero occupancy flow As the bus keeps moving, the entire bus turns blue, and you can also see that the network predicts the precise curvature of the bus This is a very complicated problem for a traditional object detection network, as you'll have to see whether I'm going to use one cuboid or perhaps two to feed the curvature But for an occupancy network, since all we care about is the occupancy in the visible space, we'll be able to model the curvature precisely Besides the voxel grid, the occupancy network also produces a drivel surface The drivel surface has both 3D geometry and semantics. They are very useful for control, especially on hilly and curvy roads The surface and the voxel grid are not predicted independently. Instead, the voxel grid actually aligns with the surface implicitly Here, we are at a hill quest where you can see the 3D geometry of the surface being predicted nicely Planner can use this information to decide perhaps we need to slow down more for the hill quest And as you can also see, the voxel grid aligns with the surface consistently Besides the voxels and the surface, we're also very excited about the recent breakthrough in Neural Radiance Field or NERF We're looking into both incorporating some of the last NERF features into occupancy network training as well as using our network output as the input state for NERF As a matter of fact, Ashok is very excited about this.
Paragraph 20
This has been his personal weekend project for a while About these NERFs, because I think the academia is building out of these foundation models for language using tons of large data sets for language But I think for vision, NERFs are going to provide the foundation models for computer vision because they are grounded in geometry And geometry gives us a nice way to supervise these networks and freezes off the requirement to define an ontology And the supervision is essentially free because you just have to differentially render these images So I think in the future, this occupancy network idea where images come in and then the network produces a consistent volumetric representation of the scene That can then be differentially rendered into any image that was observed I personally think it's a future of computer vision and we do some initial work on it right now But I think in the future, both at Tesla and in academia, we will see that this combination of one-shot prediction of volumetric occupancy will be the future That's my personal bet Thanks Ashok So here's an example early result of a 3D reconstruction from our free data Instead of focusing on getting perfect RGB reproduction in image space, our primary goal here is to accurately represent the world in 3D space for driving And we want to do this for all our free data over the world in all weather and lighting conditions And obviously this is a very challenging problem and we're looking for you guys to help Finally, the occupancy network is trained with large auto-labeled data sets without any human in the loop And with that, I'll pass to Tim to talk about what it takes to train this network Thanks Phil Alright, hey everyone Let's talk about some training infrastructure So we've seen a couple of videos, no four or five I think and care more and worry more about a lot more clips on that So we've been looking at the occupancy networks just from Phil Just Phil's videos, it takes 1.4 billion frames to train that network What you just saw and if you have 100,000 GPUs, it would take one hour But if you have one GPU, it would take 100,000 hours So that is not a humane time period that you can wait for your training job to run, right? We want to ship faster than that So that means you're going to need to go parallel So you need a more compute for that That means you're going to need a supercomputer So this is why we've built in-house three supercomputers comprising of 14,000 GPUs Where we use 10,000 GPUs for training and around 4,000 GPUs for auto-labeling All these videos are stored in 30 petabytes of a distributed managed video cache You shouldn't think of our data sets as fixed Let's say as you think of your image net or something, you know, with like a million frames You should think of it as a very fluid thing So we've got half a million of these videos flowing in and out of this cluster These clusters every single day And we track 400,000 of these kind of Python video instantiations every second So that's a lot of calls We're going to need to capture that in order to govern the retention policies of this distributed video cache So underlying all of this is a huge amount of infra, all of which we build and manage in-house So you cannot just buy, you know, 14,000 GPUs and then 30 petabytes of Flash NVMe And you just put it together and let's go train It actually takes a lot of work and I'm going to go into a little bit of that What you actually typically want to do is you want to take your accelerator So that could be the GPU or dojo, which we'll talk about later And because that's the most expensive component, that's where you want to put your bottleneck And so that means that every single part of your system is going to need to outperform this accelerator And so that is really complicated That means that your storage is going to need to have the size and the bandwidth to deliver all the data down into the nodes These nodes need to have the right amount of CPU and memory capabilities to feed into your machine learning framework This machine learning framework then needs to hand it off to your GPU and then you can start training But then you need to do so across hundreds or thousands of GPU in a reliable way in lockstep And in a way that's also fast, so you're also going to need an interconnect Extremely complicated We'll talk more about dojo in a second So first I want to take you through some optimizations that we've done on our cluster So we're getting in a lot of videos and video is very much unlike, let's say, training on images or text Which I think is very well established Video is quite literally a dimension more complicated And so that's why we needed to go end to end from the storage layer down to the accelerator Optimize every single piece of that Because we train on the photon count videos that come directly from our fleet We train on those directly, we do not post-process those at all The way it's just done is we seek exactly to the frames we select for our batch We load those in including the frames that they depend on, so these are your eye frames or your key frames We package those up, move them into shared memory, move them into a double bar from the GPU And then use the hardware decoder that's only accelerated to actually decode the video So we do that on the GPU natively, and this is all in a very nice PyTorch extension Doing so unlocked more than 30% training speed increase for the occupancy networks And freed up basically a whole CPU to do any other thing You cannot just do training with just videos, of course you need some kind of a ground truth And that is actually an interesting problem as well The objective for storing your ground truth is that you want to make sure you get to your ground truth That you need in the minimal amount of file system operations And load in the minimal size of what you need in order to optimize for aggregate cross cluster throughput Because you should see a compute cluster as one big device which has internally fixed constraints and thresholds So for this we rolled out a format that is native to us that's called small We use this for our ground truth, our feature cache and any inference outputs So a lot of tensors that are in there And so just a cartoon here, let's say this is your table that you want to store Then that's how that would look out if you rolled out on disk So what you do is you take anything you'd want to index on, so for example video timestamps You put those all in the header so that in your initial header read you know exactly where to go on disk Then if you have any tensors you're going to try to transpose the dimensions to put a different dimension last as the contiguous dimension And then also try different types of compression Then you check out which one was most optimal and then store that one This is actually a huge tip if you do feature caching Unintelligible output from the machine learning network Rotate around the dimensions a little bit, you can get up to 20% increase in efficiency of storage Then when you store that we also order the columns by size So that all your small columns and small values are together So that when you seek for a single value you're likely to overlap with a read on more values which you'll use later So that you don't need to do another file system operation So I could go on and on, I just went on, touched on two projects that we have internally This is actually part of a huge continuous effort to optimize the compute that we have in-house So accumulating and aggregating through all these optimizations We now train our occupancy networks twice as fast just because it's twice as efficient And now if we add in a bunch more compute and go parallel we can now train this in hours instead of days And with that I'd like to hand it off to the biggest user of compute, John Hi everybody, my name is John Emmons, I lead the autopilot vision team I'm going to cover two topics with you today, the first is how we predict lanes And the second is how we predict the future behavior of other agents on the road In the early days of autopilot we modeled the lane detection problem as an image space instant segmentation task Our network was super simple though, in fact it was only capable of predicting lanes from a few different kinds of geometries Specifically it would segment the ego lane, it could segment adjacent lanes, and then it had some special casing for forks and merges This simplistic modeling of the problem worked for highly structured roads like highways But today we're trying to build a system that's capable of much more complex maneuvers Specifically we want to make left and right turns at intersections where the road topology can be quite a bit more complex and diverse When we try to apply this simplistic modeling of the problem here, it just totally breaks down Taking a step back for a moment, what we're trying to do here is to predict the sparse set of lane instances and their connectivity And what we want to do is to have a neural network that basically predicts this graph where the nodes are the lane segments And the edges encode the connectivity between these lanes So what we have is our lane detection neural network, it's made up of three components In the first component we have a set of convolutional layers, attention layers, and other neural network layers That encode the video streams from our eight cameras on the vehicle and produce a rich visual representation We then enhance this visual representation with a coarse road level map data Which we encode with a set of additional neural network layers that we call the lane guidance module This map is not an HD map, but it provides a lot of useful hints about the topology of lanes inside of intersections, the lane counts on various roads, and a set of other attributes that help us The first two components here produce a dense tensor that sort of encodes the world But what we really want to do is to convert this dense tensor into a sparse set of lanes and their connectivity We approach this problem like an image captioning task where the input is this dense tensor and the output text is predicted into a special language that we developed at Tesla for encoding lanes and their connectivity In this language of lanes, the words and tokens are the lane positions in 3D space In the ordering of the tokens, encrypted modifiers in the tokens encode the connected relationships between these lanes By modeling the task as a language problem, we can capitalize on recent autoregressive architectures and techniques from the language community for handling the multiple-diality of the problem We're not just solving the computer vision problem at Autopilot, we're also applying the state-of-the-art in language modeling and machine learning more generally I'm now going to dive into a little bit more detail of this language component What I have depicted on the screen here is a satellite image which sort of represents the local area around the vehicle The set of nose and edges is what we refer to as the lane graph, and it's ultimately what we want to come out of this neural network We start with a blank slate We're going to want to make our first prediction here at this green dot This green dot's position is encoded as an index into a course grid which discretizes the 3D world Now we don't predict this index directly because it would be too computationally expensive to do so There's just too many grid points and predicting a categorical distribution over this has both implications at training time and test time So instead what we do is we discretize the world coarsely first, we predict the heat map over the possible locations, and then we latch in the most probable location Condition on this, we then refine the prediction and get the precise point Now we know where the position of this token is, but we don't know it's tight In this case though, it's a beginning of a new lane So we predict it as a start token And because it's a start token, there's no additional attributes in our language We then take the predictions from this first forward pass, and we encode them using a learned positional embedding Which produces a set of tensors that we combine together Which is actually the first word in our language of lanes We add this to the first position in our sentence here We then continue this process by predicting the next lane point in a similar fashion Now this lane point is not the beginning of a new lane, it's actually a continuation of the previous lane So it's a continuation token type Now it's not enough just to know that this lane is connected to the previously predicted lane We want to encode its precise geometry, which we do by regressing a set of spline coefficients We then take this lane, we encode it again, and add it as the next word in the sentence We continue predicting these continuation lanes until we get to the end of the prediction grid We then move on to a different lane segment So you can see that cyan dot there Now it's not topologically connected to that pink point It's actually forking off of that green point there So it's got a fork type And fork tokens actually point back to previous tokens from which their fork originates So you can see here the fork point predictor is actually the index zero So it's actually referencing back to a token that is already predicted, like you would in language We continue this process over and over again until we've enumerated all of the tokens in the lane graph And then the network predicts the end of sentence token Yeah, I just wanted to note that the reason we do this is not just because we want to build something complicated It almost feels like a Turing complete machine here with neural networks though Is that we try simple approaches, for example, trying to just segment the lanes along the road or something like that But then the problem is when there's uncertainty, say you cannot see the road clearly And there could be two lanes or three lanes and you can't tell A simple segmentation-based approach would just draw both of them It's kind of a 2.5 lane situation And the post-processing algorithm would hilariously fail when the predictions are such Yeah, the problems don't end there I mean, you need to predict these connective lanes inside of intersections Which is just not possible with the approach that Ashok's mentioning Which is why we had to upgrade to this sort of approach Yeah, when it overlaps like this, segmentation would just go haywire But even if you try very hard to put them on separate layers, it's just a really hard problem But language just offers a really nice framework for getting a sample from a posterior As opposed to trying to do all of this in post-processing But this doesn't actually stop for just autopilot, right? John, this can be used for optimists Yeah, I guess they wouldn't be called lanes But you could imagine, sort of in this stage here That you might have sort of paths that sort of encode the possible places that people could walk Yeah, basically if you're in a factory or in a home setting, you can just ask the robot Okay, please route to the kitchen or please route to some location in the factory And then we predict a set of pathways that would go through the aisles, take the robot And say, okay, this is how you get to the kitchen It just really gives us a nice framework to model these different paths That simplify the navigation problem for the downstream planner Alright, so ultimately what we get from this lane detection network Is a set of lanes in their connectivity, which comes directly from the network There's no additional step here for sparsifying these dense predictions into sparse ones This is just a direct unfiltered output of the network Okay, so I talked a little bit about lanes I'm going to briefly touch on how we model and predict the future paths and other semantics on objects So I'm just going to go really quickly through two examples The video on the right here, we've got a car that's actually running a red light and turning in front of us What we do to handle situations like this is we predict a set of short time horizon future trajectories on all objects We can use these to anticipate the dangerous situation here And apply whatever breaking and steering actions required to avoid a collision In the video on the right, there's two vehicles in front of us The one on the left lane is parked, apparently it's being loaded, unloaded I don't know why the driver decided to park there But the important thing is that our neural network predicted that it was stopped Which is the red color there The vehicle in the other lane, as you notice, also is stationary But that one's obviously just waiting for that red light to turn green So even though both objects are stationary and have zero velocity It's the semantics that is really important here So that we don't get stuck behind that awkwardly parked car Predicting all of these agent attributes presents some practical problems when trying to build a real-time system We need to maximize the frame rate of our object section stack So that autopilot can quickly react to the changing environment Every millisecond really matters here To minimize the inference latency, our neural network is split into two phases In the first phase, we identified the locations in 3D space where agents exist In the second stage, we then pull out tensors at those 3D locations Append it with additional data that's on the vehicle And then we do the rest of the processing This specification step allows the neural network to focus compute on the areas that matter most Which gives us superior performance for a fraction of the latency cost So, putting it all together The autopilot vision stack predicts more than just the geometry and kinematics of the world It also predicts a rich set of semantics, which enables safe and human-like driving I'm now going to hand things off to Sri who will tell us how we run all these cool neural networks on our FSD computer Thank you Hi everyone, I'm Sri Today I'm going to give a glimpse of what it takes to run these FSD networks in the car And how do we optimize for the inference latency? Today I'm going to focus just on the FSD lanes network that John just talked about So, when we started this track, we wanted to know if we can run this FSD lanes network natively on the trip engine Which is our in-house neural network accelerator that we built in the FSD computer When we built this hardware, we kept it simple and we made sure it can do one thing ridiculously fast Dense dot products But this architecture is autoregressive and iterative Where it crunches through multiple attention-attention blocks in the inner loop Producing sparse points directly at every step So, the challenge here was how can we do this sparse point prediction and sparse computation on a dense dot product engine Let's see how we did this on the trip So, the network predicts the heat map of most probable spatial locations of the point To do this on trip, we actually built a lookup table in SRAM And we engineered the dimensions of this embedding such that we could achieve all of this thing with just matrix multiplication Not just that, we also wanted to store this embedding into a token cache So that we don't recompute this for every iteration, rather reuse it for future point prediction Again, we put some tricks here where we did all these operations just on the dot product engine It's actually cool that our team found creative ways to map all these operations on the trip engine In ways that were not even imagined when this hardware was designed But that's not the only thing we had to do to make this work We actually implemented a whole lot of operations and features to make this model compilable To improve the intate accuracy as well as to optimize performance All of these things helped us run this 75 million parameter model just under 10 millisecond of latency Consuming just 8 watts of power But this is not the only architecture running in the car There are so many other architectures, modules and networks we need to run in the car To give a sense of scale, there are about a billion parameters of all the networks combined Producing around 1000 neural network signals So we need to make sure we optimize them jointly and such that we maximize the compute utilization Throughput and minimize the latency So we built a compiler just for neural networks that shares the structure to traditional compilers As you can see, it takes the massive graph of neural nets with 150k nodes and 375k connection Takes this thing, partitions them into independent subgraphs And compiles each of those subgraphs natively for the inference devices Then we have a neural network linker which shares the structure to traditional linker Where we perform this link time optimization There we solve an offline optimization problem with compute memory and memory band with constraints So that it comes with an optimized schedule that gets executed in the car On the runtime, we designed a hybrid scheduling system which basically does heterogeneous scheduling on one SOC And distributed scheduling across both the SOCs to run these networks in a model parallel fashion To get 100 tops of compute utilization, we need to optimize across all the layers of software Right from tuning the network architecture, the compiler, all the way to implementing a low latency high bandwidth RDMA link Across both the SOCs and in fact going even deeper to understanding and optimizing the cache coherent and non-coherent data path of the accelerator in the SOC This is a lot of optimization at every level in order to make sure we get the highest frame rate and as every millisecond counts here And this is just the visualization of the neural networks that are running in the car This is our digital brain essentially As you can see these operations are nothing but just the matrix multiplication, convolution to name a few real operations running in the car To train this network with a billion parameters, you need a lot of labeled data So Egan is going to talk about how do we achieve this with the auto labeling pipeline Thank you Sri Hi everyone, I'm Egan Zhang and I'm leading a geometric vision at autopilot So yeah, let's talk about auto labeling So we have several kinds of auto labeling frameworks to support various types of networks But today I'd like to focus on the awesome lanes net here So to successfully train and generalize this network to everywhere, we think we went tens of millions of trips from probably one million intersection or even more Than how to do that So it is certainly achievable to source sufficient amount of trips because we already have, as Tim explained earlier, we already have like 500,000 trips per day cache rate However, converting all those data into a training form is a very challenging technical problem To solve this challenge, we've tried various ways of manual and auto labeling So from the first column to the second, from the second to the third, each advance provided us nearly 100x improvement in throughput But still, we run an even better auto labeling machine that can provide us good quality, diversity and scalability To meet all these requirements, despite the huge amount of engineering effort required here, we've developed a new auto labeling machine powered by multi-trip reconstruction So this can replace 5 million hours of manual labeling with just 12 hours on cluster for labeling 10,000 trips So how we solved? There are three big steps. The first step is high precision trajectory and structural recovery by multi-camera, visual, inertial, or geometry So here, all the features including ground surface are inferred from videos by neural networks, then tracked and reconstructed in the vector space So the typical trip rate of this trajectory in car is like 1.3 centimeter per meter and 0.45 milliliter per meter, which is pretty decent considering its compact compute requirement Then the recovered surface and road details are also used as a strong guidance for the later manual verification stuff This is also enabled in every FSD vehicle, so we get preprocessed trajectories and structures along with the trip data The second step is multi-trip reconstruction, which is the big and core piece of this machine So the video shows how the previously shown trip is reconstructed and aligned with other trips, basically other trips from different vehicles, not the same vehicle So this is done by multiple internal steps like course alignment, pairwise matching, joint optimization, then further surface refinement In the end, the human analyst comes in and finalizes the label So each heavy steps are already fully parallelized on the cluster, so the entire process usually takes just a couple of hours The last step is actually auto-labeling the new trips So here we use the same multi-trip alignment engine, but only between pre-built reconstruction and each new trip So it's much, much simpler than fully reconstructing all the clips altogether That's why it only takes 30 minutes per trip to auto-label instead of several hours of manual labeling And this is also the key of scalability of this machine This machine easily scales as long as we have available compute and trip data So about 50 trips were newly auto-labeled from this scene and some of them are shown here, so 53 from different vehicles So this is how we capture and transform the space-time slices of the world into the network supervision One thing I'd like to note is that Jagan just talked about how we auto-label our lanes We have auto-labels for almost every task that we do, including our planner And many of these are fully automatic, there's no humans involved For example, for objects, all the kinematics, the shapes, the futures, everything just comes from auto-labeling And the same is true for our occupancy too, and we have really just built a machine around this Yeah, so if you can go back one slide One more, it says parallelized on cluster So that sounds pretty straightforward, but it really wasn't Maybe it's fun to share how something like this comes about So a while ago we didn't have any auto-labeling at all, and then someone makes a script It starts to work, it starts working better, until you reach a volume that's pretty high And we clearly need a solution And so there were two other engineers in our team who were like, you know, that's an interesting, you know, thing What we needed to do was build a whole graph of essentially Python functions that would need to run one after the other First you pull the clip, then you do some cleaning, then you do some network inference, then another network inference Until you finally get this But so you need to do this at a large scale, so I tell them we probably need to shoot for, you know, 100,000 clips per day Or like 100,000 items, that seems good And so the engineers said, well, we can do, you know, a bit of post-gres and a bit of elbow grease, we can do it Meanwhile, we are a bit later and we're doing 20 million of these functions every single day Again, we pull in around half a million clips and on those we run a ton of functions, each of these, in a streaming fashion And so that's kind of the backend infra that's also needed to not just run training, but also auto-labeling Yeah, it really is like a factory that produces labels and production lines, yield, quality, inventory Like all of these same concepts applied to this label factory that applies for, you know, the factory for our cars That's right Okay, thanks, Tim and Ashok So, yeah, so concluding this section, I'd like to share a few more challenging and interesting examples for network for sure And even for humans, probably So from the top, there's like examples for like lack of lights, case or foggy night or roundabout and occlusions by heavy occlusions by parked cars And even rainy night with rain drops on camera lenses These are challenging, but once their original scenes are fully reconstructed by other clips, all of them can be auto-labeled So that our cars can drive even better through these challenging scenarios So, now, let me pass the mic to David to learn more about how Sim is creating the new world on top of these labels Thank you Thank you, Yegan My name is David and I'm going to talk about simulation So simulation plays a critical role in providing data that is difficult to source and or hard to label However, 3D scenes are notoriously slow to produce Take for example, the simulated scene playing behind me A complex intersection from Market Street in San Francisco It would take two weeks for artists to complete And for us, that is painfully slow However, I'm going to talk about using Yegan's automated ground truth labels along with some brand new tooling that allows us to procedurally generate this scene in many like it in just five minutes That's an amazing a thousand times faster than before So let's dive in to how a scene like this is created We start by piping the automated ground truth labels into our simulated world creator tooling inside the software Houdini Starting with road boundary labels, we can generate a solid road mesh and re-topologize it with the lane graph labels This helps inform important road details like cross-road slope and detailed material blending Next, we can use the line data and sweep geometry across its surface and project it to the road, creating lane paint decals Next, using median edges, we can spawned island geometry and populate it with randomized foliage This drastically changes the visibility of the scene Now the outside world can be generated through a series of randomized heuristics Modular building generators create visual obstructions while randomly placed objects like hydrants can change the color of the curves while trees can drop leaves below it obscuring lines or edges Next, we can bring in map data to inform positions of things like traffic traffic lights or stop signs We can trace along its normal to collect important information like number of lanes and even get accurate street names on the signs themselves Next, using lane graph, we can determine lane connectivity and spawn directional road markings on the road and their accompanying road signs And finally, with lane graph itself, we can determine lane adjacency and other useful metrics to spawn randomized traffic permutations inside our simulator And again, this is all automatic, no artist in the loop and happens within minutes And now this sets us up to do some pretty cool things Since everything is based on data and heuristics, we can start to fuzz parameters to create visual variations of the single ground truth It can be as subtle as object placement and random material swapping to more drastic changes like entirely new biomes or locations of environment like urban, suburban, or rural This allows us to create infinite, targeted permutations for specific ground truths that we need more ground truth for And all this happens within a click of a button And we can even take this one step further by altering our ground truth itself Say John wants his network to pay more attention to directional road markings to better detect an upcoming captive left turn lane We can start to procedurally alter our lane graph inside the simulator to help create entirely new flows through this intersection to help focus the network's attention to the road markings to create more accurate predictions And this is a great example of how this tooling allows us to create new data that can never be collected from the real world And the true power of this tool is in its architecture and how we can run all tasks in parallel to infinitely scale So you saw the tile creator tool in action converting the ground truth labels into their counterparts Next we can use our tile extractor tool to divide this data into geo hash tiles about 150 meter square in size We then save out that data into separate geometry and instance files This gives us a clean source of data that's easy to load and allows us to be rendering engine agnostic for the future Then using a tile loader tool we can summon any number of those cache tiles using a geo hash ID Currently we're doing about these 5x5 tiles or 3x3 usually centered around fleet hotspots or interesting lane graph locations And the tile loader also converts these tile sets into U assets for consumption by the unreal engine and gives you a finished product from what you saw in the first slide And this really sets us up for size and scale And as you can see on the map behind us we can easily generate most of San Francisco city streets And this didn't take years or even months of work but rather two weeks by one person We can continue to manage and grow all this data using our PDG network inside of the tooling This allows us to throw compute at it and regenerate all these tile sets overnight This ensures all environments are consistent, quality and features which is super important for training since new ontologies and signals are constantly released And now to come full circle, because we generated all these tile sets from ground truth data They contain all the weird intricacies from the real world We can combine that with the procedural, visual and traffic variety to create limitless, targeted data for the network to learn from And that concludes the SIM section, I'll pass it to Kate to talk about how we can use all this data to improve autopilot Thank you Thanks David, hi everyone, my name is Kate Park and I'm here to talk about the data engine Which is the process by which we improve our neural networks via data We're going to show you how we deterministically solve interventions via data And walk you through the life of this particular clip In this scenario, autopilot is approaching a turn and incorrectly predicts that crossing vehicle as stopped for traffic and thus a vehicle that we would slow down for In reality, there's nobody in the car, it's just awkwardly parked We've built this tooling to identify the mispredictions, correct the label and categorize this clip into an evaluation set This particular clip happens to be one of 126 that we've diagnosed as challenging parked cars at turns Because of this infra, we can curate this evaluation set without any engineering resources custom to this particular challenge case To actually solve that challenge case requires mining thousands of examples like it And it's something Tesla can trivially do We simply use our data sourcing infra, request data and use the tooling shown previously to correct the labels By surgically targeting the mispredictions of the current model, we're only adding the most valuable examples to our training set We surgically fix 13,900 clips and because those were examples where the current model struggles We don't even need to change the model architecture, a simple weight update with this new valuable data is enough to solve the challenge case So you see we no longer predict that crossing vehicle as stopped, as shown in orange, but parked, as shown in red In academia, we often see that people keep data constant, but at Tesla it's very much the opposite We see time and time and again that data is one of the best if not the most deterministic lever to solving these interventions We just showed you the data engine loop for one challenge case, namely these parked cars at turns But there are many challenge cases even for one signal of vehicle movement We apply this data engine loop to every single challenge case we've diagnosed, whether it's buses, curvy roads, stopped vehicles, parking lots And we don't just add data once, we do this again and again to perfect the semantic In fact, this year we updated our vehicle movement signal five times and with every weight update trained on the new data We push our vehicle movement accuracy up and up This data engine framework applies to all our signals, whether they're 3D, multi-cam video, whether the data is human labeled, auto-labeled, or simulated Whether it's an offline model or an online model And Tesla is able to do this at scale because of the fleet advantage, the infra that our NG team has built, and the labeling resources that feed our networks To train on all this data, we need a massive amount of compute, so I'll hand it off to Pete and Ganesh to talk about the Dojo supercomputing platform Thank you Thank you, Katie Thanks everybody, thanks for hanging in there, we're almost there My name is Pete Bannon, I run the custom silicon and low voltage teams at Tesla And my name is Ganesh Renke, I run the Dojo program Thank you I'm frequently asked, why is a car company building a supercomputer for training?
Paragraph 21
And this question fundamentally misunderstands the nature of Tesla At its heart, Tesla is a hardcore technology company All across the company, people are working hard in science and engineering to advance the fundamental understanding and methods that we have available to build cars, energy solutions, robots, and anything else that we can do to improve the human condition around the world It's a super exciting thing to be a part of, and it's a privilege to run a very small piece of it in the semiconductor group Tonight we're going to talk a little bit about Dojo and give you an update on what we've been able to do over the last year But before we do that, I wanted to give a little bit of background on the initial design that we started a few years ago When we got started, the goal was to provide a substantial improvement to the training latency for our autopilot team Some of the largest neural networks they train today run for over a month, which inhibits their ability to rapidly explore alternatives and evaluate them So a 30X speedup would be really nice if we could provide it at a cost competitive and energy competitive way To do that, we wanted to build a chip with a lot of arithmetic units that we could utilize at a very high efficiency And we spent a lot of time studying whether we could do that using DRAM, various packaging ideas, all of which failed And in the end, even though it felt like an unnatural act, we decided to reject DRAM as the primary storage medium for this system And instead focus on SRAM embedded in the chip SRAM provides, unfortunately, a modest amount of capacity, but extremely high bandwidth and very low latency, and that enables us to achieve high utilization with the arithmetic units Those choices, that particular choice led to a whole bunch of other choices For example, if you want to have virtual memory, you need page tables, they take up a lot of space, we didn't have space, so no virtual memory So we also don't have interrupts, the accelerator is a bare bonds, raw piece of hardware that's presented to a compiler and the compiler is responsible for scheduling everything that happens in a deterministic way So there's no need or even desire for interrupts in the system We also chose to pursue model parallelism as a training methodology, which is not the typical situation most machines today use data parallelism, which consumes additional memory capacity, which we obviously don't have So all of those choices led us to build a machine that is pretty radically different from what's available today We also had a whole bunch of other goals, one of the most important ones was no limits So we wanted to build a compute fabric that would scale in an unbounded way for the most part, I mean obviously there's physical limits now and then But pretty much if your model was too big for the computer, you just had to go buy a bigger computer, that's what we were looking for Today the way machines are packaged, there's a pretty fixed ratio of for example GPU, CPUs and DRAM capacity and network capacity And we really wanted to disaggregate all that so that as models evolved, we could vary the ratios of those various elements and make the system more flexible to meet the needs of the autopilot team And it's so true, no limits philosophy was our guiding star all the way, all of our choices were centered around that And to the point that we didn't want traditional data center infrastructure to limit our capacity to execute these programs at speed That's why we integrated vertically our data center, the entire data center by doing a vertical integration of the data center We could extract new levels of efficiency, we could optimize power delivery, cooling and as well as system management across the whole data center stack Rather than doing box by box and integrating that, those boxes into data centers And to do this, we also wanted to integrate early to figure out limits of scale for our software workloads So we integrated Dojo environment into our autopilot software very early and we learned a lot of lessons And today Bill Chang will go over our hardware update as well as some of the challenges that we faced along the way And Rajiv Kurian will give you a glimpse of our compiler technology as well as go over some of our cool results Great Thanks Pete, thanks Ganesh I'll start tonight with a high level vision of our system that will help set the stage for the challenges and the problems we're solving And then also how software will then leverage this for performance Now our vision for Dojo is to build a single unified accelerator, a very large one Software would see a seamless compute plane with globally addressable, very fast memory and all connected together with uniform high bandwidth and low latency Now to realize this, we need to use density to achieve performance Now we leverage technology to get this density in order to break levels of hierarchy all the way from the chip to the scale out systems Now silicon technology has done this for decades Chips have followed Moore's law for density integration to get performance scaling Now a key step in realizing that vision was our training tile Probably can we integrate 25 dies at extremely high bandwidth but we can scale that to any number of additional tiles by just connecting them together Now last year we showcased our first functional training tile and at that time we already had workloads running on it And since then the team here has been working hard and diligently to deploy this at scale Now we've made amazing progress and had a lot of milestones along the way And of course we've had a lot of unexpected challenges But this is where our fail fast philosophy has allowed us to push our boundaries Now pushing density for performance presents all new challenges One area is power delivery Here we need to deliver the power to our compute die and this directly impacts our top line compute performance But we need to do this at unprecedented density We need to be able to match our die pitch with a power density of almost 1 amp per millimeter squared And because of the extreme integration this needs to be a multi-tiered vertical power solution And because there's a complex heterogeneous material stack up we have to carefully manage the material transition Especially CTE Now why does the coefficient of thermal expansion matter in this case? CTE is a fundamental material property and if it's not carefully managed that stack up would literally rip itself apart We started this effort by working with vendors to develop this power solution But we realized that we actually had to develop this in-house Now to balance schedule and risk we built quick iterations to support both our system bring up in software development And also to find the optimal design and stack up that would meet our final production goals And in the end we were able to reduce CTE over 50% and meet our performance by 3x over our initial version Now needless to say finding this optimal material stack up while maximizing performance at density is extremely difficult Now we did have unexpected challenges along the way Here's an example where we pushed the boundaries of integration that led to component failures This started when we scaled up to larger and longer workloads and then intermittently a single site on a tile would fail Now they started out as recoverable failures but as we pushed some much higher and higher power these would become permanent failures Now to understand this failure you have to understand why and how we build our power modules Solving density at every level is the cornerstone of actually achieving our system performance Now because our XY plane is used for high bandwidth communication everything else must be stacked vertically This means all other components other than our die must be integrated into our power modules Now that includes our clock and our power supplies and also our system controllers Now in this case the failures were due to losing clock output from our oscillators And after an extensive debug we found that the root cause was due to vibrations on the module from piezoelectric effects Our nearby capacitors Now singing caps are not a new phenomenon and in fact very common in power design But normally clock chips are placed in a very quiet area of the board and often not affected by power circuits But because we needed to achieve this level of integration these oscillators need to be placed in very close proximity Now due to our switching frequency and then the vibration resonance created It caused out of plane vibration on our MEMS oscillator that caused it to crack Now the solution to this problem is a multi-prong approach We can reduce the vibration by using soft terminal caps We can update our MEMS part with a lower Q factor for the out of plane direction And we can also update our switching frequency to push the resonance further away from these sensitive bands Now in addition to the density at the system level we've been making a lot of progress at the infrastructure level We knew that we had to read examine every aspect of the data center infrastructure in order to support our unprecedented power and cooling density We brought in a fully custom designed CDU to support Dojo's dense cooling requirements And the amazing part is we're able to do this at a fraction of the cost versus buying off the shelf and modifying it And since our Dojo cabinet integrates enough power and cooling to match an entire row of standard IT racks We need to carefully design our cabinet and infrastructure together And we've already gone through several iterations of this cabinet to optimize this And earlier this year we started low testing our power and cooling infrastructure And we were able to push it over 2 megawatts before we tripped our substation and got a call from the city Now last year we introduced only a couple of components of our system The custom D1 die and the training tile, but we teased the exit pod as our end goal We'll walk through the remaining parts of our system that are required to build out this exit pod Now the system tray is a key part of realizing our vision of a single accelerator It enables us to seamlessly connect tiles together, not only within the cabinet, but between cabinets We can connect these tiles at very tight spacing across the entire accelerator And this is how we achieve our uniform communication This is a laminated bus bar that allows us to integrate very high power, mechanical and thermal support, and an extremely dense integration It's 75 millimeters in height and supports 6 tiles at 135 kilograms This is the equivalent of 3 to 4 fully loaded high performance racks Next we need to feed data to the training tiles This is where we've developed the Dojo interface processor It provides our system with high bandwidth DRAM to stage our training data And it provides full memory bandwidth to our training tiles using TTP, our custom protocol that we use to communicate across our entire accelerator It also has high speed Ethernet that helps us extend this custom protocol over standard Ethernet And we provide native hardware support for this with little to no software overhead And lastly we can connect to it through a standard Gen4 PCIe interface Now we pair 20 of these cards per tray and that gives us 640 gigabytes of high bandwidth DRAM And this provides our disaggregated memory layer for our training tiles These cards are a high bandwidth ingest path both through PCIe and Ethernet They also provide a high-ratex Z-connectivity path that allows shortcuts across our large Dojo accelerator Now we actually integrate the host directly underneath our system tray These hosts provide our ingest processing and connect to our interface processors through PCIe These hosts can provide hardware video decoder support for video-based training And our user applications land on these hosts so we can provide them with the standard X86 Linux environment Now we can put two of these assemblies into one cabinet and pair it with redundant power supplies that do direct conversion of three-phase 480-volt AC power to 52-volt DC power Now by focusing on density at every level we can realize the vision of a single accelerator Now starting with the uniform nodes on our custom D1 die we can connect them together in our fully integrated training tile And then finally seamlessly connecting them across cabinet boundaries to form our Dojo accelerator And all together we can house two full accelerators in our Exapod for a combined one exa-flop of ML compute Now all together this amount of technology and integration has only ever been done a couple of times in the history of compute Next we'll see how software can leverage this to accelerate their performance Thanks Bill, my name is Rajiv and I'm going to talk some numbers So our software stack begins with the PyTorch extension that speaks to our commitment to run standard PyTorch models out of the box We're going to talk more about our JIT compiler and the ingest pipeline that feeds the hardware with data Abstractly, performance is tops times utilization times accelerator occupancy We've seen how the hardware provides peak performance is the job of the compiler to extract utilization from the hardware while code is running on it And it's the job of the ingest pipeline to make sure that data can be fed at a throughput high enough for the hardware to not ever starve So let's talk about why communication-bound models are difficult to scale But before that let's look at why ResNet 50-like models are easier to scale You start off with a single accelerator, run the forward and backward passes, followed by the optimizer Then to scale this up you run multiple copies of this on multiple accelerators And while the gradients produced by the backward pass do need to be reduced and this introduces some communication, this can be done pipeline with the backward pass This setup scales fairly well, almost linearly For models with much larger activations we run into a problem as soon as we want to run the forward pass The batch size that fits in a single accelerator is often smaller than the batch norm surface So to get around this researchers typically run this setup on multiple accelerators in sync batch norm mode This introduces latency bound communication to the critical path of the forward pass and we already have a communication bottleneck And while there are ways to get around this they usually involve tedious manual work best suited for a compiler And ultimately there's no skirting around the fact that if your state does not fit in a single accelerator you can be communication bound And even with significant efforts from our ML engineers we see such models don't scale linearly The doger system was built to make such models work at high utilization The high density integration was built to not only accelerate the compute bound portions of a model but also the latency bound portions Like a batch norm or the bandwidth bound portions like a gradient all reduced or a parameter all gathered A slice of the doger mesh can be carved out to run any model The only thing users need to do is to make the slice large enough to fit a batch norm surface for their particular model After that the partition presents itself as one large accelerator freeing the users from having to worry about the internal details of execution And as the job of the compiler to maintain this abstraction Fine grain synchronization primitives in uniform low latency makes it easy to accelerate all forms of parallelism across integration boundaries Tensors are usually stored sharded in SRAM and replicated just in time for a layer's execution We depend on the high doger bandwidth to hide this replication time Tensor replication and other data transfers are overlapped with compute and the compiler can also recompute layers when it's profitable to do so We expect most models to work out of the box As an example we took the recently released stable diffusion model and got it running on dojo in minutes Out of the box the compiler was able to map it in a model parallel manner on 25 dojo dies Here are some pictures of a Cybertruck on Mars generated by stable diffusion running on dojo Looks like it still has some ways to go before matching the Tesla design studio team So we've talked about how communication bottlenecks can hamper scalability Perhaps an asset test of a compiler and the underlying hardware is executing a cross die batch norm layer Like mentioned before this can be a serial bottleneck The communication phase of a batch norm begins with nodes computing their local mean and standard deviations Then coordinating to reduce these values, then broadcasting these values back and then they resume their work in parallel So what would an ideal batch norm look like on 25 dojo dies? Let's say the previous less activations are already split across dies We would expect the 350 nodes on each die to coordinate and produce die local mean and standard deviation values Ideally these would get further reduced with the final value ending somewhere towards the middle of the tile We would then hope to see a broadcast of this value radiating from the center Let's see how the compiler actually executes a real batch norm operation across 25 dies The communication trees were extracted from the compiler and the timing is from a real hardware one We're about to see 8,750 nodes on 25 dies coordinating to reduce and then broadcast the batch norm mean and standard deviation values Die local reduction followed by global reduction towards the middle of the tile Then the reduced value broadcast radiating from the middle accelerated by the hardware's broadcast facility This operation takes only 5 microseconds on 25 dojo dies The same operation takes 150 microseconds on 24 GPUs This is an orders of magnitude improvement over GPUs And while we talked about an already used operation in the context of a batch norm It's important to reiterate that the same advantages apply to all other communication primitives And these primitives are essential for large scale training So how about full model performance? So while we think that ResNet 50 is not a good representation of real world Tesla workloads It is a standard benchmark, so let's start there We are already able to match the 100 die for die However, perhaps a hint of dojo's capabilities is that we're able to hit this number with just a batch of 8 per die But dojo was really built to tackle larger complex models So when we set out to tackle real world workloads, we looked at the usage patterns of our current GPU cluster And two models stood out, the autolabeling networks, a class of offline models that are used to generate ground truth And the occupancy networks that you heard about The autolabeling networks are large models that have high arithmetic intensity While the occupancy networks can be ingest bound We chose these models because together they account for a large chunk of our current GPU cluster usage And they would challenge the system in different ways So how do we do on these two networks? The results we're about to see were measured on multi die systems for both the GPU and dojo, but normalized to per die numbers On our autolabeling network, we're already able to surpass the performance of an A100 With our current hardware running on our older generation VRMs On our production hardware with our newer VRMs, that translates to doubling the throughput of an A100 And our model showed that with some key compiler optimizations, we could get to more than 3x the performance of an A100 We see even bigger leaps on the occupancy network Almost 3x with our production hardware, with room for more So what does that mean for Tesla? With a current level of compiler performance, we could replace the ML compute of 1, 2, 3, 4, 5 and 6 GPU boxes with just a single dojo tile And this dojo tile costs less than one of these GPU boxes What it really means is that networks that took more than a month to train now take less than a week Alas, when we measure things, it did not turn out so well.
Paragraph 22
At the PyTorch level, we did not see our expected performance out of the gate And this timeline chart shows our problem. The teeny, tiny little green bars, that's the compile code running on the accelerator The row is mostly white space where the hardware is just waiting for data With our dense ML compute, dojo hosts effectively have 10x more ML compute than the GPU hosts. The data loader is running on this one host Simply couldn't keep up with all that ML hardware So to solve our data loader scalability issues, we knew we had to get over the limit of this single host The Tesla transport protocol moves data seamlessly across hosts, tiles and ingest processors So we extended the Tesla transport protocol to work over Ethernet. We then built the dojo network interface card, the D-NIC, to leverage TTP over Ethernet This allows any host with a D-NIC card to be able to DMA2 and from other TTP endpoints So we started with the dojo mesh, then we added a tier of data loading hosts equipped with the D-NIC card We connected these hosts to the mesh via an Ethernet switch. Now every host in this data loading tier is capable of reaching all TTP endpoints in the dojo mesh via hardware accelerated DMA After these optimizations went in, our occupancy went from 4% to 97% So the data loading sections have reduced drastically and the ML hardware has kept busy We actually expect this number to go to 100% pretty soon After these changes went in, we saw the full expected speed up from the PyTorch layer and we were back in business So we started with hardware design that breaks through traditional integration boundaries in service of our vision of a single giant accelerator We've seen how the compiler and ingest layers build on top of that hardware So after proving our performance on these complex real-world networks, we knew what our first large-scale deployment would target Our high arithmetic intensity auto-labeling networks Today that occupies 4,000 GPUs over 72 GPU racks With our dense computer and our high performance, we expect to provide the same throughput with just 4 dojo cabinets And these 4 dojo cabinets will be part of our first exapod that we plan to build by quarter one of 2023 This will more than double Tesla's auto-labeling capacity The first exapod is part of a total of 7 exapods that we plan to build in Palo Alto right here across the wall And we have a display cabinet from one of these exapods for everyone to look at 6 tiles densely packed on a tray, 54 petaflops of compute, 640 gigabytes of high bandwidth memory with power and host defeated A lot of compute And we're building out new versions of all our cluster components and constantly improving our software to hit new limits of scale We believe that we can get another 10x improvement with our next generation hardware And to realize our ambitious goals, we need the best software and hardware engineers So please come talk to us or visit tesla.com. Alright, so hopefully that was enough detail And now we can move to questions And guys, I think the team can come out on stage We really wanted to show the depth and breadth of Tesla in artificial intelligence, compute hardware, robotics actuators And try to really shift the perception of the company away from, you know, a lot of people think we're like just a car company Or we make cool cars, whatever But most people have no idea that Tesla is arguably the leader in real world AI hardware and software And that we're building what is arguably the most radical computer architecture since the Kray-1 supercomputer And I think if you're interested in developing some of the most advanced technology in the world that's going to really affect the world in a positive way Tesla's the place to be So yeah, let's fire away with some questions I think there's a mic at the front and a mic at the back Just throw mics at people Jump all for the mic Yeah, hi, thank you very much I was impressed here I was impressed very much by Optimus, but I wonder why did not driven the hand Why did you choose a tendon-driven approach for the hand?
Paragraph 23
Because tendons are not very durable And why spring-loaded? Cool, awesome, yes, that's a great question You know, when it comes to any type of actuation scheme, there's trade-offs between, you know, whether or not it's a tendon-driven system or some type of linkage-based system Keep the mic close to your mouth A little bit closer, hear me? Cool Yeah, the main reason why we went for a tendon-based system is that, you know, first we actually investigated some synthetic tendons, but we found that metallic boating cables are, you know, a lot stronger One of the advantages of these cables is that it's very good for part reduction We do want to make a lot of these hands, so having a bunch of parts, a bunch of small linkages ends up being, you know, a problem when you're making a lot of something One of the big reasons that, you know, tendons are better than linkages in a sense is that you can be anti-backlash So anti-backlash essentially, you know, allows you to not have any gaps or, you know, stuttering motion in your fingers Spring-loaded, mainly what spring-loaded allows us to do is allows us to have active opening So instead of having to have two actuators to drive the fingers closed and then open, we have the ability to, you know, have the tendon drive them closed and then the springs passively extend And this is something that's seen in our hands as well, right? We have the ability to actively flex and then we also have the ability to extend Yeah I mean, our goal with Optimus is to have a robot that is maximally useful as quickly as possible So there's a lot of ways to solve the various problems of a humanoid robot And we're probably not barking up the right tree on all the technical solutions And I should say that we're open to evolving the technical solutions that you see here over time, they're not locked in stone But we have to pick something, and we want to pick something that's going to allow us to produce the robot as quickly as possible and have it, like I said, be useful as quickly as possible We're trying to follow the goal of fastest path to a useful robot that can be made at volume And we're going to test the robot internally at Tesla in our factory and just see, like, how useful is it Because you have to have a, you've got to close the loop on reality to confirm that the robot is in fact useful And, yeah, so we're just going to use it to build things And we're confident we can do that with the hand that we have currently designed But I'm sure there'll be hand version 2, version 3, and we may change the architecture quite significantly over time Hi, the Optimus robot is really impressive, you did a great job, bipedal robots are really difficult But what I noticed might be missing from your plan is to acknowledge the utility of the human spirit And I'm wondering if Optimus will ever get a personality and be able to laugh at our jokes while it folds our clothes Yeah, absolutely. I think we want to have really fun versions of Optimus And so that Optimus can both be utilitarian and do tasks, but can also be kind of like a friend and a buddy And hang out with you, and I'm sure people will think of all sorts of creative uses for this robot And, you know, the thing, once you have the core intelligence and actuators figured out Then you can actually, you know, put all sorts of costumes, I guess, on the robot I mean, you can make the robot look, you can skin the robot in many different ways And I'm sure people will find very interesting ways to, yeah, versions of Optimus Thanks for the great presentation I wanted to know if there was an equivalent to interventions in Optimus It seems like labeling through moments where humans disagree with what's going on is important And in a humanoid robot, that might be also a desirable source of information Yeah, I think we will have ways to remote operate the robot and intervene when it does something bad Especially when we are training the robot and bringing it up And hopefully we, you know, design it in a way that we can stop the robot from, if it's going to hit something We can just, like, hold it and it will stop, it won't, like, you know, crush your hand or something And those are all intervention data Yeah, and we can learn a lot from our simulation systems, too Where we can check for collisions and supervise that those are bad actions Yeah, I mean, so Optimus, we went over time for it to be, you know, an android, the kind of android that you've seen in sci-fi movies Like Star Trek, The Next Generation, like data But obviously we could program the robot to be less robot-like and more friendly And, you know, you can obviously learn to emulate humans and feel very natural So as AI in general improves, we can add that to the robot And, you know, it should be obviously able to do simple instructions or even intuit what it is that you want So you could give it a high level instruction and then it can break that down into a series of actions And take those actions Hi, yeah, it's exciting to think that with the Optimus you will think that you can achieve orders of magnitude of improvement in economic output That's really exciting And when Tesla started, the mission was to accelerate the advent of renewable energy or sustainable transport So with the Optimus, do you still see that mission being the mission statement of Tesla or is it going to be updated with, you know, mission to accelerate the advent of, I don't know, infinite abundance or limitless economy Yeah, it is not strictly speaking, Optimus is not strictly speaking directly in line with accelerating sustainable energy To the degree that it is more efficient at getting things done than a person, it does, I guess, help with sustainable energy But I think the mission effectively does somewhat broaden with the advent of Optimus to, you know, I don't know, making the future awesome So, you know, I think you look at Optimus and I know about you, but I'm excited to see what Optimus will become And, you know, this is like, you know, if you could, I mean, you can tell like any given technology, do you want to see what it's like in a year, two years, three years, four years, five years, ten? I'd say for sure, you definitely want to see what's happened with Optimus Whereas, you know, a bunch of other technologies are, you know, sort of plateaued About name names here, but, you know, so, I think Optimus is going to be incredible in like five years, ten years like mind-blowing And I'm really interested to see that happen, and I hope you are too I have a quick question here, Justin, and I was wondering, like, are you planning to extend like conversational capabilities for the robot?
Paragraph 24
And my second full-on question to that is, what's like the end goal? What's the end goal with Optimus? Yeah, Optimus would definitely have conversational capabilities So, you'd be able to talk to it and have a conversation, and it would feel quite natural So, from an end goal standpoint, I don't know, I think it's going to keep evolving, and I'm not sure where it ends up, but some place is interesting for sure And, you know, we always have to be careful about the, you know, don't go down the terminator path That's a, you know, I thought maybe we should start off with a video of like the terminator starting off with this, you know, skull crushing But that might be, you know, people might not get too seriously So, you know, we do want Optimus to be safe, so we are designing in safeguards where you can locally stop the robot And, you know, with like basically a localized control ROM that you can't update over the internet Which I think that's quite important, essential, frankly So, like a localized stop button or remote control, something like that, that cannot be changed But, I mean, it's definitely going to be interesting, it won't be boring Okay, yeah, I see today you have a very attractive product with Dojo and its applications So, I'm wondering what's the future for the Dojo platform? So, you know, like provide like infrastructure and service like AWS or you will like sell the chip like the NVIDIA So, basically, what's the future? Because I say you use 7nm, so the developer cost is like easily over 10 million US dollars How do you make the business like business wise? Dojo is a very big computer and actually will use a lot of power and need a lot of cooling So, I think it's probably going to make more sense to have Dojo operate in like an Amazon Web Services manner Than to try to sell it to someone else So, that would be the most efficient way to operate Dojo is just have it be a service that you can use That's available online and that where you can train your models way faster and for less money And as the world transitions to software 2.0 And that's on the bingo card Someone I know has to know to drink 5 tequila So, let's see, software 2.0 will use a lot of neural net training So, it kind of makes sense that over time as there's more neural net stuff People will want to use the fastest, lowest cost neural net training system So, I think there's a lot of opportunity in that direction Hi, my name is Ali Jahanian Thank you for this event, it's very inspirational My question is, I'm wondering what is your vision for humanoid robots that understand our emotions and art And can contribute to our creativity Well, I think you're already seeing robots that at least are able to generate very interesting art Like Dali and Dali 2 And I think we'll start seeing AI that can actually generate even movies that have coherence Like interesting movies and tell jokes So, it's quite remarkable how fast AI is advancing at many companies besides Tesla We're headed for a very interesting future And yeah, so, any guys want to comment on that?
Paragraph 25
Yeah, I guess the Optimus Robot can come up with physical art, not just digital art You can ask for some dance moves in text or voice and then you can produce those in the future So, it's a lot of physical art, not just digital art Oh, yeah, computers can absolutely make physical art, yeah, 100% Yeah, like dance, play soccer or whatever you... I mean, it needs to get more agile over time, for sure Thanks so much for the presentation Now, for the Tesla Autopilot slides, I noticed that the models that you were using were heavily motivated by language models And I was wondering what the history of that was and how much of an improvement it gave I thought that that was a really interesting, curious choice to use language models for the lane transitioning So, there are sort of two aspects for why we transition to language modeling So, the first... Talk loud and close Okay, got it Yeah, so the language models help us in two ways The first way is that it lets us predict lanes that we couldn't have otherwise As Ashok mentioned earlier, basically when we predicted lanes in sort of a dense 3D fashion You can only model certain kinds of lanes, but we want to get those criss-crossing connections inside of intersections It's just not possible to do that without making it a graph prediction If you try to do this with dense segmentation, it just doesn't work Also, the lane prediction is a multimodal problem Sometimes you just don't have sufficient visual information to know precisely how things look on the other side of the intersection So you need a method that can generalize and produce coherent predictions You don't want to be predicting two lanes and three lanes at the same time You want to commit to one in a general model like these language models provides that Hi Hi, my name is Giovanni Yeah, thanks for the presentation. It's really nice I have a question for FSD team For the neural networks, how do you test... How do you do unit tests, software unit tests on that? Do you have a bunch or I don't know, mid-thousands or...
Paragraph 26
Yes, cases where the neural network that after you train it, you have to pass it Before you release it as a product, right? Yeah, what's your software unit testing strategies for this, basically? Yeah, glad you asked. There's like a series of tests that we have defined starting from unit tests for software itself But then for the neural network models, we have VAP sets defined where you can define... If you just have a large test set, that's not enough what we find We need like sophisticated VAP sets for different failure modes And then we queate them and grow them over the time of the product So over the years, we have like hundreds of thousands of examples where we have been failing in the past That we have curated and so for any new model, we test against the entire history of these failures And then keep adding to this test set On top of this, we have shadow modes where we ship these models in silent to the car And we get data back on where they are failing or succeeding And there's an extensive QA program It's very hard to ship for regression There's like nine levels of filters before it hits customers But then we have really good infra to make this all efficient I'm one of the QA testers, so I have QA the car... Yeah, QA tester Yeah, so I'm constantly in the car just being queuing like whatever the latest alpha build is that doesn't totally crash Yeah, finds a lot of bugs Hi, great event.
Paragraph 27
I have a question about foundational models for autonomous driving We have all seen that big models that really can... When you scale up with data and model parameter from GP3 to POM, it can actually now do reasoning Do you see that it's essential scaling up foundational models with data and size And then at least you can get a teacher model that potentially can solve all the problems And then you distill to a student model Is that how you see foundational models relevant for autonomous driving? That's quite similar to our auto labeling models So we don't just have models that run in the car We train models that are entirely offline that are extremely large that can't run in real time on the car So we just run those offline on the servers producing really good labels that can then train the online networks So that's one form of distillation of these teacher-student models In terms of foundation models, we are building some really, really large datasets that are multiple petabytes And we are seeing that some of these tasks work really well when we have these large datasets Kinematics, like I mentioned, video in, all the kinematics out of all the objects and up to the fourth derivative And people thought we couldn't do detection with cameras Detection, depth, velocity, acceleration And imagine how precise these have to be for these higher-order derivatives to be accurate And this all comes from these kind of large datasets and large models So we are seeing the equivalent of foundation models in our own way for geometry and kinematics and things like those Do you want to add anything, John? Yeah, I'll keep it brief Basically, whenever we train on a larger dataset, we see big improvements in our model performance And basically, whenever we initialize our networks with some pre-training steps from some other auxiliary tasks We basically see improvements The self-supervised or supervised with large datasets both help a lot Hi, so at the beginning, Elon said that Tesla was potentially interested in building artificial general intelligence systems Given the potentially transformative impact of technology like that It seems prudent to invest in technical AGI safety expertise specifically I know Tesla does a lot of technical, narrow AI safety research I was curious if Tesla was intending to try to build expertise in technical artificial general intelligence safety specifically Well, I mean, if we start looking like we're going to be making a significant contribution to artificial general intelligence Then we'll for sure invest in safety on big believer in AI safety I think there should be an AI sort of regulatory authority at the government level Just as there is a regulatory authority for anything that affects public safety So we have regulatory authority for aircraft and cars and sort of food and drugs Because they affect public safety and AI also affects public safety So I think, and this is not really something that government I think understands yet I think there should be a referee that is trying to ensure public safety for AGI And you think of like, well, what are the elements that are necessary to create AGI? Like the accessible dataset is extremely important And if you've got a large number of cars and humanoid robots processing petabytes of video data and audio data from the real world Just like humans, that might be the biggest dataset, probably is the biggest dataset Because in addition to that, you can obviously incrementally scan the internet But what the internet can't quite do is have millions or hundreds of millions of cameras in the real world Like I said, with audio and other sensors as well So I think we probably will have the most amount of data And probably the most amount of training power Therefore probably we will make a contribution to AGI Hey, I noticed the semi was back there, but we haven't talked about it too much I was just wondering for the semi truck, what are the changes you're thinking about from a sensing perspective? I imagine there's very different requirements obviously than just a car And if you don't think that's true, why is that true?
Paragraph 28
No, I think basically you can drive a car Think about what drives any vehicle, it's a biological neural net with eyes With cameras essentially What is your primary sensors are? Two cameras on a slow gimbal, a very slow gimbal That's your head So if a biological neural net with two cameras on a slow gimbal can drive a semi truck Then if you've got like eight cameras with continuous 360 degree vision Operating at a higher frame rate and a much higher reaction rate Then I think it is obvious that you should be able to drive a semi or any vehicle much better than human Hi, my name is Akshay, thank you for the event Assuming Optimus would be used for different use cases and would evolve at different speeds for these use cases Would it be possible to sort of develop and deploy different software and hardware components independently And deploy them in Optimus so that the overall feature development is faster for Optimus Okay, we did not comprehend Unfortunately our neural net did not comprehend the question Next question Hi, I want to switch the gear to the autopilot So when you guys plan to roll out the FSD beta to countries other than US and Canada And also my next question is what's the biggest bottleneck or the technology or barrier you think in the current autopilot stack And how you envision to solve that to make the autopilot is considerably better than human in terms of performance matrix Like safety assurance and the human confidence I think you also mentioned for the FSD V11 you are going to combine the highway and the city as a single stack And some architectural big improvements, can you maybe expand a bit on that, thank you Well, that's a whole bunch of questions We're hopeful to be able to, I think from a technical standpoint FSD beta should be possible to roll out FSD beta worldwide by the end of this year But for a lot of countries we need regulatory approval And so we are somewhat gated by the regulatory approval in other countries But I think from a technical standpoint it will be ready to go to a worldwide beta by the end of this year And there's quite a big improvement that we're expecting to release next month That will always be especially good at assessing the velocity of fast moving cross traffic And a bunch of other things So, anyone want to elaborate? I guess so, there used to be a lot of differences between production autopilot and the full self driving beta But those differences have been getting smaller and smaller over time I think just a few months ago we now use the same vision only object detection stack in both FSD and in the production autopilot on all vehicles There's still a few differences, the primary one being the way that we predict lanes right now So we upgraded the modeling of lanes so that it could handle these more complex geometries like I mentioned in the talk In production autopilot we still use a simpler lane model But we're extending our current FSD beta models to work in all sort of highway scenarios as well The version of FSD beta that I drive actually does have the integrated stack So it uses the FSD stack both in city streets and highway and it works quite well for me But we need to validate it in all kinds of weather like heavy rain, snow, dust And just make sure it's working better than the production stack across a wide range of environments But we're pretty close to that I think it's, I don't know, maybe, it'll definitely be before the end of the year and maybe November Yeah, in our personal drives, the FSD stack on highway drives already way better than the production stack we have And we do expect to also include the parking lot stack as a part of the FSD stack before the end of this year So that will basically bring us to, you sit in the car in the parking lot and drive till the end of the parking lot at a parking spot before the end of this year And in terms of the fundamental metric to optimize against is how many miles between a necessary intervention So just massively improving how many miles the car can drive in full autonomy before an intervention is required that is safety critical So, yeah, that's the fundamental metric that we're measuring every week and we're making radical improvements on that Hi, thank you, thank you so much for the presentation, very inspiring My name is Daisy, I actually have a non-technical question for you I'm curious, if you are back to your 20s, what are some of the things you wish you knew back then? What are some advice you would give to your younger self? Well, I'm trying to figure out something useful to say Yeah, a joint Tesla would be one thing Yeah, I think just trying to expose yourself to as many smart people as possible I don't read a lot of books You know, I did do that though So, I think there's some merit to just also not being necessarily too intense And enjoying the moment a bit more, I would say to 20-something me Just to stop and smell the roses occasionally would probably be a good idea You know, it's like when we were developing the Falcon 1 rocket on the Quageline Atoll And we had this beautiful little island that we were developing the rocket on And not once during that entire time did I even have a drink on the beach I'm like, I should have had a drink on the beach, that would have been fine Thank you very much I think you have excited all of the robotics people with Optimus This feels very much like 10 years ago in driving But as driving has proved to be harder than it actually looked 10 years ago What do we know now that we didn't 10 years ago that would make, for example, AGI on a humanoid come faster? Well, I mean, it seems to me that AGI is advancing very quickly Hardly a week goes by without some significant announcement And, yeah, I mean, at this point, like, AI seems to be able to win at almost any rule-based game It's able to create extremely impressive art Engage in conversations that are very sophisticated, you know, write essays And these just keep improving And there's so many more talented people working on AI And the hardware is getting better AI is on a super, like, a strong exponential curve of improvements Independent of what we do at Tesla And obviously we'll benefit somewhat from that exponential curve of improvement with AI Like, Tesla just also has to be very good at actuators Motors gearboxes, controllers, power electronics, batteries, sensors And, you know, really, like, I'd say the biggest difference between the robot on four wheels And the robot with arms and legs is getting the actuators right It's an actuators and sensors problem And obviously, how you control those actuators and sensors But it's, yeah, actuators and sensors and how you control the actuators I don't know, we have to have, like, the ingredients necessary to create a compelling robot And we're doing it, so...
Paragraph 29
Hi, Ilan You are actually bringing the humanity to the next level Literally, Tesla, and you are bringing the humanity to the next level So, you said Optimus Prime, Optimus will be used in next Tesla factory My question is, will a new Tesla factory be fully run by Optimus program? And when can general public order a humanoid? Yeah, I think it'll, you know, we're going to start Optimus with very simple tasks in the factory You know, like maybe just, like, loading a part, like you saw in the video You know, carrying a part from one place to another Or loading a part into one of our more conventional robot cells to, you know, that welds body together So we'll start, you know, just trying to, how do we make it useful at all? And then gradually expand the number of situations where it's useful And I think that number of situations where Optimus is useful will grow exponentially Like really, really fast In terms of when people can order one, I don't know, I think it's not that far away Well, I think you mean, when can people receive one? So, I don't know, I'm like, I'd say probably within three years And not more than five years Within three to five years, you could probably receive an Optimus I feel the best way to make the progress for AGI is to involve as many smart people across the world as possible And given the size and resource of Tesla compared to robot companies And given the state of humanoid research at the moment Would it make sense for the kind of Tesla to sort of open source some of the simulation hardware parts? I think Tesla can still be the dominant platformer where it can be something like an Android OS Or like an iOS stuff for the entire humanoid research Would that be something that rather than keeping the Optimus to just Tesla researchers Or the factory itself can open it and let the whole world explore humanoid research?
Paragraph 30
I think we have to be careful about Optimus being potentially used in ways that are bad Because that is one of the possible things to do So I think we would provide Optimus where you can provide instructions to Optimus But where those instructions are governed by some laws of robotics that you cannot overcome So not doing harm to others and I think probably quite a few safety related things with Optimus We'll just take maybe a few more questions and then thank you all for coming Questions, one deep and one broad On the deep for Optimus, what's the current and what's the ideal controller bandwidth? And then in the broader question, there's this big advertisement for the depth and breadth of the company What is it uniquely about Tesla that enables that? Anyone want to tackle the bandwidth question? So the technical bandwidth of the... Close to your mouth and loud For the bandwidth question, you have to understand or figure out what is the task that you want it to do And if you took a frequency transform of that task, what is it that you want your limbs to do? And that's where you get your bandwidth from It's not a number that you can specifically just say you need to understand your use case And that's where the bandwidth comes from What is the broad question?
Paragraph 31
The breadth and depth thing, I can answer the breadth and depth On the bandwidth question, I think we probably will just end up increasing the bandwidth Which translates to the effective dexterity and reaction time of the robot It's safe to say it's not one hertz and maybe you don't need to go all the way to 100 hertz But maybe 10, 25, I don't know Over time, I think the bandwidth will increase quite a bit Or translate it to dexterity and latency You'd want to minimize that over time Minimize latency, maximize dexterity In terms of breadth and depth, I guess we're a pretty big company at this point So we've got a lot of different areas of expertise that we necessarily had to develop In order to make electric cars and then in order to make autonomous electric cars Tesla is like a whole series of startups basically And so far they've almost all been quite successful So we must be doing something right And I consider one of my core responsibilities in running the company Is to have an environment where great engineers can flourish And I think in a lot of companies, I don't know, maybe most companies If somebody's a really talented driven engineer, they're unable to actually Their talents are suppressed at a lot of companies And some of the companies that the engineering talent is suppressed In a way that is maybe not obviously bad But where it's just so comfortable and you paid so much money The output you actually have to produce is so low that it's like a honey trap So there's a few honey trap places in Silicon Valley Where they don't necessarily don't seem like bad places for engineers But you have to say like a good engineer went in and what did they get out And the output of that engineering talent seems very low Even though there seem to be enjoying themselves That's why I call it there's a few honey trap companies in Silicon Valley Tesla is not a honey trap that we're demanding and it's like You're going to get a lot of shit done and it's going to be really cool And it's not going to be easy But if you are a super talented engineer Your talents will be used I think to a greater degree than anywhere else You know, SpaceX also that way Hi Ilan, I have two questions So both to the autopilot team So the thing is like I have been following your progress for the past few years So today you have made changes on like the lane detection Like you said that previously you were doing instant semantic segmentation Now you guys are built transfer models for like building the lanes So what are some other common challenges which you guys are facing right now Like which you are solving in future as a curious engineer So that like we as a researcher can work on those Start working on those And the second question is like I'm really curious about the data engine Like you guys have like told a case like where the car is stopped So how are you finding cases which is very much similar to that from the data which you have So a little bit more on the data engine would be great I'll answer the first question using occupancy network as an example So what you saw in the presentation did not exist a year ago So we only spent one year on time We actually shipped more than 12 occupancy network And to have a one foundation model actually to represent the entire physical world Around everywhere and you always condition is actually really really challenging So only over a year ago we're kind of like driving a 2D world If there's a wall and if there's a curve we kind of represent with the same static edge Which is obviously you know not ideal right There's a big difference between a curve and a wall when you drive you make different choices right So after we realized that we have to go to 3D We have to basically rethink the entire problem and think about how we address that So this will be like one example of a challenges we have we have a conquer in the past year Yeah to answer the question about how we actually source examples of the tricky stopped cars There's a few ways to go about this but two examples are one we can trigger for disagreements within our signals So let's say that parked bit flickers between parked and driving We'll trigger that back and the second is we can leverage more of the shadow mode logic So if the customer ignores the car but we think we should stop for it we'll get that data back too So these are just different like various trigger logic that allows us to get those data campaigns back Hi Thank you for the amazing presentation thanks so much So there are a lot of companies that are focusing on the AGI problem And one of the reasons why it's such a hard problem is because the problem itself is so hard to define Several companies have several different definitions they focus on different things So what is Tesla how's Tesla defining the AGI problem and what are you focusing on specifically Well we're not actually specifically focused on AGI I'm simply saying that AGI is seems likely to be an emergent property of what we're doing Because we're creating the oldies autonomous cars and autonomous humanoids That are actually with a truly gigantic data stream that's coming in and being processed It's by far the most amount of real world data and data you can't get by just searching the internet Because you have to be out there in the world and interacting with people and interacting with the roads And just you know it's Earth is a big place and reality is messy and complicated So I think it's sort of like it just seems likely to be an emergent property If you've got tens or hundreds of millions of autonomous vehicles and maybe even a comparable number of humanoids Maybe more than that on the humanoid front Well that's just the most amount of data and if that video is being processed It just seems likely that the cars will definitely get way better than human drivers And the humanoid robots will become increasingly indistinguishable from humans perhaps And so then like I said you have this emergent property of AGI And arguably humans collectively are sort of a superintelligence as well Especially as we improve the data rate between humans The thing like that seems way back in the early days the internet was like the internet was like humanity acquiring a nervous system Where now all of a sudden any one element of humanity could know all of the knowledge of humans by connecting to the internet Almost all the knowledge or certainly a huge part of it Whereas previously we would exchange information by osmosis Like in order to transfer data so you would have to write a letter Someone would have to carry the letter by person to another person And then a whole bunch of things in between and then it was like Yeah I mean it's insanely slow when you think about it And even if you were in the Library of Congress you still didn't have access to all the world's information And you certainly couldn't search it and obviously very few people are in the Library of Congress So I mean one of the great sort of equality elements Like the internet has been the biggest equalizer in history in terms of access to information and knowledge And any student of history I think would agree with this Because you know you go back a thousand years there were very few books And books would be incredibly expensive but only a few people knew how to read And even a small number of people even had a book Now look at it like you can access any book instantly You can learn anything basically for free It's pretty incredible So you know I was asked recently what period of history would I prefer to be at the most And my answer was right now This is the most interesting time in history and I read a lot of history So let's do our best to keep that going And to go back to one of the earlier questions I would ask The thing that's happened over time with respect to Tesla autopilot is that the neural nets have gradually absorbed more and more software And in the limit of course you could simply take the videos as seen by the car And compare those to the steering inputs from the steering wheel and pedals Which are very simple inputs And in principle you could train with nothing in between Because that's what humans are doing with the biological neural net You could train based on video and what trains the video is the moving of the steering wheel and the pedals With no other software in between We're not there yet but it's gradually going in that direction Alright, one last question How are you going? I think we've got a question at the front here Hello, they're right there We'll do two questions, fine They're here Thanks for such a great presentation We'll do your question last Okay, cool With FSD being used by so many people How do you evaluate the company's risk tolerance in terms of performance statistics And do you think there needs to be more transparency or regulation from third parties As to what's good enough and defining thresholds for performance across many miles The number one design requirement at Tesla is safety And that goes across the board So in terms of the mechanical safety of the car We have the lowest probability of injury of any cars ever tested by the government For just a passive mechanical safety Essentially crash structure and airbags and what not We have the highest rating for active safety as well And I think it's going to get to the point where the active safety is so ridiculously good It's just absurdly better than a human And then with respect to autopilot We do publish broadly speaking the statistics on miles driven With cars that have no autonomy Tesla cars with no autonomy With hardware one, hardware two, hardware three And then the ones that are in FSD beta And we see steady improvements all along the way And sometimes there's this dichotomy of Should you wait until the car is three times safer than a person before deploying any technology But I think that's actually morally wrong At the point at which you believe that adding autonomy reduces injury and death I think you have a moral obligation to deploy it Even though you're going to get sued and blamed by a lot of people Because the people whose lives you saved don't know that their lives are saved And the people who do occasionally die or get injured Definitely know, or their state does, that there was a problem with autopilot That's why you have to look at the numbers in total miles driven How many accidents occurred, how many accidents were serious, how many fatalities And we've got well over three million cars on the road So that's a lot of miles driven every day And it's not going to be perfect But what matters is that it is very clearly safer than not deploying it Yeah So, I think, last question I think, yeah, thanks The last question here Okay, hi So, I do not work on hardware So maybe the hardware team and you guys can enlighten me Why is it required that there be symmetry in the design of Optimus? Because humans, we have handedness, right? We use some set of muscles more than others Over time there's wear and tear, right? So maybe you'll start to see some joint failures or some actuator failures more Over time, I understand that this is extremely pre-stage Also, we as humans have based so much fantasy and fiction Over superhuman capabilities Like all of us don't want to walk right over there We want to extend our arms and like we have all these, you know A lot of fantasy, fantastical designs So considering everything else that is going on In terms of batteries and intensity of compute Maybe you can leverage all those aspects into coming up with something Well, I don't know, more interesting in terms of the robot that you're building And I'm hoping you're able to explore those directions Yeah, I think it would be cool to have like, you know, make Inspector Gadget real That would be pretty sweet So, yeah, I mean, right now we just want to make basic humanoid work well And our goal is to pass this path to a useful humanoid robot I think this will ground us in reality, literally And ensure that we are doing something useful Like one of the hardest things to do is to be useful To actually, and then to have high utility under the curve Like how much help did you provide to each person on average And then how many people did you help? The total utility Like trying to actually ship useful product that people like To a large number of people is so insanely hard It boggles the mind You know, that's why I can say like, man, there's a hell of a difference between a company that has shipped product And one has not shipped product This is night and day And then even once you ship product, can you make the cost, the value of the output Worth more than the cost of the input Which is, again, insanely difficult, especially with hardware So, but I think over time I think it would be cool to do creative things And have like eight arms and whatever And have different versions And maybe, you know, there'll be some hardware Like companies that are able to add things to an optimist Like maybe we, you know, add a power port or something like that Or attach them, you can add attachments to your optimist Like you can add them to your phone There could be a lot of cool things that could be done over time And there could be maybe an ecosystem of small companies that, or big companies that Make add-ons for optimists So, with that, I'd like to thank the team for their hard work You guys are awesome And thank you all for coming And for everyone online, thanks for tuning in And I think this will be one of those great videos where you can like If you can fast forward to the bits that you find most interesting But we try to give you a tremendous amount of detail Literally so that you can look at the video at your leisure And you can focus on the parts that you find interesting and skip the other parts So, thank you all, and we'll do this, try to do this every year And we might do a monthly podcast even So, but I think it'll be great to sort of bring you along for the ride And like show you what cool things are happening And yeah, thank you Alright, thanks Thank you