Andrew Feldman
Analyst · UBS
Thank you, Sean. Thank you all for joining us today. Q2 was a strong quarter. We completed our public offering, but we did not let that distract us from execution. We delivered record core revenue and beat guidance on all metrics; core revenue, core gross margins and core operating margin. Looking forward, we see unbound demand for fast inference. The market is realizing that speed is not a benchmark item. Speed changes user engagement, it changes agentic performance, and it changes AI productivity. Fast inference unlocks new applications and new markets. As we've shared with you previously, 2026 is a foundation building year for Cerebras. We've made excellent progress on multiple fronts in the past 7 weeks since our last earnings call, preparing us for a massive 2027, 2028 and 2029 as we deliver on the $25 billion of RPO we currently have on our books. With the benefit of that progress, we expect to more than triple our core revenues in '27 and continue to grow at multiples in the years following. We think of progress in terms of capacity, capabilities and customers. We're expanding capacity by adding new contracts for data centers around the world, expanding manufacturing capabilities and collaborating with our vendors to ensure supply and to support our extraordinary growth. We're advancing our capabilities by inventing new technology that extends our performance and throughput and our power efficiency. And we're expanding our customer base by accelerating AI productivity in existing markets like coding and agentic flows and pioneering new areas like security, where speed opens up entirely new opportunities. On the capacity front, data center space continues to be the bottleneck for the entire industry, and we are no exception. The faster we and our customers bring on new data centers, the faster we grow. So over the last 7 months, we've had an all-out push to secure and build out data centers. We have 2 advantages. First, because we are serving inference, we do not need gigawatt footprint locations like those needed for training clusters. This gives us much more flexibility to scale up capacity across a multitude of locations around the world. Second, we built a repeatable process for site selection, cluster deployment and customer activation, which is an operational muscle required to turn gigawatts into production tokens at a global scale. I'm pleased to report that our push has been very successful. We now have data centers either up or under contract in Alabama, Dallas, Denver, Minneapolis, Santa Clara, Stockton and outside the U.S. in France, Finland, Manitoba, Montreal, Norway, Saskatchewan and Toronto. In total, over the last 7 months, we have secured more than 600 megawatts of data center capacity that is either live now or will be delivered by the end of 2027. And while this isn't nearly enough to meet our demand, our data center pipeline of new opportunities for expansion continues to grow and is now measured in gigawatts. To put it in perspective, as we continue to build out our first-party cloud, it will be among the largest non-hyperscale AI clouds. And whereas at the end of 2025, we were on a steep learning curve, today, I'm happy to report that we're pretty good at data center build-out with a clear path to becoming excellent. Other key dimensions of capacity include manufacturing and supply chain. Here, we've successfully increased our manufacturing capacity and are building up new factories with Flex and Sanmina and expect to increase our manufacturing capacity by more than 10x in 2026 and continue that expansion in 2027, again, preparing us for the exceptional growth expected in the years ahead. Our partnership with our supply chain vendors has also turned into a significant advantage. TSMC has once again come through, and we have the wafers needed to fuel our growth. Our ability to get wafer supply also benefits from the fact that we were able to deliver industry-leading performance while running on TSMC's 5-nanometer node, where wafers are less expensive and supply is less constrained. Our decades-long relationships with our supply partners reinforces our confidence that we can deliver on our growth plans going forward. These relationships are rare and valuable, particularly in times of short supply. Finally, recall that most of the critical supply chain constraints currently faced by the industry don't apply to us. For example, we don't use HBM memory, CoWoS packaging or require 3-nanometer fab capacity. On the capabilities front, in the second quarter, we delivered support for OpenAI's GPT-5.6 Sol, the largest and most capable of the frontier models. In fact, Cerebras serves 5.6 Sol at a speed that is 10x faster. With GPT-5.6 Sol, this lays to rest any of the remaining concerns regarding our ability to support large frontier models. Being a partner for the delivery of GPT-5.6 Sol and serving it to our cloud speaks to the maturity of our software stack. It takes millions of system hours of production hardening to get to the point where one can deliver hyperscale quality and reliability. We're proud that our inference cloud can meet the requirements of the most demanding customers. Our collaboration on serving models at the Frontier has opened up new and significant strategic advantage previously only available to NVIDIA. Closed-source frontier models include a continual stream of new insights and new AI techniques. Serving these models allows us to see into the future and to prepare for it. Our road map from the hardware through the software stack now reflects what we're seeing and will give us a compounding advantage in the years to come. Continuing on the theme of capabilities, let's turn to disaggregation. We now have disaggregated inference solutions with 2 of the leading chip companies, AMD with their Helios and AWS with Trainium. Disaggregation expands the market for both the GPU provider and for Cerebras. Disaggregation enables GPUs to participate in a market currently foreclosed to them, namely fast inference. Disaggregation enables Cerebras to expand our opportunity to those customers who are more price sensitive and expands the profitability of our data centers. Let's see how this works. As with any compute market, as inference grows and matures, opportunities for specialization emerge. Disaggregation is a form of specialization that is particularly well suited for workloads with well-known traffic patterns. In these cases, disaggregation delivers advantage by separating inference into 2 stages, prefill and decode and using different processors for each stage. Prefill processes the input from the user or agent. It is a parallelizable workload. As a result, prefill is well suited for GPUs and their HBM-based memory architectures. Decode generates the output tokens. It's the harder technical problem and is the bulk of the computational work in a disaggregated solution. It is sequential and memory bandwidth intensive and is particularly well suited for our wafer scale engine. The prefill and decode processors need to be linked to create the end-to-end solution. And this is where standards-based I/O and open engagement strategy has made integration easy and straightforward for Cerebras. A few weeks ago, we announced a partnership with AMD to build disaggregated inference solutions. The solutions combine their Helios racks with our CS systems. The combined solution maintains Cerebras speed while increasing throughput by 5x. To understand how powerful this is, it's important to understand the difference between speed and throughput. Speed is a measure per user. It's measured in tokens per second per user. It is how fast your query is answered or how long it takes an agent to finish a task. Here, it is on the X-axis. Throughput, on the other hand, is the total number of tokens the solution can produce per second. It is measured by adding up all the tokens across all the simultaneous users. Here, it is shown as it's generally done on the Y-axis. Speed is critical for user experience. Throughput is critical for inference economics. GPU solutions can support high throughput, but only at low speeds. When configured to support even moderate speeds, GPU throughput drops precipitously. This is true not just for GPUs, but also for ASICs and all solutions that use HBM. The HBM memory architecture forces a trade-off between throughput and speed. SRAM-based architectures like Cerebras are the exact opposite. We support blisteringly fast tokens, but at moderate throughput. So GPUs want to get faster without giving up throughput. Cerebras wants more throughput without giving up speed. Herein is the strength of our disaggregated solution. It delivers Cerebras speed with 5x higher throughput. Increasing throughput by 5x while keeping our industry-leading speed has a profound impact on the economics of token generation. It means up to 5x as many high-speed, high-value tokens are made by each Cerebras system. More tokens per system at lower cost means more revenue and more gross margin. More tokens generated per CS system also means more tokens per watt, making each data center more profitable. Perhaps most important in a data center constrained environment, the disaggregated solution allows us to serve more of the demand that we have in RPO. Finally, we believe this disaggregation approach makes performance and economic sense with any GPU. For operators who have already deployed large footprints of GPUs, disaggregation with Cerebras offers them an opportunity to create meaningful leverage built on their existing investments by pairing some portion of those GPUs with Cerebras solutions, dramatically improving the value and usefulness of their data center footprint. Continuing on the capabilities theme, let's turn to our road map. Our engineering execution is continuing at pace. We expect to deliver new systems that double our speed each year for the next several years. Remember, we're doubling our performance starting with a 15x performance advantage over everyone else in the industry. In addition, while keeping the performance crown over the next 18 months, we plan to deliver solutions that increase throughput by more than 20x. Next week at our annual Supernova Conference, we will be unveiling the CS4, our fourth-generation system. It will be a great event with lots of product announcements, so I recommend you attend. Finally, we are currently on track to launch our CS5 in the second half of 2027. Looking even further out, our invention engine is humming. We have significant partnerships with the U.S. government for delivery of stacked memory solutions as well as integrated wafer-scale optical solutions. In the years ahead, you can expect to see inventions from us in chip and chip architecture as well as all elements of system design, including packaging, I/O and power delivery. To summarize the capability section, we expect to continue to deliver pioneering advances in product and technology to drive up speed and throughput, reduce the power use per token and slash the cost per token of our solution. Now let's turn to the customer front. Fast tokens are in demand and command a premium at market and fast tokens with frontier intelligence are only available through OpenAI Cerebras partnership. Our work with AWS continues, and we expect to have solutions generally available in Q1 2027 through AWS' Bedrock platform. This AWS partnership expands our market opportunity and provides us with global reach through an industry leader who is trusted by nearly every enterprise in the world. Our discussions with other hyperscalers are also going well. We expect to produce first revenue starting in mid-2027 and ramp through 2028 and beyond. And with all of this progress, I think it's important to keep in mind that our $25 billion in RPO does not reflect any backlog of business from AWS or any other hyperscaler at this time. Our business outside of OpenAI and the hyperscalers continues to grow nicely. For example, in Q2, we signed 6 deals north of $30 million. AI coding continues its rapid rate of growth. In our experience, no one says, I'm happy with slow tokens when coding. So not surprisingly, in the coding category, our footprint continues to grow. We signed new agreements with public companies such as Figma and startup leaders such as Cognition, and we extended our presence in Europe, a major win at Lovable. Agentic flows are growing quickly and the value of speed compounds as agentic operations rapidly evolves toward multistep, multi-agent solutions. Companies as diverse as Block, AlphaSense and GSK signed new agreements during the second quarter with Cerebras to leverage fast inference to provide their custom-made agents. Fast AI also opens up new markets, extending the TAM for Cerebras. Security is one such example. Our recent win with CrowdStrike is an application that only exists if AI is fast. Fast AI enables AI-based security devices to sit in line with enterprise traffic and use LLMs to secure traffic so quickly that nobody notices. Fast AI enables an LLM to provide security that is invisible to users. The AI provides the security. The speed creates the invisibility that enables the security to avoid delay and disruption. We expect this type of security to become the norm given the rapidly evolving threat landscape. Enterprises will soon expect vast swaths of their traffic to be inspected in this way, creating massive new opportunities made possible exclusively through fast AI. Frontier labs, hyperscalers, leading chip makers, the fastest-growing start-ups and massive enterprises are all now customers and partners of Cerebras and benefit from our blazing fast inference. To summarize, overall, a strong quarter. We went public in a successful IPO. We beat on all metrics; core revenue, core margins and core operating margins. We made progress in each of our key domains; capacity, capability, and customers. These are the foundations on which we'll achieve our goals of massive growth in 2027 and 2028 and continue this exceptional rate of growth in '29 and beyond. And with that, I'll turn things over to Bob. Bob?