Morgan Stanley Models 25%-50% AI Returns via GPU Rentals and Token APIs
Key Takeaways
Morgan Stanley projects 25%-50% capital returns on AI infrastructure by 2028, analyzing three monetization paths: IaaS GPU rentals, proprietary Model APIs, and third-party compute leasing. The report highlights Microsoft, Amazon, Google, and Meta as key b
Woofun AI reports that Morgan Stanley has released a comprehensive financial model detailing the potential profitability of Generative AI infrastructure, identifying Microsoft, Amazon, Google, and Meta as the primary entities positioned to capture value. The analysis dissects the capital expenditure landscape, arguing that despite massive upfront investments in hardware and data centers, these hyperscalers can achieve substantial returns if specific operational thresholds are met. This framework addresses investor skepticism regarding whether current spending will yield tangible profits or merely inflate depreciation and energy costs. The core thesis rests on the ability of major tech platforms to leverage their existing customer bases and distribution channels to monetize idle compute power through structured service offerings.
The report, dated July 27, establishes a baseline projection that AI infrastructure investments can generate a 25%-50% capital return under optimal conditions. This range serves as a critical benchmark for evaluating the efficiency of current spending trends. The calculation assumes that key variables such as GPU utilization rates, rental pricing, token throughput speeds, and API pricing structures remain stable at assumed levels. By defining this return corridor, Morgan Stanley provides a quantitative anchor for assessing the viability of the ongoing build-out. The projection implies that the industry is moving beyond speculative spending toward a phase where unit economics must justify continued capital allocation.
This shift marks a transition from growth-at-all-costs to disciplined return-on-investment monitoring.
Monetization is categorized into three distinct pathways, each with varying margins and risk profiles. The first path involves renting out GPU compute as Infrastructure as a Service (IaaS), leveraging existing data center assets. The second path focuses on providing Model APIs through proprietary infrastructure, where revenue is tied to token consumption rather than raw compute hours. The third path entails renting third-party compute to offer Model APIs, a model more susceptible to margin compression due to external rental costs. These pathways are not mutually exclusive; however, their suitability depends on a company’s asset base. Large tech platforms with integrated ecosystems are better suited for the first two paths, while pure-play model companies may rely more heavily on the third. The structural advantage lies in vertical integration, which reduces dependency on external compute providers.
The scale of this infrastructure expansion is unprecedented, with Morgan Stanley projecting that hyperscaler capital expenditures will exceed $1.4 trillion by 2028. This financial commitment underscores the strategic importance of AI in the broader technology sector. Alongside the financial metrics, physical capacity is expected to double from 2025 levels to approximately 120GW by 2028. This surge in capacity creates a significant challenge: ensuring that newly deployed hardware is fully utilized rather than sitting idle. The gap between capacity deployment and revenue generation is bridged by inference workloads, enterprise software adoption, and cloud service integration. If demand does not keep pace with supply, the high fixed costs of depreciation and energy will erode profitability. Therefore, the timeline from 2025 to 2028 is critical for establishing sustainable utilization rates.
In the first monetization path, IaaS unit economics are modeled around a baseline of 1GW of capacity, which corresponds to approximately 410,000 NVIDIA GB300 GPUs. The model assumes a 75% utilization rate and an hourly rental rate of $8.5 per GPU. Under these conditions, the GPU rental business generates approximately $22.9 billion in revenue per GW-year. The incremental EBIT margin is projected at about 67%, yielding a capital return of around 31%. These figures suggest that IaaS is not merely a low-margin hardware rental business but a high-utilization infrastructure operation. The profitability hinges on maintaining high occupancy rates and sustaining rental prices. If demand softens or competition drives prices down, the margin profile could deteriorate rapidly. The advantage for major cloud providers lies in their existing relationships and platform integration, which facilitate direct revenue conversion.
Woofun AI data shows. The second path, proprietary Model API economics, offers higher potential returns due to value-added services. The baseline scenario assumes that 65% of capacity is dedicated to inference, with each GPU processing approximately 2,750 tokens per second. Pricing is set at $1.75 per million tokens, reflecting a blended rate across different model tiers. In this scenario, the Model API’s incremental EBIT margin reaches about 75%, with a capital return ranging from 40%-46%.
This superior margin structure arises because revenue is linked to token consumption, which scales with user engagement and application complexity. Unlike raw compute rentals, API pricing captures the value of model intelligence and optimization. The trend indicates that embedding AI capabilities into existing products, such as search and office software, creates a high-frequency, billable revenue stream. This approach allows companies to monetize inference more effectively than simple hardware leasing.
The third path involves third-party compute leasing, where model companies rent infrastructure to provide APIs. In the baseline scenario, GPU rental costs are approximately $7.75 per hour. This cost structure results in an incremental EBIT margin of about 31%, leading to an after-tax ROI or profit margin of around 25%. While still profitable, this path is less lucrative than owning infrastructure because the renter must cover both compute costs and model development expenses.
The narrower margin reflects the competitive pressure on pure-play model companies, which lack the vertical integration of hyperscalers. These entities are more vulnerable to fluctuations in rental prices and competition from open-source models. The 25% return serves as a floor for profitability, but it requires efficient model design and strong pricing power to maintain. Companies without proprietary infrastructure must compete on model quality and application differentiation to survive.
Morgan Stanley maintains favorable ratings for the key beneficiaries of this infrastructure build-out. Microsoft is rated Overweight with a target price of $600, reflecting its strong position in both cloud infrastructure and enterprise software. Alphabet, the parent company of Google, has a target price of $400, driven by its search and cloud dominance. Meta is listed as a Top Pick with a target price of $775, highlighting its aggressive AI integration in advertising and social platforms.
Amazon is also included in the beneficiary group, though specific target prices vary across public sources. These ratings underscore the belief that these companies can leverage their scale and customer access to monetize AI investments. The stock valuations incorporate expectations of future cash flows from inference and API services. Investors are pricing in the potential for these firms to convert capital expenditures into sustained earnings growth.
Risks to these projections are concentrated around utilization thresholds and competitive dynamics. The model assumes a 75% utilization rate, which is a high bar that may not be met if enterprise AI adoption slows. Idle GPUs directly drag down returns, as fixed costs continue to accrue without corresponding revenue.
Additionally, the rise of open-source models and price wars among cloud providers could compress token pricing. If the revenue per million tokens declines, the profit margins for both proprietary and third-party API providers will erode. The sensitivity of returns to these variables means that small deviations in utilization or pricing can significantly impact profitability. Companies must continuously optimize their models and infrastructure to maintain efficiency. The threat of ASICs and next-generation GPUs further complicates the landscape, potentially reducing unit compute costs and altering competitive advantages.
Ultimately, the 25%-50% return projection serves as a yardstick rather than a guaranteed outcome. The model spreadsheet illustrates the potential upside if hyperscalers can maintain high utilization and effectively monetize token consumption through APIs and cloud services.
However, if rental rates fall, utilization drops, or inference adoption lags, the projected returns will remain theoretical. The divergence between the model’s assumptions and market reality will determine the actual financial performance of these investments. This analysis highlights the critical importance of operational discipline in the AI infrastructure cycle. As the industry matures, the focus will shift from capacity building to efficiency and monetization. The companies that succeed will be those that can bridge the gap between massive capital expenditure and tangible revenue generation.
Comments
No comments yet.