US Firms Adopt Chinese AI Models, Facing Integration Hurdles
Key Takeaways
Silicon Valley startups are replicating Chinese open-source architectures and aggressive pricing strategies to compete. However, while China scales AI agents in enterprise workflows, US firms face funding shortages and regulatory barriers, signaling a his
Woofun AI reports that a strategic reversal has emerged in the global artificial intelligence landscape, with Silicon Valley entities increasingly adopting Chinese open-source frameworks.
This shift is epitomized by Mira Murati’s Thinking Machines Lab, which recently released its inaugural model, Inkling. The technical documentation for Inkling explicitly states that its underlying architecture is derived from DeepSeek-V3, while its subsequent training data was generated by Kimi K2.5. This indicates that some of the most affluent and influential actors in Silicon Valley are now utilizing the foundational structures and synthetic data produced by Chinese open-source models. Just two years ago, the direction of technological dependency would have been reversed, with Chinese firms looking to US innovations. Now, the flow of architectural inspiration and data generation strategies has inverted, marking a significant pivot in the industry's core development patterns.
The imitation extends beyond mere architectural borrowing to include the aggressive pricing strategies pioneered by Chinese companies. Earlier this year, Arcee AI’s Trinity-Large-Thinking model achieved a score of 91.9 points on the PinchBench benchmark. This performance was only 1.4 points lower than Anthropic’s Opus 4.6, yet the cost of using Trinity-Large-Thinking was merely 4% of the price of Opus 4.6. Mark McQuade, CEO of Arcee AI, stated, "We’re striving to help the U.S. catch up to and surpass China."
This sentiment reflects a broader industry recognition that the competitive advantage lies not just in model weights but in cost-efficiency and accessibility. The architecture can be replicated, training methods can be copied, and data can be generated in bulk to train other models. Once weights are made public, developers can build upon them, and pricing strategies inevitably follow suit as competitors attempt to match the value proposition. The traditional model of success, built on large parameter counts, high hash rates, and closed-source exclusivity, is crumbling under the pressure of these new, cost-effective alternatives.
A pivotal moment in this shift occurred on July 17, when Moonshot AI released Kimi K3. This model boasts 2.8 trillion total parameters and a context length of 1 million tokens, establishing it as the largest open-source model in the world. On the Frontend Code Arena rankings, Kimi K3 scored 1,679 points, outperforming both Claude Fable 5 and GPT-5.6 Sol. This marked the first instance where an open-source model surpassed closed-source counterparts in performance metrics, taking first place in six out of seven frontend categories. Elon Musk commented on the results, writing, "Impressive." Despite its high performance, Kimi K3’s API is not inexpensive, charging 100 yuan per million output tokens.
However, in SuperCLUE’s tests, its task scores were 18% higher than second-place models, while its average cost per task was 16% lower. This disparity highlights that the true measure of a model’s value is not the price per million tokens but the total cost to complete a task from start to finish. The wider the gap between performance and cost, the less meaningful the nominal price tag becomes.
The immediate impact of Kimi K3’s release was a strain on infrastructure, with Moonshot AI suspending new C-end subscriptions just two days later due to request volumes exceeding expectations within 48 hours. On July 27, the full weights of Kimi K3 were made public, along with technical reports and supporting infrastructure. Shortly thereafter, DeepSeek entered the fray with the official release of V4 Flash on August 2. This model features 284 billion total parameters, with only 13 billion activated per inference. Its performance on the DeepSWE programming test rose dramatically from 7.
3 to 54.4. The pricing strategy was equally aggressive; in May, V4-Pro’s price had already dropped to one-fourth of its original level, marking the fourth price cut by DeepSeek in a single month. When V4 Flash launched, its cost-effectiveness surprised the market, reinforcing the trend that Chinese models are becoming increasingly competitive while their prices continue to fall. Foreign competitors find themselves in a difficult position, as users have already experienced affordable, high-performance options and are reluctant to return to premium pricing models.
The disruption of global market norms continued as Alibaba announced it would release the weights for Qwen3.8-Max within 16 days. The Max series, previously reserved for closed-source flagship products, is now opening up to the public. Its international price is approximately 40% of Opus 5’s input cost and only 24% of its output cost. This rapid succession of releases and price cuts demonstrates a coordinated shift away from scarcity pricing, where companies relied on exclusive capabilities to command high prices.
Instead, firms like Kimi, DeepSeek, and Qwen are pursuing a strategy of improving performance, making weights available, lowering prices, and quickly integrating models into products. This approach undermines the traditional business model of large model companies, which previously leveraged small differences in rankings to justify huge price gaps. The goal is no longer just to have the best model but to have the most accessible and integrated solution.
Woofun AI data shows that despite technical success and government ties, US startups like Arcee AI face significant funding struggles. By the end of 2025, Arcee AI had invested almost all its funds in development, spending $20 million on 2,048 B300 chips over 33 days with a team of just 30 employees. The resulting Trinity Large model had 400 billion total parameters and 13 billion activated per inference, with nearly half of its training data being synthetic. In July, Arcee signed a partnership with the U.S. Department of Energy to develop a model for scientific research.
Despite having a strong team, impressive results, low costs, and government contracts, Arcee struggled to secure additional funding. To date, it has raised only $50 million, with a valuation of $240 million. In the current competitive landscape, such amounts are insufficient to cover basic operational costs. Mark McQuade noted that "Almost all top-tier VCs rejected us," indicating that US open-source models are being held back by their own investment committees before they can even compete directly with Chinese models.
The bias against open-source models in venture capital is evident in the funding trends of the first quarter of 2026, where global AI startups raised $255.5 billion. Nearly two-thirds of this capital went to just three deals: OpenAI’s $122 billion, Anthropic’s $30 billion, and xAI’s $20 billion. Almost all the money flowed into closed-source companies. Joe Floyd, an early investor in Arcee and partner at Emergence Capital, reported hearing similar rejection reasons from many VCs: "I don’t want this model to succeed. I refuse to invest because it could undermine my investments in Anthropic and OpenAI." In contrast, chip manufacturers like NVIDIA are willing to support open-source initiatives.
In March, NVIDIA disclosed in SEC filings that it planned to invest $26 billion over the next five years to support open weights. Reflection AI, Poolside, and Thinking Machines Lab all received funding from NVIDIA. Jensen Huang understands that closed-source models make money from APIs, while open-source models drive GPU sales. The cheaper the models, the more developers use them, boosting NVIDIA’s revenue. On July 24, Jensen Huang created an X Corp account and posted an open letter titled "Open Weights and America’s Leadership in AI," arguing that US leadership depends on an ecosystem that can spread across various sectors.
Political pushback against restrictions on Chinese open-source models has also emerged, with founders of nearly 200 US startups jointly writing to Trump’s administration. The letter opposed restrictions on Chinese open-source models, arguing that affordable Chinese models have become practical tools and that restrictions would force companies to buy expensive US APIs instead. While Silicon Valley debates the adoption of open weights, China has moved forward with the integration of AI agents.
The phrase "the beginning era of AI applications" has been repeated for three years, but true application phase entry is determined by user willingness to pay and corporate integration. This year, Claude Code and Cursor both achieved annual revenues of over $1 billion, with AI coding tools gaining near 50% penetration among individual developers. On the corporate side, Sandhill Research found that the adoption rate of AI agents among Chinese companies rose from 17.
3% at the end of 2024 to 40.3% by mid-2026. Robin Li introduced the metric DAA (daily active agents) to track this phenomenon, shifting focus from daily active users to the number of agents working on tasks. According to CCTV, China’s average daily token usage rose from around 100 billion at the beginning of 2024 to 140 trillion by the end of March, a more than thousand-fold increase in just over two years.
As models become cheaper, compute scarcity has intensified. In March, Tencent Cloud, Alibaba Cloud, and Baidu Smart Cloud successively raised their compute prices by about 30% within ten days. When AI agents became popular in the first half of the year, rental prices for inference computing power rose by over 40%. This is because the way people use AI has changed; agents need to handle context, use multiple tools, verify results, and restart if mistakes are made, leading to dramatic increases in token consumption.
Price cuts have not shrunk the market but have allowed previously unaffordable demands to be met. In 1981, IBM opened up the PC architecture, leading to a surge of compatible machines and falling hardware prices. The biggest profits were reaped not by machine manufacturers but by companies like Lotus 1-2-3, WordPerfect, and later Microsoft, which thrived on top of those machines. Today, the pattern is similar. While people focus on minor differences in underlying capabilities, companies above that level are competing for users, entry points, and workflows.
US companies can study DeepSeek’s architecture or use data generated by Kimi, but applications are not easy to copy. There are no downloadable weights for applications; success depends on expertise accumulated over years. Models are integrating into apps like DingTalk, Lark, and WeCom in China, and Gmail, Slack, and Salesforce in the US. From 2022 to 2025, the US imposed strict chip restrictions on China, banning H100 and B300 chips. Despite this, Chinese companies developed the world’s largest open-source model.
By 2026, the US began discussing restrictions on Chinese model weights, with OpenAI and Anthropic expressing concerns to regulators and Congress examining whether model weights can continue to cross borders. Insiders complain that chip bans restrict supply, but model bans first hit the costs of US startups. The joint letter clarified that restrictions cannot stop dissemination; developers will download all available models before restrictions take effect. Chips are physical products that can be blocked, but weights can be downloaded, mirrored, quantized, fine-tuned, and distilled infinitely.
The era of "Innovation in US, Scale in China" is ending, replaced by a new dynamic where Chinese companies are redefining the upper end of the technology chain. For thirty years, China has won by adapting technologies invented elsewhere—mobile payments, e-commerce, food delivery, live streaming—to meet the needs of over a billion people. The underlying infrastructure was provided by others: transistors from Bell Labs, x86 from Intel, TCP/IP funded by the US government, Windows, Android, and iOS from the West Coast, and TensorFlow and PyTorch from Google and Meta. Now, an exception has appeared.
Thinking Machines Lab’s technical documentation shows that a prominent Silicon Valley company used a Chinese company’s architecture and data generated by another Chinese model. Silicon Valley is starting to directly follow technical paths proven by Chinese companies. While Silicon Valley has long dominated the upper end of the technology chain, from chips to IDEs, the results matter. Whoever can propose new architectures, reduce costs, and inspire imitation has the right to redefine the chain. Chinese companies are now appearing on this list. Silicon Valley may catch up eventually, with plenty of engineers, capital, and chips, but applications integrated into real businesses will not wait.
Whoever integrates first into corporate workflows will gain access to customers, data, and opportunities for further improvement. The starting line was never the same. The era of AI applications will not begin with a grand ceremony but will appear in financial statements, soaring usage volumes, and programmers running out of quota. Widespread adoption is only a matter of time, as noted by Mark McQuade. This marks a historic reversal in tech leadership, where the race shifts from model weights to enterprise integration, and China’s scale in application integration challenges US dominance in foundational innovation.
Comments
No comments yet.