Original title: "Silicon Valley Copying China's Homework, Next Question is AI Applications"
Original author: Dongcha Beating
Silicon Valley has started copying China's homework.
Recently, a batch of low-cost open models has emerged in the U.S., aiming to catch up with China's open-source models, with prices also pushed down to the same level.
Mira Murati's Thinking Machines Lab released its first model, Inkling, last month. The official description mentioned two things: the underlying architecture references DeepSeek-V3, and the training data was generated by Kimi K2.5.
The most expensive people in Silicon Valley are using the skeleton of Chinese open-source models, fed with data generated by Chinese models. Two years ago, this would have happened in reverse.
Earlier this year, Arcee AI's Trinity-Large-Thinking scored 91.9 on PinchBench, just 1.4 points lower than Anthropic's Opus 4.6, but priced at only 4% of the latter.
Arcee AI's CEO Mark McQuade said, "We are working hard to let the U.S. catch up and surpass China."
What Silicon Valley is copying now is not just models, but also the prices set by Chinese companies.
Architectures can be referenced, training methods can be replicated, and data can be generated in bulk by other models. Once the weights are released, developers can continue to develop. Pricing is simpler; as long as one company cuts prices, the rest will eventually have to follow.
In the past three years, large model companies have built towering structures on parameters, computing power, and closed-source models, but now this building is about to collapse.
### Sixteen Days
In the early hours of July 17, Moonlight released Kimi K3. With a total of 2.8 trillion parameters and a context of 1 million tokens, it is the largest open model in the world.
On the Frontend Code Arena front-end programming leaderboard, K3 scored 1679 points, followed by Claude Fable 5 and GPT-5.6 Sol. This is the first time an open model has ranked ahead of closed-source models. It took first place in six out of seven front-end subcategories.
Elon Musk left a comment under the evaluation results saying, "Impressive."

K3's API is not cheap, charging 100 yuan per million output tokens. However, in the horizontal testing of SuperCLUE, its task score was 18% higher than the second place, and the average cost for completing the same batch of tasks was actually 16% lower.
**Whether a model is expensive or not should not be judged by the price per million tokens, but by how much it costs to complete a task from start to finish.** The further apart these two numbers are, the less meaningful the pricing becomes.
Two nights later, Moonlight suspended new user subscriptions for the C-end. The request volume in 48 hours far exceeded expectations, and computing power was insufficient.
On July 27, K3 released its complete weights, along with technical reports and supporting infrastructure.
Next up was DeepSeek.
The market had been waiting for the official version of V4. At the end of April, DeepSeek released two preview versions, V4-Pro and V4-Flash, both supporting 1 million tokens of context, which the official referred to as the "Era of Million Context Accessibility." After two or three months of silence, on August 2, V4 Flash was officially released and weights were opened. With a total of 284 billion parameters, only 13 billion are activated for a single inference, and the score in DeepSWE programming tests jumped from 7.3 to 54.4.
The price drop was even more drastic. In May, V4-Pro was directly reduced to a quarter of its original price, marking the fourth price adjustment by DeepSeek in a month. When Flash was officially launched, everyone was still surprised by its cost-performance ratio.
"You get what you pay for" was the old rule. Now, Chinese models are becoming increasingly competitive while charging less and less. Foreign competitors are certainly feeling the pressure, but there’s not much they can do. Users have already seen cheaper and more effective options, and no one wants to go back to being overcharged.
Subsequently, Alibaba announced it would release the weights for Qwen3.8-Max next week. The Max series, which had always held a closed-source flagship position, is also starting to open up. Its international price is about 40% of Opus 5's input and only 24% of its output.
These events occurred within a short span of 16 days.
In the past, model companies relied on scarcity pricing, where only they could produce, and you had to pay their price. A few points difference on the leaderboard could lead to price differences of dozens of times.
Kimi, DeepSeek, and Qwen have different technical routes, but they are all doing the same thing: pushing capabilities up, opening weights, cutting prices, and quickly integrating them into products.
### Want to Copy Homework, But No One is Paying
People in the U.S. have also begun to understand this trend.
By the end of 2025, Arcee AI had nearly all its funds invested. In 33 days, they acquired 2048 B300 chips for 20 million dollars, with only 30 people in the entire company. The resulting Trinity Large has a total of 400 billion parameters, activating 13 billion for a single inference, with nearly half of the training data being synthetic. In July of this year, Arcee signed a collaboration with the U.S. Department of Energy to develop a model aimed at scientific research.
The team, achievements, costs, and government orders are all in place. Yet, the door to financing remains firmly shut.
Arcee has raised a total of 50 million dollars, with a valuation of 240 million dollars. In today's large model race, this amount of money is not enough to even get a seat at the table.
CEO Mark McQuade stated, "Almost all top-tier VCs have rejected us."
The U.S. open models have not yet faced off against China directly, but have already been blocked by their own investment committees.

In the first quarter of 2026, global AI startups raised 255.5 billion dollars, with nearly two-thirds concentrated in three deals: OpenAI's 122 billion, Anthropic's 30 billion, and xAI's 20 billion. Almost all the money flowed into closed-source companies.
Joe Floyd, an early investor in Arcee and a partner at Emergence Capital, said he has heard similar rejection reasons from many VCs:
"I do not want this model to succeed. I do not want to invest because it would undermine my investments in Anthropic and OpenAI."
The ones truly willing to spend money on open models are actually chip sellers.
In March of this year, Nvidia disclosed in SEC filings that it plans to invest 26 billion dollars over the next five years to support open weights. Reflection AI, Poolside, and Thinking Machines Lab have all received funding from them.
Jensen Huang understands this better than anyone. **Closed-source models make money from APIs, while open models burn cards.** The cheaper the models are sold, the more developers use them, and the better Nvidia's cards sell.
On the night of July 24, Jensen Huang registered an X account, and his first tweet was a retweet of an open letter titled "Open Weights and America's Leadership in AI." The letter stated that whether the U.S. can continue to lead this industry cannot just be judged by having the most advanced models, but also by having an ecosystem that can spread across various industries.
Later, nearly 200 founders of U.S. startups jointly wrote to the Trump administration, opposing the ban on Chinese open models. The most interesting part of this letter is that it contains no flattering words. They did not speak well of China, nor did they wave the flag of openness; their reason was simply that the cheap Chinese models have become production tools, and if they were banned, they would have to go back to buying expensive U.S. APIs.
McQuade also mentioned in an interview that large-scale popularization is just a matter of time.
Demand Created by Low Prices
Silicon Valley is still debating whether to pursue open weights, while China has already begun to break through to the next level.
The phrase "Year of AI Applications" has been shouted many times over the past three years. However, to determine whether a technology has truly entered the application phase, two things must be considered: whether users are willing to spend money and whether companies are integrating it into their business.
This year, both Claude Code and Cursor have annual revenues exceeding 1 billion dollars, with the penetration rate of AI programming tools among individual developers approaching half.
The changes are even more significant on the enterprise side. According to research by the Dune Intelligence Institute, the adoption rate of AI Agents among Chinese enterprises rose from 17.3% at the end of 2024 to 40.3% by mid-2026.
Li Yanhong proposed a new metric called DAA, Daily Active Agent Count. Previously, internet companies counted how many people opened their products daily; in the future, they will count how many Agents are working for people daily and how many results they actually deliver.
The usage curve is even more astonishing. According to CCTV's statistics, China's average daily token usage rose from about 100 billion at the beginning of 2024 to 140 trillion by the end of March this year, more than a thousand times in just over two years.
Models are becoming cheaper, but computing power is becoming increasingly scarce. In March this year, Tencent Cloud, Alibaba Cloud, and Baidu Smart Cloud successively raised computing power prices within ten days, with an increase of about 30%. As the demand for intelligent agents exploded in the first half of the year, the rental price for inference computing power surged by over 40%.
The reason lies in the fact that the way people use AI has changed. Previously, asking one question and getting one answer did not consume many tokens. Now, for an Agent to complete a task, it needs to carry context, adjust several tools, check back and forth, and if it makes a mistake, it has to start over. The unit price has decreased, but the consumption has increased exponentially.
Price cuts have not shrunk the market; instead, they have brought in previously unaccounted demand.
In 1981, IBM opened up the PC architecture, and compatible machines quickly flooded the market, driving hardware prices down. Ultimately, the real profits were not made by the manufacturers of the machines, but by software companies like Lotus 1-2-3, WordPerfect, and later Microsoft.
Today, the path is almost identical. While everyone is still focused on the underlying capabilities, companies on top have already begun to compete for users, entry points, and workflows.
American companies can study the architecture of DeepSeek, utilize data generated by Kimi, and spend $20 million to train a well-performing open model. However, applications cannot be easily replicated.
Applications do not have downloadable weights or benchmarks. What determines their performance is the expertise accumulated by a company over years of trial and error.
In the past, large companies tended to treat models as standalone products, waiting for users to open a chat window specifically for them. Now, models are beginning to integrate into DingTalk, Feishu, and WeChat Work, with agents embedded directly into existing organizations and processes.
In the U.S., an office agent might grow along Gmail, Slack, and Salesforce. In China, it will first enter through DingTalk, Feishu, and WeChat Work, then connect to financial software, supply chains, factory systems, and government platforms. Even if both sides use the same model, the final products will not be the same.
After Prices Drop
From 2022 to 2025, the U.S. has progressively tightened chip restrictions on China. H100s are not sold, B300s are not sold, and advanced processes and manufacturing equipment are also being regulated. Under these conditions, Chinese companies have developed the world's largest open model.
By 2026, the U.S. will begin discussing restrictions on the weights of Chinese models. Reports indicate that OpenAI and Anthropic have expressed concerns to regulators, and Congress is studying whether model weights can continue to flow across borders.
The knife hasn't fallen yet, but insiders are already crying out in pain.
The chip ban restricts the supply for Chinese companies, but the model ban primarily cuts into the costs of American startups. The joint letter clearly states that bans cannot stop dissemination; they will only force developers to download all available models before the door closes.
Chips are physical; they can restrict orders, logistics, and foundries. Once weights flow out, they can be infinitely downloaded, mirrored, quantized, fine-tuned, and distilled.
What’s even harder to control is the price.
The U.S. still leads in cutting-edge closed-source models. OpenAI and Anthropic remain at the forefront, and training the top models still relies on NVIDIA's chips, a fact that won't change in the short term.
However, when it comes to applications, the outcome is rarely determined by a single strongest model. It depends on whether the model is affordable, whether companies have a ready digital entry point, and whether developers can truly integrate into business processes.
Looking back over the past thirty years, China has won many times, but the methods have been largely the same.
Mobile payments, e-commerce, food delivery, and live streaming have all taken technologies invented elsewhere and integrated them into the daily lives of over a billion people, creating astonishingly large businesses. The foundational layer has long been provided by others. Transistors came from Bell Labs, x86 belongs to Intel, TCP/IP was funded by the U.S. government, and Windows, Android, and iOS all originated on the West Coast. Even the most commonly used deep learning frameworks, TensorFlow and PyTorch, were developed by Google and Meta, respectively.
"Invention in the U.S., scale in China"—this phrase has been around for thirty years with few real exceptions.
Now, an exception has emerged.
The technical documentation from Thinking Machines Lab is very clear. One of the most prominent and wealthiest new companies in Silicon Valley has referenced the architecture of Chinese companies at the foundational level and used data generated by another Chinese model during training.
In the past, Silicon Valley sourced capacity, recruited talent, or sought markets from China. Now, it is directly following the validated technological paths of Chinese companies.
The admiration for Silicon Valley certainly has its roots. Over the past few decades, the most important foundational technologies in the computer industry have indeed largely emerged from there. From the chips in machines to the IDEs programmers open daily, Silicon Valley has long stood at the upstream of the technology chain.
But in the end, the tech world looks at results. Whoever can propose new architectures, lower costs, and inspire imitation among peers has the right to redefine the upstream.
Now, the names of Chinese companies are starting to appear on this list.
Silicon Valley will likely catch up in the future; there are enough engineers, capital, and chips there. However, applications that have already integrated into real business will not wait for anyone. Whoever first integrates into enterprise processes will gain customers, data, and the opportunity for the next round of improvements. The starting line has never been the same.
The era of AI applications will not open with a grand ceremony. It will first manifest in financial reports, skyrocketing usage, and a programmer once again exhausting their quota in the middle of the night.
Mass adoption is just a matter of time. This statement was made by an American.
Original link
This content is provided for general informational purposes only and doesn't constitute financial, investment, legal, or tax advice. Any events, rewards, online promotions, or related information mentioned herein should not be considered a recommendation, solicitation, or invitation to purchase, sell, trade, or otherwise deal in any crypto assets. Crypto assets are highly volatile and may result in loss. The availability of WEEX services, products, and related events may vary by region. You are responsible for ensuring that your participation is in accordance with applicable local laws and regulations.