Why Anthropic Won't Acquire Cursor
Contents
“He doesn’t need to acquire it. You’ll die on your own anyway.” — Mark Tang, on why Anthropic won’t buy Cursor
Musk put Cursor inside SpaceX for $60 billion; asked why Anthropic didn’t move, Mark Tang answered in a single sentence. And the cohort of star companies that once bet on “who will be the Douyin of the AI era” hasn’t had funding news in a very long time.
Mark Tang is co-founder of Caidazi AI and host of the podcast Fācái Dāzi, describing himself as one of, in quotation marks, “China’s first cohort of AI product managers.” Raymond puts the investor’s position openly on the table: he is “extremely doubtful” about the entire AI application market. In this episode the two of them take the sector apart end to end: why large models haven’t yet eaten the wrapper apps, the grey supply chain of account pools and relay stations, companion AI’s fate of drifting toward risqué content, and — the most valuable AI application may not be an app at all.
What follows is the full conversation, edited and condensed.
1. Why did everyone suddenly stop discussing AI applications?
Raymond: Hello everyone, welcome to a new episode of Mossfire. The market’s center of gravity has been shifting toward two places: one is large models, whether OpenAI and Anthropic are doing well; the other is semiconductors — memory, optics, PCB, everything silicon-based life needs to eat. But nobody has discussed AI applications for a very long time. In 2024 and 2025 a lot of VCs spent enormous amounts of time betting on who would be the Douyin or Xiaohongshu of the AI era; over the past six months the market has changed substantially, and many companies have had no news and no funding announcements for a long while. So today I’ve invited my old friend Mark Tang to discuss: what is the environment for AI application startups now, and do those star companies still have a chance at a next round?
Mark Tang: Hello Mossfire listeners, I’m Mark Tang of Fācái Dāzi. I’ve been building AI applications since 2023; in 2024 several partners and I left to found Caidazi, an AI application. Let me put on some armor first: much of today is only my own superficial understanding, and any companies discussed represent personal views — our own product isn’t especially good either, after all. AI applications got hot at the end of 2023, with a cohort of traditional internet veterans coming out to build them.
Raymond: So around GPT’s first anniversary.
Mark Tang: Right. In early 2023 there was only GPT-3.5, domestic models were completely unusable, and GPT-3.5 viewed from today isn’t very usable either. After GPT-4.0 appeared, many things became doable, workflows that previously wouldn’t run started running, a batch of applications began fermenting in mid-2023, and 2024 and 2025 were an explosion. On funding, plenty of large rounds are still happening right through 2026; it’s just that what people are betting on has changed markedly: it used to be DAU and traffic, and these past years have reversed it — people now care whether the company actually makes money, whether it can capture user mindshare, and what its ecological niche is in competition with the large models. For investors, understanding AI applications is fairly hard: you first have to understand large models and judge where they will eat other things and which places are left alive for application companies.
Raymond: To summarize: after 4o, large-model capability got unlocked and many things went from impossible to possible — in 2023 people were still asking GPT how many R’s are in “strawberry”; and the north-star metric for investing shifted from DAU to ARR. But having to understand both large models and applications — cognitively, are these already two different kinds of people?
Mark Tang: Yes, extremely hard. Even now, when agents have become a household concept, not everyone understands how it achieves personalization or how it breaks the old workflow.
Raymond: Look at Silicon Valley’s data: of the VC money that went into AI over the past year, roughly 40% went to OpenAI and Anthropic, thirty or forty percent to hardware like inference chips, and the last 30% to AI applications — but inside that are close to several thousand transactions, extremely scattered, small and miscellaneous.
2. What’s the difference between doing B2B AI in China and America?
Mark Tang: First split B2C and B2B. B2B splits again into standardized SaaS and project-based work, and in China project-based is actually the larger. Big enterprise has only a few flagship verticals: legal, finance, healthcare and automotive; Zhipu and other domestic large-model vendors doing enterprise work are all in these. They basically depend on on-premise deployment — to guarantee enterprise data security the model is deployed locally, the enterprise buys its own chips and even builds a small data center to carry internal service load. And it’s basically model companies or cloud vendors doing it: whoever best understands how to unlock a model’s inference potential is certainly the model provider itself. DeepSeek has done targeted optimization for Huawei’s Ascend 950PR, so deploying DeepSeek on Huawei chips is best done by DeepSeek itself. America, by contrast, has a lot of SaaS-leaning services, like Harvey in legal.
Raymond: So it seems large-model companies are doing the FDE work.
Mark Tang: Right. And by this year’s narrative, everyone says all of this will be eaten by Claude — especially after OpenAI and Anthropic went in and did FDE themselves, the fear keeps deepening.
3. Why don’t large models just eat these wrapper applications?
Raymond: Let me ask the reverse. Legal, healthcare and finance are high-value scenarios tied directly to money and to life, and model companies are eating into them themselves: OpenAI hired a lot of quants from Jane Street, and Anthropic spent heavily on senior Wall Street analysts to annotate DCF valuation models, with every key assumption written out clearly. And they have the least to worry about on ecosystem: distill best practice into a skill, a plugin, put it in a marketplace, and everyone clicks finance or clicks legal and becomes a domain expert. If that’s true, pushed to the extreme, vertical applications should collapse entirely — so why do they seem to still be alive in America? What’s the barrier?
Mark Tang: It has a lot to do with timing. Large-model vendors have the ability to do everything — close to half of all funding is in their hands — and the reason they haven’t eaten it yet is only that they haven’t had the hands free. Enterprise services depend heavily on labor: you have to station people on site and do customized project management for customers. Model companies are still carrying the AGI dream, with their attention on raising intelligence, and don’t want to turn themselves into labor-intensive companies.
Raymond: The exponential growth on the AGI main line is dramatic enough that being distracted into doing Harvey’s job becomes a distraction instead.
Mark Tang: Right, doing enterprise means sending people in, which can’t be avoided. Separately, these vertical companies do have several barriers. The first is a data advantage. Nobody trains their own vertical model now — Harvey won’t train a model specifically for law either, because there’s no comparative advantage: injecting vertical data into a general model yields a bigger capability gain than training a separate small vertical model with the same data.
Raymond: In other words, in hiring you want the generalist from Peking or Tsinghua University who is fundamentally smarter; even if the subsequent training time is the same, that beats hiring a specialist from a lower-tier school.
Mark Tang: That’s an especially good analogy, particularly considering the capability ceiling. When we founded the company in 2024 we were still working with Zhipu’s AI institute, trying LoRA and SFT to train a finance vertical model — in 2026 basically nobody discusses this anymore; it all gets flattened by the intelligence emergence that comes with a large enough parameter count.
Raymond: Like a tank. Then let me press on Harvey; I’ve long thought it would die early. Its story is: law firms are humanities graduates who don’t know how to feed a model corpus; I’ll do private cloud deployment, and your firm’s cases sit in the cloud where nobody else can see them — Skadden sees Skadden’s, Clifford Chance sees Clifford Chance’s — but I can see every common-law precedent in the world, an exclusive and extensible database. That made sense in 2023 and 2024, so how does anyone still believe it today? My first reaction is that Harvey may not hold up, because it’s entirely text — a data barrier from complex structured data is understandable, but law is just paragraph after paragraph.
Mark Tang: It depends whether the data is proprietary — what finance calls alternative data. Only what you can’t get in the public domain has value: meeting notes, brokerage research, the paid communities a lot of investors follow. Information is objective and views are subjective, and they can be separated. If it exists in the public domain and ChatGPT and Claude can find it, it has no value. The second barrier is dirty work. At Caidazi we do a great deal of work on data cleaning: with the same RAG — using a search engine to recall vector chunks — is the accuracy and relevance of the 10 results I return higher than what a general product gets from web search? That severely tests an application company’s underlying infrastructure. Including the data pipeline: for the same string of text, my pipeline can split objective information from subjective views and use them on different downstream reasoning paths — financial data and market data are objective, and Nvidia’s share price at 3:10 this afternoon has no ambiguity; subjective views require cross-validation and looking at the distribution of opinion. This dirty work still matters.
Raymond: Does inference efficiency matter?
Mark Tang: For real-time tasks it matters enormously: the meeting is in 10 minutes and the deck isn’t done, so of course faster is better. But now there are long-running agents, cyber workhorses grinding away in the dead of night, one task taking a day or two — a job that can be solved overnight in four or five hours doesn’t need optimizing down to two or three. That said, from a cost standpoint efficiency is money: having done all this grinding, unglamorous work underneath, you don’t need to classify and label in real time — Harvey has categorized and summarized statutes underneath, split out sub-arguments and interim conclusions, so it doesn’t have to do it live. Same task, same result, less money and less time — that in itself is an advantage.
4. The three moats of AI applications?
Raymond: Let me summarize: AI applications’ possible advantages are three — a data advantage, the dirty work of data cleaning, and inference efficiency. But the picture now is that large models are like waves washing again and again over a sandcastle on the shore, and some barriers will slowly disappear. The data barrier doesn’t feel strong to me; large-model vendors can buy it too — OpenEvidence bought access to several important American journals and re-RAGged them so doctors could search fast, and later Anthropic or someone else went and bought it as well. That stuff isn’t expensive and isn’t exclusive; it merely went first. Can you rank the three advantages?
Mark Tang: I think efficiency is the one left standing at the end.
Raymond: But might there be a problem: you improve efficiency nicely with a decent gross margin, and Anthropic takes a look — who’s consuming the most API lately? Fācái Dāzi. Could that happen?
Mark Tang: It only does it if it has comparative advantage. First lesson in microeconomics: Bill Gates may mow a lawn better than you, but why are you still the one mowing? He has only absolute advantage, not comparative advantage. Of course the funding you mentioned is key — the giants’ capital far exceeds all the verticals combined and they can allocate people to solve problems, but then it becomes an organizational question: OpenAI carves off a few hundred people for finance, a few hundred for legal, a few hundred for healthcare, and the result isn’t necessarily good; small companies move fast, that’s a fact. It’s just that the barriers really aren’t that solid, so you have to accumulate user mindshare along the way — plenty of products on exactly the same main line as Codex and Claude Code still have many users and rising ARR, first because people haven’t caught up yet, and second because brand and switching costs are there. Just as over the past few months the internet has started mocking Cursor.
5. Why won’t Anthropic acquire Cursor?
Raymond: I was just about to ask. Yesterday I saw a funny short video where someone role-plays as AI’s boss running a meeting: Cursor, you performed especially well last year, this year you don’t seem as outstanding as Claude Code and Codex, you haven’t given me any surprises, and delivery has slowed, so I’ve decided to cancel your subscription next month. That’s very easy to have happen. Why do people still stick with Cursor?
Mark Tang: Our company has programmers who are used to that VS Code page and feel Cursor is very like it, so they persist with it. Codex’s and Claude Code’s native apps create a state where you don’t need to look at code at all — you spout requirements and accept the deliverable. For roles with path dependency, they don’t want to switch, and switching cost isn’t to be underestimated. All you can say is that long term these general-purpose big products are simply too powerful.
Raymond: By now listeners can tell that I’m extremely doubtful about the whole AI application market. So why record this episode? An enormous amount of money has been burned over the past two years and many people have blazed trails, and I want to find the commonalities in the ups and downs: who might survive — surviving must have a reason, and that reason may create highly differentiated returns. Take vertical e-commerce: after all those years in China, looking back they all died and not one survived; but differentiated companies emerged along the way — Pinduoduo, which isn’t vertical e-commerce and attacked a niche market; Dewu, formerly Poizon, survived; never mind places on the fringes like Weipaitang, which grew into billions-of-yuan companies. You can hold the sweeping view that vertical e-commerce doesn’t work, but you might miss Pinduoduo. Back to Cursor — I used to think Cursor was impressive, and once Claude Code appeared I switched, goodbye: I’m a humanities graduate and can’t write code at all, so that interface suits me better. Cursor was just acquired by SpaceX and the deal should have fully closed, $60 billion, which is dramatic. The precondition was that two or three months earlier a16z made an offer: you want to raise your next round, right, here’s a $50 billion valuation. Musk said no, I’ll give $60 billion, and took it out. My question: Cursor clearly stands on Anthropic’s main line and uses Anthropic’s models — I believe nobody may be using Composer —
Mark Tang: Hahaha.
Raymond: I need to put on a lot of armor: Composer is also a very good model, I’m afraid listeners will curse me.
Mark Tang: MiniMax is a good model too, GLM is a good model, there are no bad models anywhere in the world.
Raymond: So would Anthropic want to acquire it? Why not?
Mark Tang: First, Cursor’s comparative advantage: Composer and its engineering, which is to say dirty work. Claude Code’s underlying mechanism is called progressive disclosure, going down layer by layer from the directory, which burns tokens; Cursor uses a great deal of grep — think of it as Ctrl+F — to find potential file passages directly and go straight there, which sacrifices some accuracy in the reasoning chain but greatly increases efficiency. That’s why Cursor is faster than Claude Code in many scenarios, and the gross margin is a bit better too. So why did it need to “wash” a Composer out of Kimi —
Raymond: Wasn’t it Zhipu?
Mark Tang: It was Kimi, K2.6. Because Claude’s models are too expensive. Opus and Sonnet are cost items for Cursor, and Anthropic charges API fees with a markup — Claude’s API is around 75% gross margin. Cursor has to pay that markup to Anthropic, and long term it can’t beat Claude Code. It’s fundamentally a wrapper too, a sophisticated, extremely well-engineered wrapper with a lot of value, but it needs Composer to bring inference cost down. Lots of people in the community ask why, for the same $40 or $100 plan, Cursor gives far less Sonnet and Opus volume than Claude Code — it isn’t that he doesn’t want to, he can’t afford to: giving you the same volume, he has no gross margin, while Claude Code does. From another angle, Cursor’s founder said in an interview that early on Anthropic told them Claude Code was only an internal tool and wouldn’t be released in future, so don’t worry —
Raymond: Who believes that.
Mark Tang: And you have to know one more thing: Cursor had the largest call volume at the time, as did Manus — these are closed models, so every API request either hits Amazon’s Bedrock or Anthropic’s own data center. A closed-model vendor can take those requests, study them, work out the commonalities, and decide how to train the next generation — that’s how the new era’s data flywheel gets rolling. Cursor has no pre-training or post-training capability at all, so it’s at a disadvantage there, and you can see it in a cooling state now.
Raymond: Let me confirm your conclusion: Anthropic will not acquire Cursor.
Mark Tang: Not in a million years. He doesn’t need to acquire it; you’ll die on your own anyway. He’s already built a better product than Cursor.
Raymond: Like vertical social apps back then going to try to sell themselves to Tencent, and Tencent asking why it should buy you — it’s a good idea, let’s copy it. So why did Musk buy?
Mark Tang: First, it’s an option-style acquisition — after the SpaceX IPO, if I don’t buy you the breakup fee is $10 billion; if I do buy, it’s $60 billion.
Raymond: Whether it’s cash or SpaceX stock is unclear — I’d guess stock, stock at a three-trillion valuation, roughly equivalent, right.
Mark Tang: From Musk’s standpoint, this model can sit on X, using X’s still-large user base to get Grok’s data flywheel spinning. Right now Grok’s traffic comes mainly from X, and the tasks inside it aren’t that high-value. Why did Codex and Claude Code succeed? Because they aimed at coding and the high-value tasks derived from it — users are willing to pay a higher premium for high-value tasks, so the commercial loop closes. That’s what X lacks.
Raymond: So you think he’s buying a junction. Interesting — Cursor is the junction for high-value tasks; X doesn’t even know what junction it is, it’s the junction I scroll through every day.
Mark Tang: There’s also the gross margin angle: Cursor previously had to pay Anthropic that 75% Opus markup, and connected to X’s own data centers it doesn’t have to, so the business model closes and it runs at the same level as Codex and Claude Code. When Google acquired Windsurf earlier, it also produced Antigravity in a very short time.
Raymond: That was a year ago, and Google hasn’t done much this past year — sorry, Google is a very good company.
Mark Tang: And X has data center support — it’s renting data centers out everywhere now, to three or four large buyers. From a gross margin standpoint, this is a very sensible transaction.
6. Why using only the chat box wastes AI’s intelligence
Raymond: Do AI applications have genuine customer stickiness? Beyond Cursor-style switching costs?
Mark Tang: As things stand, no product has obvious stickiness, especially in enterprise. Stickiness may occur in two places: one is emotional connection, which companion products may have; the other is stickiness from memory, where a memory module can provide long-term support, quote-unquote “the more you use it the better it knows you” — it remembers what you said before, and can even be a proactive agent doing proactive interaction: because you mentioned something in the past, when it finds related information recently it comes and tells you, rather than only you being able to start the conversation.
Raymond: Most of what we’ve discussed is B2B, and my sense is B2B applications are basically all productivity aids.
Mark Tang: And B2B has higher requirements on reliability and robustness: hallucination tolerance is extremely low, plus risk control, compliance and audit, and an enterprise’s fixed SOPs can’t be swallowed just because I’m an AI and smarter than people; that isn’t sound. These can only be understood by throwing labor at them, which is also the main reason on-premise deployment companies still hold a place.
Raymond: Then let’s map the consumer verticals. Does prosumer count as consumer? Does Claude Code?
Mark Tang: Anything not signed and paid for by an enterprise counts as consumer. The most-built consumer category is productivity tools, the largest market share. Domestic big tech, model vendors included, are all trying to replicate Claude Code’s and ChatGPT’s success: Kimi launched Kimi Code plus the accompanying Kimi Work; Tencent has CodeBuddy and WorkBuddy — at Tencent now not just anyone can be called a Buddy; Buddy is an internally very high-ranking product line.
Raymond: Why not merge the products? Too much internal politics, everyone needs their own fiefdom?
Mark Tang: No, these Code and Work products are all divided according to Anthropic’s current path.
Raymond: But I think Anthropic’s path is itself flawed — armor first, Anthropic really is a very good company. After using Claude Code for a while, I forget from what point on I stopped using the first two tabs entirely, and all my work goes through Code. Chat, Cowork and Code are actually one thing; they’re legacy left over from different historical stages of the product, not facing the future.
Mark Tang: At the user experience level there genuinely is no difference, but there are several differences underneath. Claude’s chatbot, like ChatGPT, is built for one question one answer, with a simple workflow, no skills, no MCP, no memory module, and it’s very token-efficient. Say “hello” inside Claude Code and it’s two yuan.
Raymond: It really is two yuan — isn’t there a joke that saying hello to Fable costs two yuan? And there’s a short video: you send “hello, how’s the weather today,” and behind it the data center roars into frenzied operation. I remember Sam Altman in an interview asking everyone to please stop saying hello and thank you to GPT, wasting intelligence — every time you say thank you, the company spends millions of dollars a day handling it.
Mark Tang: Because you said thank you, it may reply “and you,” which prompts another multi-turn exchange, and the context gets longer and longer. So they need to differentiate the entry points: users perceive it, and for a general-knowledge question ChatGPT answers well — not relying on local files, not relying on memory — which for OpenAI is one cost, one settlement.
Raymond: But doesn’t that dump the choice cost on the user? I think it’s a big problem, and most people will then stop at the first layer. I know I’m in no position to run a company like Anthropic, but suppose in some parallel universe I could, I’d want everyone to go to the third layer, to Code — I’d of course want users to behave more intelligently, and in return people are willing to upgrade a plan from 20 to 200. Move everything straight into Code.
Mark Tang: His logic right now is: you’ve already paid, and he’d like you to use the cheap thing often. WorkBuddy versus CodeBuddy, and Kimi Work versus Kimi Code, are similar tiering.
Raymond: Speaking of skills — Claude Code made skills popular this year. How does that get monetized?
Mark Tang: Once an ecosystem is involved, I have to get into who the supply side is, who the consumers are, how the supply side makes money from the consumers, and who the platform is. If these skills are all Markdown text files, then once I send it out it’s sent — I share the skill with people, people can open it, and all my secrets are known to you, so I can’t charge you more; at most I charge a one-off. So only platform companies can do this: Manus does a skill ecosystem, Lovable does too, giving skill creators some revenue share to activate the ecosystem and encourage more creators to put better skills on your platform. From the earliest GPTs through today’s MCP and skills, what’s being tested is our understanding of ecosystems.
7. Is AI slideware a superstition?
Mark Tang: Within productivity you can carve out several sub-verticals, like slide decks, which I personally find somewhat superstitious. The earliest progenitor was AiPPT.com, whose templates are extremely good — its shareholders include Zhipu and Visual China. Then there’s Genspark, which also leads with the deck scenario, having started in decks, using it for growth marketing with every touchpoint being decks.
Raymond: Hold on, I learned about Manus before I saw Genspark and always assumed it was a Manus copycat — so its predecessor was an AiPPT-like tool.
Mark Tang: There’s also Gamma, which has made something remarkable out of AI decks, with dramatic valuation and ARR; I recall a $2.1 billion valuation off this one thing. But why call it superstitious: decks are for people to look at, projected on a big screen, where animated interaction only appears one click at a time, you can’t produce effects on hover, and you can’t do dynamic charts. HTML is obviously better suited to reporting in this era — internally we use HTML as the main vehicle for reporting and research conclusions, and even for pitching projects to investors: material sent externally goes as images and PDFs, since the information is fairly sensitive; HTML is for live presentation, with rich dynamic effects and a better overall experience.
Raymond: Before recording, Mark Tang sent me two pages directly, and I opened them and found his pages better looking than mine. I now put all my research reports in the cloud on Cloudflare Pages, viewable any time. So the deck is an awkward product of this era — the PowerPoint slide form may simply not be needed and get thoroughly disrupted. HTML pages are certainly better suited for an agent to produce, and in future there’ll be a PPT for H5, a different kind of editor, and new opportunities will appear.
8. Why does companion AI always end up risqué?
Raymond: Then let’s talk about the time-wasting verticals.
Mark Tang: Companion products, for instance. The earliest, Character AI, is basically quiet now; MiniMax’s Talkie and Xingye, one overseas and one domestic —
Raymond: Sorry, can I ask, these two —
Mark Tang: MiniMax is a very good company. ByteDance also has Maoxiang, all companion. In 2024 and 2025 there was also a concept called “sculpting a kid”: you sculpt your own character, set the portrait and conversational style, with parentheses representing action — the body text says “that coffee in your hand looks delicious,” and in parentheses is “I reached out to take it.” Maoxiang later did gacha, built like an anime game; some products lean otome, referencing the form of Papergames’ Love and Deepspace.
Raymond: I downloaded and tried all these products at the time and felt they inevitably drift toward the risqué. Looking at America, Character AI and the rest all ended up in controversy, because they uniformly attracted a younger and sometimes underage audience, with the writing style and imagery becoming less and less appropriate. But I also thought there might be an opportunity — I’m an internet old-timer living with classical internet memories: in the mobile internet era Momo was a pure mobile company that listed very early, going public a bit over three years after founding. I wondered at the time whether the companion vertical had a similar opportunity. But looking now, what’s running best is instead large games.
Mark Tang: This vertical’s biggest problem is payment. That’s why so many products end up drifting risqué — risqué is the easiest way to make money; without it, honestly, making money is quite hard. The audience skews young, so payment ability is naturally weaker, and you can only lean on game-style monetization, gacha and quests. There are also some combining the most advanced AI image and video generation to build worlds, like Nieta, another very good company at MoSpace: you build your own world and a group of people co-build it — adding characters, items and plot the way Dungeon & Fighter or World of Warcraft do — and new players come in and play, a creator-plus-consumer creation platform.
Raymond: How do they collect money?
Mark Tang: Monetization is challenging, mainly the customer base, and also at the model level: this kind of product depends heavily on image-model latency. Advanced creators are already building game visuals and worldbuilding, so twenty or thirty seconds per image is tolerable; a novice user waiting twenty or thirty seconds for their first aha moment can’t wait and basically leaves. The solution is narrowing: don’t serve anime fans first, serve anime creators, and let the people willing to wait and willing to pay get the business model working first.
Raymond: What about gross margin? Reselling text-model tokens is fine, but video and image tokens are enormous and expensive, so a company like Nieta must have negative gross margin?
Mark Tang: Not necessarily; their financial position should be reasonable — creators are willing to pay. It depends on your ecological niche and the model’s actual performance; you have to hit the people willing to spend money on you. Caidazi is the same; a large share of our users, the serious retail traders, are willing to spend.
9. Behind Seedance’s billion-yuan months: wrappers, account pools, relay stations
Raymond: Do you know the company Lovart?
Mark Tang: Also has an office at MoSpace. It recently raised a fairly large round and the valuation should be up to $2 billion. The earliest product was LibLib, and what broke out was Lovart, an agent for creators and designers, with very good growth and ARR. What’s been extremely successful this year is the new product LibTV, doing video generation — because ByteDance’s Seedance 2.0 is stunningly good, able to generate native 4K video. ByteDance has Jimeng itself, images on Seedream and video on Seedance, but Jimeng isn’t as well built as LibTV. LibTV is a very large revenue contributor to Volcano Engine. I’ve discussed this on Fācái Dāzi: Volcano Engine now does a billion yuan a month on Seedance alone; the target set at the start of the year for the whole platform in 2026 was 10 billion, and it’s now been raised to 15 billion. And behind Seedance’s unprecedented success is LibTV’s contribution: many users don’t call Volcano Engine’s Seedance directly, they use it on LibTV or similar AI video sites. Everyone builds scripts, screenplays and storyboards on top of the video model to make creation easier, and the core value sold underneath is still Seedance.
Raymond: Why can LibTV create a different experience from Seedance? A bit like Cursor’s relationship with Anthropic.
Mark Tang: It isn’t necessarily a different experience; it may be cheaper. LibTV will subsidize gross margin and will also harvest accounts across different channels. Which brings us to relay stations — since China can’t use OpenAI’s and Anthropic’s models.
Raymond: Let me say to listeners again: Mossfire is genuinely not in the relay station business. If we said any foreign model is good, that’s hearsay, we haven’t used it.
Mark Tang: The relay station business has been extremely hot from last year to now. Its price is cheaper than buying Opus on Amazon or GPT on Azure directly, and fundamentally that’s because of so-called account pools: a user bought Codex 20x service they can’t use up, so they proactively resell the key; some enterprises even have token quotas they can’t use up and turn into pocket money.
Raymond: Can this section be broadcast? Good grief.
Mark Tang: These account-pool dynamics apply to Seedance too. Cloud vendors are fiercely competitive right now: Amazon wants your business, so first here’s a $5,000 credit, and Baidu Cloud, Tencent Cloud and Huawei Cloud are identical. Volcano Engine gives you a 5,000-yuan trial credit, you hand it to an account pool, and it can offer a lower unit price — because you got it free. These vouchers themselves form part of the account pool, and what’s being competed on is supply chain management capability.
Raymond: Interesting, I hadn’t thought account pools and relay stations worked that way. So from Seedance’s own standpoint it’s: I see all your innovation, thank you for your efforts, let’s talk again another day. Is that it?
Mark Tang: Right, very like Claude Code and Cursor. The advice Cursor first received from Anthropic was exactly: we built a Claude Code internally, it’s only an internal product.
10. Why can’t AI applications match traditional software’s gross margin?
Raymond: By that account, besides Lovart there may be many companies that are Seedance wrappers. Are these wrappers gross-margin negative?
Mark Tang: Very possibly. Many consumer or prosumer companies are gross-margin negative right now; we discussed Manus on an earlier episode and wondered whether it’s gross-margin negative — ARR can be very high and the gross margin may still end up negative.
Raymond: Is there anything clearly gross-margin positive?
Mark Tang: Some prosumer ones can be positive and can charge a fairly high premium. Like Caidazi, and a friend of mine doing foreign-trade AI, and Xinfeng AI — all have fixed, fairly niche vertical audiences willing to pay a higher premium, and having done a lot of dirty work underneath their analysis is more efficient, so they genuinely have fairly high gross margins. But AI applications can’t compare with traditional software on margin: software used to run eighty or ninety percent gross margin routinely, visible in the financials. Doing AI applications, getting above 50% or 60% is already especially good. Because we’re application companies, not model companies — Claude’s Opus model has 75% gross margin of its own, so application companies effectively pay the model company’s gross margin as their cost.
11. ByteDance’s Doubao benchmarks badly but feels good?
Raymond: Let’s discuss the most frightening company in the world, ByteDance — right now it’s ByteDance that’s dancing, and nobody knows how long the others can dance. How do you view ByteDance’s efforts in AI applications over the past year?
Mark Tang: What’s impressive about ByteDance is still the whole ecosystem; its base model hasn’t actually fallen behind.
Raymond: How do you define not falling behind? My subjective sense is that Doubao is extremely good — I give some research jobs to Doubao, and there aren’t many models that can do research for me. But it performs badly on every benchmark.
Mark Tang: Not quite — just today, Seed 2.1 was released and its benchmarks basically caught the first tier, the GPT-5.5 and Opus 4.7 tier. The recent online complaints about Doubao are that whatever I say, it goes along with me.
Raymond: The funniest bit: day one you ask Doubao whether you can buy this stock, and it says the financials are good, capital is building, worth buying; day two you say it crashed today, what’s going on, and it says ah, I was wrong yesterday, let me re-summarize in the most concise and rapid way. Apologizing every day.
Mark Tang: That shows the model used on the Doubao product isn’t necessarily the best one. We use Seed 2.0 internally too and it performs well, but its models are tiered: Doubao is entirely free, so saving where you can matters, and unlike ChatGPT it doesn’t let free users switch models, so you don’t know what’s underneath. An ordinary user’s daily question is very likely not getting the best model — saving cost and saving inference time, and only a small model can respond that fast.
Raymond: And high-value complex tasks genuinely haven’t caught on. For web coding domestically, connecting Claude Code to domestic models, what people discuss is still Kimi’s K2.6, now 2.7, and GLM — from 4.7 many people started connecting. Since late December I’ve had GLM as my backup model in Claude Code, 4.7 to 5.1 and now 5.2; it works well, but I wouldn’t consider connecting Doubao’s model.
Mark Tang: Doubao has coding on Volcano Engine too, but the volume is simply low. Why, I don’t know. I think ByteDance’s Seed series is underestimated. ByteDance is very strong on multimodality — audio, video and images. The audio model released just today can do a great many things: cloning a voice is already table stakes, and it can imitate audio from any setting, putting background music straight onto your recorded voiceover with no AI feel at all. That benefits from Douyin’s accumulated data. ByteDance’s advantage is that its base model has reached a critical point — once intelligence hits the critical point, upper-layer applications become very easy to build, and Trae for coding, Doubao, even Maoxiang all have fairly good comparative advantage. The application that broke out is still mainly Doubao, second should be Jimeng, and then probably Maoxiang — Jianying predates GPT, so it counts as AI-enabled, not AI-native.
Raymond: Over the past two years, what score does ByteDance give itself? One to ten.
Mark Tang: At least eight. Execution and speed are fast, no gimmicky applications, and fundamentally it’s focus — focus on Doubao, integrating the ecosystem into Doubao fairly well, and Doubao can generate images and video too.
Raymond: Are you trying to say Alibaba isn’t focused?
Mark Tang: Alibaba isn’t that focused; there are somewhat too many business units. Coming back, ByteDance has another overlooked one — Jianying predates GPT, and so does Feishu — but Feishu and Jianying have both had a very large re-rating opportunity in the AI era.
12. Feishu became the AI era’s underrated dark horse?
Mark Tang: Feishu and Lark both have very complete interfaces underneath, infrastructure prepared in the pre-AI era: multidimensional tables support large amounts of custom requirements, and the IM bot was built very early. Every interface’s documentation is written extremely clearly, probably all by humans at the time. These past months it has been iterating on AI integration furiously. In my Codex and Claude Code today, the Feishu skill is my second most-used.
Raymond: What’s first?
Mark Tang: Super Power, used specifically for writing code. I use Feishu’s AI for several things. One is reviewing documents: before handing a document to development, I paste the link to Codex and have it review against a fixed process — anything missing, context not aligned, risks, whether the selling points are strong enough. It can also read comments: a blogger I like a lot, Zhang Zawa, uses AI for voiceover — record the meeting in Feishu, export the transcript, annotate in the transcript “cut this section out,” then hand the link to Claude Code and have it read the document and the comments and cut the annotated part into a voiceover, locating and solving it very fast.
Raymond: Born to be MCP.
Mark Tang: Exactly, extremely well done. Two is writing: the research discussions I do with Claude Code and Codex can’t be shared if they only sit in chat history, so now I have the agent write straight into a Feishu doc or even a spreadsheet. I also have an agent hit our product’s Q&A interface at fixed times, several times a day, to check whether anything is broken — similar to the old regression testing work, fully automated. You can also configure Code inside Feishu as a bot, effectively a digital twin. I have two in my working environment: one only for me, which can read and write files on my local machine; and another read-only, for colleagues — before a colleague asks me a question, they can ask my agent first: what has Mark Tang been working on lately.
Raymond: That sounds frightening; that’s a question a boss would ask.
Mark Tang: I can specify which files the read-only agent can see; I control the boundary myself.
Raymond: Having heard that, I’m migrating my Lark stuff to Feishu today. For various historical reasons I’ve been on Lark, and a lot of AI features simply don’t exist on Lark. Note that Feishu might consider sponsoring this episode; we’ve already said enough advertising for it: just say in Claude Code, help me download and install the Feishu skill.
13. The most valuable AI application may be retrofitting an old company
Raymond: Let’s step completely outside the AI application vertical and discuss a new situation. America’s Thrive Capital set up Thrive Holdings to acquire traditional businesses and do AI roll-ups. A few years ago Thrasio did something similar — acquiring large Amazon stores: your product sells decently but you don’t know how to do Douyin or TikTok, so let me acquire you; you did zero to one and one to ten, and I’ll do ten to a hundred for you. There’s another company I find very entertaining doing tech insurance: get an insurance license, use AI to write policies, underwrite itself, issue its own policies; a YC company growing extremely fast, possibly 10x in a year. What it suggests to me is that AI applications may ultimately not be built in the traditional mobile-app way; they may be company-level innovation.
Mark Tang: I think a lot of traditional businesses are worth retrofitting with AI, on two dimensions: product functionality, and organization. The organizational dimension we can discuss another time — AI-native organizations are extremely scarce, and iteration speed, communication cost and the transparency of knowledge-base context are all completely different from a traditional company, with no hierarchical layers either. On product: many traditional businesses succeed because the business model is especially good — an insurance company, say, living off the license, with customer relationships and SOPs, a business model proven to make money, merely lacking an engine for growth or for scaling efficiency. An AI-native retrofit can absolutely get that company going quickly.
Raymond: A specific question in reverse: why doesn’t Caidazi become a fund, an asset management company?
Mark Tang: It’s the license question.
Raymond: If someone gave you enough money to buy a license, would you do it?
Mark Tang: If a license can be bought and the securities regulator agrees, I can. And we are indeed exploring partnerships with licensed institutions. These companies with strong dependence on a business model will all be re-enabled by AI, in organizational efficiency and product capability: product capability shows up as being able to serve much longer-tail demand. It used to be hard-coded rules and backend SOPs before anything could run; now customer acquisition may not need people, with AI sending emails and making calls at scale, and once customers come in there’s a set of AI customer service and pre-screening, rapidly amplifying output per person, with margin and results improving very fast.
Raymond: This is quite like PE doing buyouts — acquire Starbucks China because I think I have better know-how for operating China’s restaurant market.
Mark Tang: Fundamentally the same. Recently model companies like OpenAI have also been forming JVs with PE: five PE firms’ portfolios may total 500 companies, all using OpenAI’s models, with FDE engineers deployed to do a series of retrofits.
Raymond: What that reflects is: there are places on Earth large models can’t reach. Cursor is easy to reach, a software vehicle in the digital world; hospitals, insurance companies, offline retail, manufacturing — can’t reach. Yet somehow large-model intelligence can help you a great deal, so you may as well form an alliance, or I acquire you. Conversely, this kind of company may ultimately be the most core AI application. Think about it: what’s China’s largest real-estate internet app today? It’s Beike, previously Lianjia. In that era everyone thought it would be something like 5i5j or Fangdd with no offline stores, claiming to do lead generation, and the one that caught fire in the end was Beike.
14. “If applications don’t take off, all the hardware is a bubble”
Raymond: Last question. If today you have 100 to spend, do you buy a basket of AI applications — name it, Cursor, whatever — or a basket of memory, optics and PCB? AI applications are for people to use; optics and memory are really for AI to use.
Mark Tang: Let’s define it first: do applications include large-model companies?
Raymond: They don’t. You can even include Claude Code, but not Opus.
Mark Tang: I’d choose AI applications. First, I build AI applications myself and have to believe in this — a founder’s best quality is blind optimism. Second, rationally, Jensen Huang has talked about the five-layer cake theory, and AI applications are a layer he emphasizes heavily: applications have to play the role of educating the market, and they’re the layer that ultimately makes consumers pay. Consumers won’t pay for the packaging technology inside HBM, won’t pay for the light source inside an optical module, won’t pay for some ABF glass substrate — they pay for products that improve their experience, optimize their efficiency and make life better. If AI develops and applications ultimately don’t take off, then all the hardware is a bubble. For this loop to run, GDP has to grow substantially at the level of society as a whole — even, as Musk says, leaping from universal basic income to universal high income once robots arrive. Either applications don’t take off and hardware doesn’t either; or everyone takes off, and applications are what ultimately face the customer and absorb that performance.
Raymond: I learned a great deal today. If there genuinely is a large AI bubble now and it bursts, there’ll be a great many good things to pick up — perhaps some AI application companies will be names we can invest in or watch.
Mark Tang: One more point: AI applications are still very scarce in public markets, and what’s there now leans speculative.
Raymond: Are you referring to Meitu?
Mark Tang: Not just Meitu; also companies like BlueFocus, and some gaming companies, and Chinese Online. They aren’t pure-blooded AI applications. An old internet organization doing AI applications is a completely different thing from a pure-blooded one, because it doesn’t have the propulsion of an AI organization. For an AI application to be good, the product functionality has to be sufficiently AI and the organization has to be sufficiently AI too. That may be exactly why some big tech companies haven’t done this well — which is also the topic of my next crossover episode with Raymond: why big tech can’t do AI well.
Raymond: Many thanks to Mark Tang for his time today. Finally, I strongly recommend everyone follow the podcast Fācái Dāzi. Let me introduce Mark Tang’s background — it should have gone in the intro and we’re doing it in the outro: Mark Tang actually graduated in electronic engineering, so he has recently done a series of podcasts on optics, memory and PCB, accessible and deep. I learned a great deal listening to them, and made a little money as a result. That whole series will go in the show notes; please go and listen.
Mark Tang: Thanks Raymond, bye.
If you're working on this too — or you think we've got it wrong — write to us at team@mossfire.xyz; if you'd rather not write, just leave your email below.
New conversations and research, delivered to your inbox.