<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
    <channel>
        <title>AI on TorchTree</title>
        <link>https://torchtree.com/en/tags/ai/</link>
        <description>Recent content in AI on TorchTree</description>
        <generator>Hugo -- gohugo.io</generator>
        <language>en</language>
        <copyright>TorchTree Co., Ltd.</copyright>
        <lastBuildDate>Tue, 25 Aug 2026 12:58:43 +0800</lastBuildDate><atom:link href="https://torchtree.com/en/tags/ai/index.xml" rel="self" type="application/rss+xml" /><item>
        <title>Notion CEO&#39;s Jazz Mode: Five Deep Shifts in AI-Era Organizational Management</title>
        <link>https://torchtree.com/en/post/notion-jazz-mode-ai-org/</link>
        <pubDate>Tue, 25 Aug 2026 12:58:43 +0800</pubDate>
        
        <guid>https://torchtree.com/en/post/notion-jazz-mode-ai-org/</guid>
        <description>&lt;p&gt;In a recent conversation with Sequoia partner Brian Halligan (former HubSpot CEO), Notion CEO Ivan Zhao introduced a new concept: &lt;strong&gt;Jazz Mode&lt;/strong&gt;. He argues that after Manager Mode and Founder Mode, organizations in the AI era should operate like a jazz band: everyone has room to improvise, yet the whole still comes together in collaboration.&lt;/p&gt;
&lt;p&gt;This article doesn&amp;rsquo;t intend to recap everything Ivan said. Instead, it tries to extract a few structural changes that are easy to overlook from his remarks, and what those changes mean for teams building products with AI.&lt;/p&gt;
&lt;h2 id=&#34;from-building-bridges-to-brewing-beer-why-ai-rewrites-the-underlying-logic-of-product-development&#34;&gt;From Building Bridges to Brewing Beer: Why AI Rewrites the Underlying Logic of Product Development
&lt;/h2&gt;&lt;p&gt;Ivan used a metaphor to distinguish traditional software development from AI product development: traditional software is like building a bridge — designers draw the blueprint and engineers build to it; AI products are more like brewing beer — you can only experiment, observe, and adjust, without precisely controlling the final result.&lt;/p&gt;
&lt;p&gt;The value of this metaphor isn&amp;rsquo;t rhetorical; it exposes a structural problem that has been underestimated in AI product development: &lt;strong&gt;traditional software development is requirement-driven; AI product development is technology-driven.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;In the traditional model, product managers define requirements, designers produce solutions, and engineers implement them. The whole process starts from customer needs, with technology as the means of execution. But AI products are different. The boundaries of model capability determine what a product can do, and those boundaries change every week. If you plan your product roadmap strictly according to customer requirements, three months later what you&amp;rsquo;ve built may already be obsolete.&lt;/p&gt;
&lt;p&gt;Ivan said their internal development model has shifted from &amp;ldquo;customer-driven&amp;rdquo; to &amp;ldquo;technology-driven experimentation.&amp;rdquo; The traditional boundaries between PM, designer, and engineer have been thoroughly blurred — even product designers are among the team members who consume the most LLM tokens.&lt;/p&gt;
&lt;p&gt;This isn&amp;rsquo;t just Notion&amp;rsquo;s practice. As LLM capabilities iterate quickly, more and more AI product teams face the same dilemma: you can&amp;rsquo;t fully plan your product at the start of the year, because three months later the model&amp;rsquo;s capabilities have already changed. &lt;strong&gt;An AI product roadmap is, in essence, a collection of assumptions that keep being overturned.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;For teams building AI products, this means two things. First, you need to accept that &amp;ldquo;plans can&amp;rsquo;t keep up with change&amp;rdquo; isn&amp;rsquo;t a management problem but a consequence of the product form. Second, you need to let the people with the best judgment on your team (not just engineers) work directly with the model, because instinctive judgment about model capability is becoming a more valuable product asset than a requirements document.&lt;/p&gt;
&lt;h2 id=&#34;hierarchy-wont-disappear-but-its-function-is-being-redefined&#34;&gt;Hierarchy Won&amp;rsquo;t Disappear, But Its Function Is Being Redefined
&lt;/h2&gt;&lt;p&gt;Ivan was explicit in the conversation that he doesn&amp;rsquo;t believe in hierarchical-free organizations. His reasoning is simple: hierarchy is human nature — even chimpanzee societies have natural hierarchies, and you can&amp;rsquo;t eliminate it by forcibly flattening the org.&lt;/p&gt;
&lt;p&gt;But he also pointed out that language models are becoming the new infrastructure of organizations. Information transfer, state synchronization, and decision coordination — work that once required many middle managers — can now be partly handled by AI. Future organizations will be flatter, but hierarchy won&amp;rsquo;t disappear.&lt;/p&gt;
&lt;p&gt;Hidden beneath this is a deeper change: &lt;strong&gt;the rationale for hierarchy is shifting from &amp;ldquo;information transfer&amp;rdquo; to &amp;ldquo;allocation of judgment.&amp;rdquo;&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;In traditional organizations, one core function of hierarchy is information filtering and passing. A CEO can&amp;rsquo;t know every detail, so you need VPs; VPs need Directors; Directors need Managers. Each layer compresses information and reports upward. But AI tools (including Notion&amp;rsquo;s own products) are making information transfer more efficient, even automated. When information transfer is no longer the bottleneck, hierarchy only retains two values: decision judgment and resource allocation.&lt;/p&gt;
&lt;p&gt;This means the job of future middle managers will fundamentally change. They will no longer be transfer stations for information, but carriers of judgment in specific domains. The value of a middle manager will no longer depend on how many people they manage, but on how broadly they can make high-quality decisions.&lt;/p&gt;
&lt;p&gt;For organizational designers, this is a very practical question: if your middle managers mostly do meetings, reporting, and passing information, their roles are being eroded by AI; if your middle managers are the judgment centers of a domain, their value is actually rising.&lt;/p&gt;
&lt;h2 id=&#34;the-talent-formula-has-changed-capability-or-taste--which-is-scarcer&#34;&gt;The Talent Formula Has Changed: Capability or Taste — Which Is Scarcer?
&lt;/h2&gt;&lt;p&gt;Ivan proposed a talent formula he currently endorses most: &lt;strong&gt;Talent = Capability × Taste × Agency.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;In the past, Capability was the most important variable. An engineer&amp;rsquo;s technical skill determined their output. But in the AI era, capability is being rapidly commoditized. Tools like Claude Code, Cursor, and Copilot are making &amp;ldquo;writing code that runs&amp;rdquo; increasingly easy.&lt;/p&gt;
&lt;p&gt;When capability becomes cheap, Taste and Agency become scarce resources. Taste is what you think is good; Agency is whether you proactively push things forward.&lt;/p&gt;
&lt;p&gt;This isn&amp;rsquo;t empty philosophy. Notion has adjusted its hiring strategy along these lines: it no longer focuses mainly on work history, big-company background, or resume length, but on curiosity, optimism, energy, and proactivity. Their first-round interviews no longer review resumes; they simply ask candidates to &amp;ldquo;build something,&amp;rdquo; looking at the work before the background.&lt;/p&gt;
&lt;p&gt;Even more noteworthy is their &amp;ldquo;barbell model&amp;rdquo; engineering organization: at one end, very young engineers responsible for fast trial-and-error and efficient execution; at the other, a tiny number of super-senior architects responsible for judgment, Taste, and architecture. One senior architect guides 2–3 junior engineers, and with the help of AI Agents, each junior engineer can produce as efficiently as managing a mini-team.&lt;/p&gt;
&lt;p&gt;The underlying logic of this model is: &lt;strong&gt;AI amplifies individual execution, but not judgment.&lt;/strong&gt; A tasteful senior engineer, plus AI tools and a few junior executors, can produce what a small team used to. But if the Taste is wrong, no amount of execution is doing anything but accelerating in the wrong direction.&lt;/p&gt;
&lt;p&gt;For team managers, this means you need to reassess who in your team is the &amp;ldquo;Taste center,&amp;rdquo; and ensure those people&amp;rsquo;s judgment directly shapes product direction rather than being diluted by layers of reporting.&lt;/p&gt;
&lt;h2 id=&#34;founders-are-an-organizations-decalcifying-agent-why-notion-took-in-60-founders&#34;&gt;Founders Are an Organization&amp;rsquo;s Decalcifying Agent: Why Notion Took In 60 Founders
&lt;/h2&gt;&lt;p&gt;Notion has 50–60 former startup founders internally — an unusual number. Ivan said founders are an organization&amp;rsquo;s &amp;ldquo;decalcifying agent.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;His logic is clear: a 1,000-person company will naturally tend toward bureaucracy; that&amp;rsquo;s a physical law. But if a steady stream of founders keeps mixing in, they continuously launch new projects, challenge old processes, propose new ideas, and push organizational change — effectively injecting new vitality into the organization.&lt;/p&gt;
&lt;p&gt;This observation reveals a problem often neglected in organizational theory: &lt;strong&gt;bureaucratization isn&amp;rsquo;t a management failure, but a natural result of organizational scale.&lt;/strong&gt; Any sufficiently large organization will produce processes, norms, and hierarchy. These things ensure efficiency early on, but once they accumulate past a certain point, they become obstacles to innovation.&lt;/p&gt;
&lt;p&gt;The value of founders is that they naturally distrust process. They&amp;rsquo;re used to &amp;ldquo;doing first, then talking,&amp;rdquo; used to breaking existing frames and rethinking problems. Scattered across a large organization, they act like so many &amp;ldquo;anti-entropy nodes,&amp;rdquo; constantly resisting the force that drives organizations toward rigidity.&lt;/p&gt;
&lt;p&gt;But there&amp;rsquo;s a prerequisite: the organization must give these people enough autonomy. If founders are assimilated by existing processes after joining a large company, the &amp;ldquo;decalcifying agent&amp;rdquo; fails. Ivan said Notion&amp;rsquo;s approach is to keep these people highly autonomous, consistent with the spirit of Jazz Mode.&lt;/p&gt;
&lt;h2 id=&#34;enterprise-sales-is-the-last-fortress-ai-cant-take&#34;&gt;Enterprise Sales Is the Last Fortress AI Can&amp;rsquo;t Take
&lt;/h2&gt;&lt;p&gt;One detail in the conversation is easy to overlook: Ivan said they had previously made mistakes in the sales area — they tried to reinvent sales and failed. Later they found that many enterprise customers simply want to talk to a real person.&lt;/p&gt;
&lt;p&gt;This observation forms an interesting contrast with the current narrative of many AI companies. Over the past two years, a lot of AI startups have tried to replace the sales process with AI, from automated outbound calling to intelligent customer service to AI SDRs. But Ivan&amp;rsquo;s experience suggests that, at least in the enterprise market, trust relationships between people remain the key to closing deals.&lt;/p&gt;
&lt;p&gt;The logic behind this: enterprise procurement decisions are costly, and the consequences of error are severe. In such scenarios, customers need more than just information and efficiency; they need &amp;ldquo;someone who can be found when something goes wrong.&amp;rdquo; AI can handle information, but it&amp;rsquo;s hard for it to provide that kind of psychological security.&lt;/p&gt;
&lt;p&gt;For AI entrepreneurs, this points to a pragmatic direction: &lt;strong&gt;AI&amp;rsquo;s best position in the enterprise market may not be replacing sales, but empowering it.&lt;/strong&gt; Let salespeople use AI tools to prepare materials, analyze customers, and follow up leads more efficiently, but the final customer relationship is still maintained by humans.&lt;/p&gt;
&lt;h2 id=&#34;planning-cycles-are-shrinking-dramatically&#34;&gt;Planning Cycles Are Shrinking Dramatically
&lt;/h2&gt;&lt;p&gt;Ivan said something very direct: financial planning can be done quarterly, but product planning may need to change weekly, because model capabilities change too fast.&lt;/p&gt;
&lt;p&gt;That&amp;rsquo;s not an exaggeration. Over the past two years, the capability curve of LLMs has been nearly exponential. Tasks considered impossible when GPT-4 was released in early 2023 had been matched or surpassed by multiple models by the end of 2024. If your product planning is built on the assumption of &amp;ldquo;GPT-4-level capability,&amp;rdquo; that assumption needs updating every few months.&lt;/p&gt;
&lt;p&gt;Ivan also stressed that a CEO must experience AI firsthand — not just watch videos, talks, or summaries. He used two phrases: Feel the AI, Feel the AGI. Only by using it yourself and building with it can you know which opportunities really exist.&lt;/p&gt;
&lt;p&gt;The practical meaning of this advice: &lt;strong&gt;product decisions in the AI era increasingly rely on intuitive judgment of model capability, and that intuition can only come from hands-on use.&lt;/strong&gt; You can&amp;rsquo;t gain an accurate sense of model capability by reading secondhand information, just as you can&amp;rsquo;t learn to swim by reading a swimming tutorial.&lt;/p&gt;
&lt;h2 id=&#34;a-company-eventually-grows-into-the-image-of-its-founder&#34;&gt;A Company Eventually Grows Into the Image of Its Founder
&lt;/h2&gt;&lt;p&gt;At the end of the conversation, Ivan touched on a view: many CEOs like to imitate others — Jobs, Musk, or Brian Armstrong — but a company is ultimately a projection of its founder. If you&amp;rsquo;re a craftsman, the company becomes a craftsman culture; if you&amp;rsquo;re a salesperson, it becomes a sales culture; if you&amp;rsquo;re a jazz musician, the company also eventually becomes a jazz band.&lt;/p&gt;
&lt;p&gt;He was very candid about the quasi-religious/&amp;ldquo;cultish adoration&amp;rdquo; perception of the Notion community, even saying he &amp;ldquo;liked&amp;rdquo; it. He believes a company is, in a sense, a religion, projecting a worldview and value system onto the real world through commerce and products.&lt;/p&gt;
&lt;p&gt;The deeper meaning of this view: &lt;strong&gt;organizational culture isn&amp;rsquo;t designed; it&amp;rsquo;s an amplification of the founder&amp;rsquo;s personality.&lt;/strong&gt; You can draw a perfect org chart, but what ultimately determines organizational behavior is what the founder believes, values, and how they make decisions.&lt;/p&gt;
&lt;p&gt;Jazz Mode suits Notion not just because it&amp;rsquo;s a good management concept, but because it matches Ivan&amp;rsquo;s own personality. He doesn&amp;rsquo;t like pure delegation, pure process, or pure management, so he needs an organization that allows improvisation. If a founder is a natural controller by nature, Jazz Mode will most likely fail in their company.&lt;/p&gt;
&lt;p&gt;For entrepreneurs, this may be the question most worth thinking about: does the organizational model you&amp;rsquo;re trying to build match your own personality? If it doesn&amp;rsquo;t, even the best concept is nothing but a castle in the air.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Original link:&lt;/strong&gt; &lt;a class=&#34;link&#34; href=&#34;https://mp.weixin.qq.com/s/62feH1VU4_Im_a4k24i1ew&#34;  target=&#34;_blank&#34; rel=&#34;noopener&#34;
    &gt;Notion CEO on the New Paradigm for AI-Era Organizational Management: Jazz Mode&lt;/a&gt;&lt;/p&gt;
</description>
        </item>
        <item>
        <title>OpenCode Go vs Command Code GOAT: Which $10 Coding AI Subscription Is Better Value?</title>
        <link>https://torchtree.com/en/post/opencode-go-vs-command-code-goat/</link>
        <pubDate>Mon, 17 Aug 2026 14:00:03 +0800</pubDate>
        
        <guid>https://torchtree.com/en/post/opencode-go-vs-command-code-goat/</guid>
        <description>&lt;img src="https://getnas.s3.bitiful.net/2026/08/openai_codex_gpt-image-2-medium_20260817_215228_80fba107.png" alt="Featured image of post OpenCode Go vs Command Code GOAT: Which $10 Coding AI Subscription Is Better Value?" /&gt;&lt;p&gt;OpenCode and Command Code are both terminal-based coding agents — open-source alternatives to Claude Code. You chat with an AI in your command line and it writes code, fixes bugs, and refactors projects for you. Both let you switch between underlying models from different vendors and don&amp;rsquo;t lock you into a single provider.&lt;/p&gt;
&lt;p&gt;OpenCode&amp;rsquo;s community is much larger, with 198K GitHub stars (measured via the GitHub API) and over 7.5 million monthly active developers. Command Code&amp;rsquo;s community is smaller — 3.7K GitHub stars (measured via the GitHub API) — but it invests more in harness engineering (the layer that optimizes tool calling for AI agents), such as a built-in Read tool that automatically filters out irrelevant file content to reduce context waste.&lt;/p&gt;
&lt;p&gt;Both have launched low-cost subscription plans aimed at individual developers. This article compares two plans that both cost $10/month: OpenCode Go and Command Code GOAT.&lt;/p&gt;
&lt;h2 id=&#34;pricing-and-credit-limits&#34;&gt;Pricing and Credit Limits
&lt;/h2&gt;&lt;p&gt;Both plans use the same pricing model: you pay $10 a month and receive a credit allowance denominated in US dollars, which is consumed at each model&amp;rsquo;s actual API price. Cheaper models burn credits slowly; expensive models burn them fast.&lt;/p&gt;
&lt;table&gt;
  &lt;thead&gt;
      &lt;tr&gt;
          &lt;th&gt;Dimension&lt;/th&gt;
          &lt;th&gt;OpenCode Go&lt;/th&gt;
          &lt;th&gt;Command Code GOAT&lt;/th&gt;
      &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
      &lt;tr&gt;
          &lt;td&gt;Monthly fee&lt;/td&gt;
          &lt;td&gt;$10 ($5 first month)&lt;/td&gt;
          &lt;td&gt;$10&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;Monthly credit&lt;/td&gt;
          &lt;td&gt;$60&lt;/td&gt;
          &lt;td&gt;$70&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;Credit multiplier&lt;/td&gt;
          &lt;td&gt;6x&lt;/td&gt;
          &lt;td&gt;7x&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;5-hour limit&lt;/td&gt;
          &lt;td&gt;$12&lt;/td&gt;
          &lt;td&gt;$14&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;Weekly limit&lt;/td&gt;
          &lt;td&gt;$30&lt;/td&gt;
          &lt;td&gt;$35&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;Top-ups&lt;/td&gt;
          &lt;td&gt;Yes (Zen balance)&lt;/td&gt;
          &lt;td&gt;Yes (never expires)&lt;/td&gt;
      &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;The &amp;ldquo;credit multiplier&amp;rdquo; is how much model usage your $10 actually buys. OpenCode Go gives $60 (6x), while Command Code GOAT gives $70 (7x) — 17% more.&lt;/p&gt;
&lt;h2 id=&#34;model-library-comparison&#34;&gt;Model Library Comparison
&lt;/h2&gt;&lt;p&gt;OpenCode Go offers 19 models; Command Code GOAT offers 36+.&lt;/p&gt;
&lt;p&gt;Shared core models include: DeepSeek V4 Pro/Flash, Kimi K3/K2.7 Code/K2.6, GLM-5.3/5.2/5.1, MiMo-V2.5/V2.5-Pro, MiniMax M3/M2.7, Qwen 3.8 Max/3.7 Max/3.7 Plus/3.6 Plus, GPT 5.6 Luna, Grok 4.5, Tencent Hy3.&lt;/p&gt;
&lt;p&gt;These shared models are priced identically on both platforms. For example, DeepSeek V4 Pro is off-peak $0.66 input/$1.98 output per million tokens on both, and MiMo V2.5 is $0.14 input/$0.28 output per million tokens on both.&lt;/p&gt;
&lt;p&gt;Models exclusive to Command Code GOAT include: Gemini 3.7 Flash (limited-time 50% discount), Grok 4.6, Muse Spark 1.2, Laguna S 2.1 (free, limited capacity), Inkling, Step 3.7/3.5 Flash, Qwen 3.7 Flash, and more MiniMax variants.&lt;/p&gt;
&lt;h2 id=&#34;exclusive-discounts&#34;&gt;Exclusive Discounts
&lt;/h2&gt;&lt;p&gt;The general model pricing is the same on both platforms, but Command Code GOAT has a series of permanent or time-limited discounts:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;MiMo V2.5 Pro: permanent 99% discount.&lt;/strong&gt; It has a dedicated $30 monthly quota on GOAT, equivalent to roughly $150 of usage at the discounted rate. On OpenCode Go, this model consumes credits at standard pricing with no extra benefit.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;MiMo V2.5: permanent 98% discount.&lt;/strong&gt; Output drops from about $4.00 to $0.28 per million tokens, and input from about $0.80 to $0.14. Using this model on GOAT, a $10 subscription is equivalent to roughly $100 of standard-priced usage.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;MiniMax M3: permanent 50% discount.&lt;/strong&gt; Your credit effectively doubles.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Laguna S 2.1: completely free.&lt;/strong&gt; This is Poolside&amp;rsquo;s open-source coding model — all tokens (input, output, cached reads) cost $0, with limited capacity on a first-come, first-served basis.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Gemini 3.7 Flash: 50% off until December 31, 2026.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;OpenCode Go has no exclusive discounts; all models consume credits at standard pricing.&lt;/p&gt;
&lt;p&gt;This means that if you frequently use MiMo or MiniMax models, Command Code GOAT&amp;rsquo;s effective capacity far exceeds the nominal $70. For heavy users of these models, GOAT&amp;rsquo;s value advantage goes well beyond 17% — the real gap can be several times.&lt;/p&gt;
&lt;h2 id=&#34;api-access&#34;&gt;API Access
&lt;/h2&gt;&lt;p&gt;Both platforms offer API access, but the implementations differ.&lt;/p&gt;
&lt;p&gt;OpenCode Go&amp;rsquo;s API endpoint is &lt;code&gt;https://opencode.ai/zen/go/v1/chat/completions&lt;/code&gt; (OpenAI-compatible format). You can call it directly with an API key, or configure it as a custom provider in other coding agents.&lt;/p&gt;
&lt;p&gt;Command Code GOAT similarly offers API endpoints compatible with both OpenAI Chat Completions and Anthropic Messages, with a single API key used for both the CLI and external calls.&lt;/p&gt;
&lt;p&gt;Both platforms&amp;rsquo; APIs can plug into other toolchains. For developers who only write code in the terminal, this difference barely matters. If your workflow involves multiple tools sharing one model subscription, both work — the difference comes down to endpoint format and integration convenience.&lt;/p&gt;
&lt;h2 id=&#34;how-to-choose&#34;&gt;How to Choose
&lt;/h2&gt;&lt;p&gt;&lt;strong&gt;Already in the Command Code ecosystem:&lt;/strong&gt; Upgrading to GOAT is the natural choice — more credit and exclusive discounts right away.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Value model diversity:&lt;/strong&gt; GOAT&amp;rsquo;s 36+ models nearly double OpenCode Go&amp;rsquo;s 19, and the extras include popular options like Gemini 3.7 Flash and Grok 4.6.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Frequently use the MiMo line:&lt;/strong&gt; GOAT&amp;rsquo;s permanent discounts are the decisive advantage. MiMo V2.5 Pro at 99% off and MiMo V2.5 at 98% off mean the same money goes several times further.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Just need a standalone CLI coding tool:&lt;/strong&gt; OpenCode Go is plenty. 19 models cover mainstream needs, and $60 of monthly credit is enough for moderate usage.&lt;/p&gt;
&lt;p&gt;The real choice depends on which models you use most. Both platforms promise not to train on your code, both accept credit cards, and extra purchased credits never expire on either.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Pricing information comes from the official OpenCode and Command Code pricing pages (August 2026); actual prices may change — refer to official announcements for the latest.&lt;/em&gt;&lt;/p&gt;
&lt;h2 id=&#34;sources&#34;&gt;Sources
&lt;/h2&gt;&lt;ul&gt;
&lt;li&gt;&lt;a class=&#34;link&#34; href=&#34;https://opencode.ai/docs/go/&#34;  target=&#34;_blank&#34; rel=&#34;noopener&#34;
    &gt;OpenCode Go official docs&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a class=&#34;link&#34; href=&#34;https://commandcode.ai/docs/plans/goat&#34;  target=&#34;_blank&#34; rel=&#34;noopener&#34;
    &gt;Command Code GOAT plan docs&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a class=&#34;link&#34; href=&#34;https://commandcode.ai/docs/resources/pricing-limits&#34;  target=&#34;_blank&#34; rel=&#34;noopener&#34;
    &gt;Command Code pricing and limits&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a class=&#34;link&#34; href=&#34;https://github.com/anomalyco/opencode&#34;  target=&#34;_blank&#34; rel=&#34;noopener&#34;
    &gt;OpenCode GitHub repository&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a class=&#34;link&#34; href=&#34;https://github.com/CommandCodeAI/command-code&#34;  target=&#34;_blank&#34; rel=&#34;noopener&#34;
    &gt;Command Code GitHub repository&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
</description>
        </item>
        <item>
        <title>Advanced Pi Agent Configuration: AGENTS.md, Model Switching, and Thinking Levels in Practice</title>
        <link>https://torchtree.com/en/post/pi-agent-configuration-guide/</link>
        <pubDate>Wed, 22 Jul 2026 12:14:08 +0800</pubDate>
        
        <guid>https://torchtree.com/en/post/pi-agent-configuration-guide/</guid>
        <description>&lt;img src="https://getnas.s3.bitiful.net/2026/07/pi-agent-config-cover.png" alt="Featured image of post Advanced Pi Agent Configuration: AGENTS.md, Model Switching, and Thinking Levels in Practice" /&gt;&lt;p&gt;In the article &lt;a class=&#34;link&#34; href=&#34;https://hitorch.cn/pi-agent-setup-guide/&#34;  target=&#34;_blank&#34; rel=&#34;noopener&#34;
    &gt;Pi Coding Agent in Practice: From Installation to Everyday Use&lt;/a&gt;, I covered Pi&amp;rsquo;s installation, first-time configuration, and its core extension packages. That content lets beginners get up and running quickly, but Pi&amp;rsquo;s real flexibility lives in its configuration files.&lt;/p&gt;
&lt;p&gt;Pi&amp;rsquo;s core is just 418 lines of TypeScript, and by default it only gives the model four tools (read, write, edit, bash). All of its advanced behavior — which model to use, how large a context, how deep to think — is controlled through external configuration files. Understanding how these files relate to each other and how they&amp;rsquo;re prioritized is the key step in taking Pi from &amp;ldquo;usable&amp;rdquo; to &amp;ldquo;actually good.&amp;rdquo;&lt;/p&gt;
&lt;h2 id=&#34;how-many-layers-does-pis-configuration-have-and-what-does-each-one-manage&#34;&gt;How many layers does Pi&amp;rsquo;s configuration have, and what does each one manage?
&lt;/h2&gt;&lt;p&gt;Pi&amp;rsquo;s configuration system uses a layered, additive design. Once you understand what each layer is responsible for and its order of precedence, you won&amp;rsquo;t run into &amp;ldquo;I changed it but nothing happened.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Global configuration&lt;/strong&gt; lives in &lt;code&gt;~/.pi/agent/&lt;/code&gt;, affecting all projects. &lt;strong&gt;Project configuration&lt;/strong&gt; lives in a project&amp;rsquo;s &lt;code&gt;.pi/settings.json&lt;/code&gt;, affecting only the current project. Nested objects in the project config are merged with the global config rather than replacing it entirely.&lt;/p&gt;
&lt;p&gt;Beyond the two layers of settings.json, Pi also uses four specialized configuration files, each with a different purpose:&lt;/p&gt;
&lt;table&gt;
  &lt;thead&gt;
      &lt;tr&gt;
          &lt;th&gt;Config file&lt;/th&gt;
          &lt;th&gt;Location&lt;/th&gt;
          &lt;th&gt;Purpose&lt;/th&gt;
          &lt;th&gt;Load timing&lt;/th&gt;
      &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
      &lt;tr&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;AGENTS.md&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
          &lt;td&gt;Project root or &lt;!-- raw HTML omitted --&gt;~/.pi/agent/&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
          &lt;td&gt;Project context and coding instructions, injected into the system prompt&lt;/td&gt;
          &lt;td&gt;Auto-loaded at startup&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;APPEND_SYSTEM.md&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;~/.pi/agent/&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
          &lt;td&gt;Global behavior rules, appended to the end of the system prompt&lt;/td&gt;
          &lt;td&gt;Loaded at startup&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;settings.json&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
          &lt;td&gt;Global or &lt;!-- raw HTML omitted --&gt;.pi/&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
          &lt;td&gt;Model selection, UI theme, compaction strategy, retries, and other runtime parameters&lt;/td&gt;
          &lt;td&gt;Loaded at startup&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;models.json&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;~/.pi/agent/&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
          &lt;td&gt;Custom models and Providers (Ollama, vLLM, etc.)&lt;/td&gt;
          &lt;td&gt;Reloaded each time you open &lt;!-- raw HTML omitted --&gt;/model&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;auth.json&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;~/.pi/agent/&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
          &lt;td&gt;API keys and OAuth credentials (permission 0600)&lt;/td&gt;
          &lt;td&gt;Read on demand&lt;/td&gt;
      &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Below I&amp;rsquo;ll break down best practices for each configuration file in turn.&lt;/p&gt;
&lt;h2 id=&#34;how-to-write-an-agentsmd-that-actually-works&#34;&gt;How to write an AGENTS.md that actually works
&lt;/h2&gt;&lt;p&gt;&lt;code&gt;AGENTS.md&lt;/code&gt; is the primary entry point for Pi to understand a project&amp;rsquo;s context. When Pi starts, it looks in several locations and merges their contents into the system prompt: it loads &lt;code&gt;~/.pi/agent/AGENTS.md&lt;/code&gt; first (global instructions), then walks up through parent directories, and finally loads the &lt;code&gt;AGENTS.md&lt;/code&gt; in the current directory.&lt;/p&gt;
&lt;p&gt;In other words, &lt;strong&gt;the global AGENTS.md defines your typical tech stack and general conventions as a developer, while the project AGENTS.md defines that specific project&amp;rsquo;s constraints and workflow.&lt;/strong&gt;&lt;/p&gt;
&lt;h3 id=&#34;what-to-put-in-the-global-agentsmd&#34;&gt;What to put in the global AGENTS.md
&lt;/h3&gt;&lt;p&gt;The global &lt;code&gt;~/.pi/agent/AGENTS.md&lt;/code&gt; is a good place to record the tech-stack preferences you use day to day. The global AGENTS.md that DeepakNess shares on his blog is a great reference example:&lt;/p&gt;
&lt;p&gt;This content comes from DeepakNess&amp;rsquo;s article, &lt;a class=&#34;link&#34; href=&#34;https://deepakness.com/blog/pi-agent-setup/&#34;  target=&#34;_blank&#34; rel=&#34;noopener&#34;
    &gt;Setting Up and Using the Pi Coding Agent&lt;/a&gt;. It doesn&amp;rsquo;t try to be overly specific — instead it gives broad tech-stack hints while asking Pi to defer to the project-level AGENTS.md first.&lt;/p&gt;
&lt;h3 id=&#34;how-to-organize-a-project-agentsmd&#34;&gt;How to organize a project AGENTS.md
&lt;/h3&gt;&lt;p&gt;The project-level AGENTS.md needs to be more precise. Here&amp;rsquo;s a template tailored for a TypeScript project, based on the recommendations in the &lt;a class=&#34;link&#34; href=&#34;https://pi.dev/docs/latest&#34;  target=&#34;_blank&#34; rel=&#34;noopener&#34;
    &gt;official Pi documentation&lt;/a&gt;:&lt;/p&gt;
&lt;p&gt;If you manage multiple projects, it&amp;rsquo;s worth keeping a template for each project type. Copy it over each time you start a new project and adjust the tech-stack fields as needed.&lt;/p&gt;
&lt;h3 id=&#34;remember-to-reload-after-changes&#34;&gt;Remember to reload after changes
&lt;/h3&gt;&lt;p&gt;Every time you modify AGENTS.md, you need to run &lt;code&gt;/reload&lt;/code&gt; or restart Pi for the change to take effect. This operation isn&amp;rsquo;t triggered often, but it&amp;rsquo;s easy to forget. It&amp;rsquo;s best to test immediately after writing a new rule to confirm Pi&amp;rsquo;s behavior changed as expected.&lt;/p&gt;
&lt;h2 id=&#34;what-does-append_systemmd-control&#34;&gt;What does APPEND_SYSTEM.md control?
&lt;/h2&gt;&lt;p&gt;&lt;code&gt;~/.pi/agent/APPEND_SYSTEM.md&lt;/code&gt; is appended to the end of the system prompt, and it takes precedence over &lt;code&gt;AGENTS.md&lt;/code&gt;. This means its instructions override what came before.&lt;/p&gt;
&lt;p&gt;It&amp;rsquo;s a good fit for defining &lt;strong&gt;behavioral guidelines that apply across projects&lt;/strong&gt;, especially constraints about how the agent works and how it interacts with the user. Drawing on official recommendations and community practice, a typical APPEND_SYSTEM.md looks like this:&lt;/p&gt;
&lt;p&gt;These rules ensure Pi maintains a consistent way of working across projects, without having to restate everything at the start of every conversation.&lt;/p&gt;
&lt;h2 id=&#34;settingsjson-common-options-explained&#34;&gt;settings.json: common options explained
&lt;/h2&gt;&lt;p&gt;&lt;code&gt;settings.json&lt;/code&gt; has two layers: global (&lt;code&gt;~/.pi/agent/settings.json&lt;/code&gt;) and project (&lt;code&gt;.pi/settings.json&lt;/code&gt;). Nested objects in the project layer merge with the global layer. Drawing on the &lt;a class=&#34;link&#34; href=&#34;https://pi.dev/docs/latest/settings&#34;  target=&#34;_blank&#34; rel=&#34;noopener&#34;
    &gt;official Pi Settings documentation&lt;/a&gt;, here are the settings most worth knowing:&lt;/p&gt;
&lt;h3 id=&#34;models-and-thinking-levels&#34;&gt;Models and thinking levels
&lt;/h3&gt;&lt;p&gt;&lt;code&gt;defaultProvider&lt;/code&gt; and &lt;code&gt;defaultModel&lt;/code&gt; control the model Pi uses by default at startup. &lt;code&gt;defaultThinkingLevel&lt;/code&gt; sets the thinking depth, with options including &lt;code&gt;off&lt;/code&gt;, &lt;code&gt;minimal&lt;/code&gt;, &lt;code&gt;low&lt;/code&gt;, &lt;code&gt;medium&lt;/code&gt;, &lt;code&gt;high&lt;/code&gt;, &lt;code&gt;xhigh&lt;/code&gt;, and &lt;code&gt;max&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;&lt;code&gt;enabledModels&lt;/code&gt; is a high-impact setting. It defines the list of models cycled through by &lt;code&gt;Ctrl+P&lt;/code&gt;, and supports wildcards. If you leave it unset, &lt;code&gt;Ctrl+P&lt;/code&gt; iterates over every available model for that provider, which hurts the experience significantly.&lt;/p&gt;
&lt;p&gt;&lt;code&gt;thinkingBudgets&lt;/code&gt; lets you customize the token budget for each thinking level. The numbers above follow the defaults given in the official Pi documentation. Whether to adjust them depends on your model and task: the higher the budget, the deeper the thinking, and the more tokens consumed.&lt;/p&gt;
&lt;h3 id=&#34;context-compaction&#34;&gt;Context compaction
&lt;/h3&gt;&lt;p&gt;Compaction is Pi&amp;rsquo;s core mechanism for handling long contexts. When the session approaches the context limit, Pi automatically summarizes older messages to free up space for subsequent conversation.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;code&gt;reserveTokens&lt;/code&gt;: the number of tokens reserved for the LLM&amp;rsquo;s response (default 16384). The smaller this value, the sooner compaction happens.&lt;/li&gt;
&lt;li&gt;&lt;code&gt;keepRecentTokens&lt;/code&gt;: the number of recent tokens kept without being summarized (default 20000). Kept messages stay intact, ensuring the most recent discussion isn&amp;rsquo;t lost.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;If you regularly handle long sessions, you can raise &lt;code&gt;keepRecentTokens&lt;/code&gt;. Note, though, that this reduces compaction efficiency and may hit the model&amp;rsquo;s context-window ceiling sooner.&lt;/p&gt;
&lt;h3 id=&#34;retry-strategy&#34;&gt;Retry strategy
&lt;/h3&gt;&lt;p&gt;The retry configuration is split into two layers: agent-level retries (handled by Pi itself) and provider-level retries (handled by the API SDK). The official documentation recommends keeping &lt;code&gt;retry.provider.maxRetries&lt;/code&gt; at 0, because provider-level retries can burn quota before you even see a rate-limit error.&lt;/p&gt;
&lt;p&gt;&lt;code&gt;baseDelayMs&lt;/code&gt; controls the initial delay of exponential backoff: 2s → 4s → 8s. For tasks that need to run stably over a long time (for example, batch data crawling), you can reasonably increase this value.&lt;/p&gt;
&lt;h3 id=&#34;project-trust-mode&#34;&gt;Project trust mode
&lt;/h3&gt;&lt;p&gt;When Pi first starts in a project, it asks whether to trust that project&amp;rsquo;s &lt;code&gt;.pi/&lt;/code&gt; directory. This mechanism exists to prevent malicious project plugins from auto-loading.&lt;/p&gt;
&lt;p&gt;Available values include &lt;code&gt;ask&lt;/code&gt; (ask every time, the default), &lt;code&gt;always&lt;/code&gt; (auto-trust), and &lt;code&gt;never&lt;/code&gt; (never trust). In CI or automation scenarios, you can use the &lt;code&gt;-a&lt;/code&gt; / &lt;code&gt;--approve&lt;/code&gt; flag to skip the prompt.&lt;/p&gt;
&lt;h2 id=&#34;adding-custom-models-with-modelsjson&#34;&gt;Adding custom models with models.json
&lt;/h2&gt;&lt;p&gt;If your model isn&amp;rsquo;t among Pi&amp;rsquo;s built-in 20-plus Providers, you can add it via &lt;code&gt;~/.pi/agent/models.json&lt;/code&gt;. This file supports Ollama, LM Studio, vLLM, OpenRouter, Cloudflare AI Gateway, and any OpenAI-compatible API endpoint.&lt;/p&gt;
&lt;p&gt;Referencing the full configuration notes in the &lt;a class=&#34;link&#34; href=&#34;https://pi.dev/docs/latest/models&#34;  target=&#34;_blank&#34; rel=&#34;noopener&#34;
    &gt;official Pi Models documentation&lt;/a&gt;, here are the three most common scenarios:&lt;/p&gt;
&lt;h3 id=&#34;scenario-1-a-local-model-via-ollama&#34;&gt;Scenario 1: a local model via Ollama
&lt;/h3&gt;&lt;p&gt;&lt;code&gt;apiKey&lt;/code&gt; being set to &lt;code&gt;&amp;quot;ollama&amp;quot;&lt;/code&gt; is just a placeholder. Ollama doesn&amp;rsquo;t validate the API key, but Pi needs an auth value to show the model in &lt;code&gt;/model&lt;/code&gt;. The two switches under &lt;code&gt;compat&lt;/code&gt; target Ollama&amp;rsquo;s characteristics: it doesn&amp;rsquo;t support the developer role or the reasoning_effort parameter.&lt;/p&gt;
&lt;p&gt;The &lt;code&gt;name&lt;/code&gt; field gives a human-readable label. Pi uses this value both in the model selector and when matching the &lt;code&gt;--model&lt;/code&gt; mode.&lt;/p&gt;
&lt;h3 id=&#34;scenario-2-openrouter-routing-configuration&#34;&gt;Scenario 2: OpenRouter routing configuration
&lt;/h3&gt;&lt;p&gt;OpenRouter lets you set routing preferences among multiple API providers. The configuration below follows the OpenRouter example in the official Pi documentation:&lt;/p&gt;
&lt;p&gt;The &lt;code&gt;openRouterRouting&lt;/code&gt; object is passed through verbatim to the &lt;code&gt;provider&lt;/code&gt; field of the OpenRouter API. &lt;code&gt;order&lt;/code&gt; specifies provider priority, and &lt;code&gt;data_collection: &amp;quot;deny&amp;quot;&lt;/code&gt; declines to use your data for training.&lt;/p&gt;
&lt;h3 id=&#34;scenario-3-proxying-the-anthropic-api&#34;&gt;Scenario 3: proxying the Anthropic API
&lt;/h3&gt;&lt;p&gt;If you use a third-party proxy for the Anthropic Messages API, you can configure it like this:&lt;/p&gt;
&lt;p&gt;In &lt;code&gt;models.json&lt;/code&gt;, &lt;code&gt;apiKey&lt;/code&gt; and &lt;code&gt;headers&lt;/code&gt; support three value-resolution modes: a direct literal, &lt;code&gt;$ENV_VAR&lt;/code&gt; environment-variable interpolation, or &lt;code&gt;!command&lt;/code&gt; command execution. Bitdoze notes in the &lt;a class=&#34;link&#34; href=&#34;https://www.bitdoze.com/pi-coding-agent-setup-guide/&#34;  target=&#34;_blank&#34; rel=&#34;noopener&#34;
    &gt;Pi Coding Agent Setup Guide&lt;/a&gt; that Pi supports Ollama, LM Studio, vLLM, and any OpenAI-compatible endpoint. This extensibility is a major advantage over comparable tools.&lt;/p&gt;
&lt;h2 id=&#34;managing-api-credentials-with-authjson&#34;&gt;Managing API credentials with auth.json
&lt;/h2&gt;&lt;p&gt;&lt;code&gt;~/.pi/agent/auth.json&lt;/code&gt; stores the API keys and OAuth tokens for all providers. Its permission is set to &lt;code&gt;0600&lt;/code&gt;, allowing only the current user to read and write it.&lt;/p&gt;
&lt;p&gt;auth.json supports three ways of resolving keys:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Literal&lt;/strong&gt;: use the API key string directly&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Environment-variable interpolation&lt;/strong&gt;: &lt;code&gt;&amp;quot;$MY_KEY&amp;quot;&lt;/code&gt; or &lt;code&gt;&amp;quot;${KEY_PREFIX}_${KEY_SUFFIX}&amp;quot;&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Shell command&lt;/strong&gt;: &lt;code&gt;&amp;quot;!security find-generic-password -ws &#39;anthropic&#39;&amp;quot;&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The official documentation notes that auth.json takes precedence over environment variables. This means if you&amp;rsquo;ve set both a &lt;code&gt;DEEPSEEK_API_KEY&lt;/code&gt; environment variable and a deepseek entry in auth.json, the latter overrides the former.&lt;/p&gt;
&lt;p&gt;One practical tip is to use a shell command to read the credential from the system keychain, avoiding writing your API key in plaintext to any file:&lt;/p&gt;
&lt;h2 id=&#34;model-switching-strategy-which-model-to-use-when&#34;&gt;Model-switching strategy: which model to use when
&lt;/h2&gt;&lt;p&gt;One of Pi&amp;rsquo;s core strengths is model-agnosticism. You can pick a different model for different tasks, and switching is instantaneous.&lt;/p&gt;
&lt;h3 id=&#34;how-to-switch&#34;&gt;How to switch
&lt;/h3&gt;&lt;table&gt;
  &lt;thead&gt;
      &lt;tr&gt;
          &lt;th&gt;Action&lt;/th&gt;
          &lt;th&gt;Shortcut / command&lt;/th&gt;
          &lt;th&gt;Notes&lt;/th&gt;
      &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
      &lt;tr&gt;
          &lt;td&gt;Open the model selector&lt;/td&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;Ctrl+L&lt;!-- raw HTML omitted --&gt; or &lt;!-- raw HTML omitted --&gt;/model&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
          &lt;td&gt;Quickly switch models&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;Cycle through models&lt;/td&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;Ctrl+P&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
          &lt;td&gt;Rotate through the &lt;!-- raw HTML omitted --&gt;enabledModels&lt;!-- raw HTML omitted --&gt; list&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;Adjust thinking level&lt;/td&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;Shift+Tab&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
          &lt;td&gt;Toggle thinking depth&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;Interrupt the current action&lt;/td&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;Escape&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
          &lt;td&gt;Cancel the running task&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;Send a steering message&lt;/td&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;Enter&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
          &lt;td&gt;Interrupt the agent&amp;rsquo;s current workflow and respond immediately&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;Send a follow-up message&lt;/td&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;Alt+Enter&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
          &lt;td&gt;Append a message after the agent finishes its work&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;Quit&lt;/td&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;Ctrl+C&lt;!-- raw HTML omitted --&gt; (press twice)&lt;/td&gt;
          &lt;td&gt;Exit Pi&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;Reference a file&lt;/td&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;@&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
          &lt;td&gt;Fuzzy-search files&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;Run a command&lt;/td&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;!&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
          &lt;td&gt;Send a command&amp;rsquo;s output to the model&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;Silent command&lt;/td&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;!!&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
          &lt;td&gt;Run a command without adding it to context&lt;/td&gt;
      &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;The shortcut table partly draws on the &lt;a class=&#34;link&#34; href=&#34;https://pi-agent.org/&#34;  target=&#34;_blank&#34; rel=&#34;noopener&#34;
    &gt;Pi Agent Chinese guide&lt;/a&gt; and DeepakNess&amp;rsquo;s setup article.&lt;/p&gt;
&lt;h3 id=&#34;a-recommended-layered-strategy&#34;&gt;A recommended layered strategy
&lt;/h3&gt;&lt;p&gt;Experience from multiple community users converges on the same pattern: use models in layers, matching capability to task complexity. The following is drawn from DeepakNess and Bitdoze&amp;rsquo;s articles:&lt;/p&gt;
&lt;table&gt;
  &lt;thead&gt;
      &lt;tr&gt;
          &lt;th&gt;Task type&lt;/th&gt;
          &lt;th&gt;Recommended model&lt;/th&gt;
          &lt;th&gt;Thinking level&lt;/th&gt;
          &lt;th&gt;Reasoning&lt;/th&gt;
      &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
      &lt;tr&gt;
          &lt;td&gt;Quick edits, file operations, batch scripts&lt;/td&gt;
          &lt;td&gt;DeepSeek V4 Flash / MiniMax M2.7&lt;/td&gt;
          &lt;td&gt;low or off&lt;/td&gt;
          &lt;td&gt;Extremely cheap, plenty for fast tasks&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;Everyday coding, small-to-medium refactors&lt;/td&gt;
          &lt;td&gt;DeepSeek V4 Pro / Qwen 3.6 Plus&lt;/td&gt;
          &lt;td&gt;medium&lt;/td&gt;
          &lt;td&gt;Balances quality and cost&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;Deep analysis, architecture design, complex debugging&lt;/td&gt;
          &lt;td&gt;DeepSeek V4 Pro / Claude Sonnet 4&lt;/td&gt;
          &lt;td&gt;high or xhigh&lt;/td&gt;
          &lt;td&gt;Needs a deeper reasoning chain&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;Visual tasks (screenshot understanding, UI analysis)&lt;/td&gt;
          &lt;td&gt;Kimi K3 / Claude&lt;/td&gt;
          &lt;td&gt;depends on the model&lt;/td&gt;
          &lt;td&gt;Proxied via pi-vision-proxy when the main model has no vision&lt;/td&gt;
      &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;DeepakNess provides a concrete data point in his article: crawling 285,000 URLs with DeepSeek V4 Flash took about 1.5 hours with a total cost of $1. That illustrates the cost-effectiveness of low thinking level plus a cheap model on batch tasks.&lt;/p&gt;
&lt;h3 id=&#34;enabledmodels-wildcards&#34;&gt;enabledModels wildcards
&lt;/h3&gt;&lt;p&gt;To make &lt;code&gt;Ctrl+P&lt;/code&gt; switching more efficient, it&amp;rsquo;s worth setting an &lt;code&gt;enabledModels&lt;/code&gt; list in settings.json:&lt;/p&gt;
&lt;p&gt;Wildcards match all qualifying models. If you only need two or three specific models, you can also write exact IDs:&lt;/p&gt;
&lt;h2 id=&#34;how-thinking-level-affects-output-quality&#34;&gt;How Thinking Level affects output quality
&lt;/h2&gt;&lt;p&gt;Pi&amp;rsquo;s thinking level is a layered parameter that controls how deeply the model reasons before answering. Based on the official Settings docs at pi.dev and the thinkingLevelMap explanation in models.md, different levels correspond to different behavioral characteristics:&lt;/p&gt;
&lt;table&gt;
  &lt;thead&gt;
      &lt;tr&gt;
          &lt;th&gt;Level&lt;/th&gt;
          &lt;th&gt;Use case&lt;/th&gt;
          &lt;th&gt;Token budget (default)&lt;/th&gt;
      &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
      &lt;tr&gt;
          &lt;td&gt;off&lt;/td&gt;
          &lt;td&gt;Simple Q&amp;amp;A, tasks that need no reasoning&lt;/td&gt;
          &lt;td&gt;no reasoning tokens&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;minimal&lt;/td&gt;
          &lt;td&gt;Very simple judgments, such as &amp;ldquo;yes/no&amp;rdquo; classification&lt;/td&gt;
          &lt;td&gt;1024&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;low&lt;/td&gt;
          &lt;td&gt;Light reasoning, e.g. formatting, simple conversions&lt;/td&gt;
          &lt;td&gt;4096&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;medium&lt;/td&gt;
          &lt;td&gt;Routine coding tasks&lt;/td&gt;
          &lt;td&gt;10240&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;high&lt;/td&gt;
          &lt;td&gt;Complex refactors, debugging&lt;/td&gt;
          &lt;td&gt;32768&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;xhigh&lt;/td&gt;
          &lt;td&gt;Deep analysis, architecture design&lt;/td&gt;
          &lt;td&gt;65536&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;max&lt;/td&gt;
          &lt;td&gt;Extremely complex multi-step reasoning&lt;/td&gt;
          &lt;td&gt;provider cap&lt;/td&gt;
      &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;The differences between levels aren&amp;rsquo;t linear. The biggest quality gain is from low to medium; from high to xhigh, the marginal returns diminish. In practice, 80% of everyday tasks get satisfactory results at the medium level.&lt;/p&gt;
&lt;p&gt;For models that support thinkingLevelMap, you can finely control in models.json which provider-side parameter each level maps to. For example, a given model might only need the high and max levels, with the middle levels skipped:&lt;/p&gt;
&lt;p&gt;This mechanism comes from the thinkingLevelMap explanation in the Pi Models documentation. When a model doesn&amp;rsquo;t support certain levels, Pi automatically jumps to the adjacent supported level.&lt;/p&gt;
&lt;h2 id=&#34;context-management-compaction-session-trees-and-manual-control&#34;&gt;Context management: compaction, session trees, and manual control
&lt;/h2&gt;&lt;p&gt;Long sessions are the norm for coding agents. Pi provides three layers of context-management mechanisms.&lt;/p&gt;
&lt;h3 id=&#34;automatic-compaction&#34;&gt;Automatic compaction
&lt;/h3&gt;&lt;p&gt;Compaction runs in the background. When the context approaches the model&amp;rsquo;s window limit, Pi automatically summarizes older messages. &lt;code&gt;compaction.reserveTokens&lt;/code&gt; controls when compaction triggers: it fires when the remaining tokens drop below this value. &lt;code&gt;compaction.keepRecentTokens&lt;/code&gt; ensures the most recent messages aren&amp;rsquo;t summarized.&lt;/p&gt;
&lt;p&gt;If you want finer control, you can trigger &lt;code&gt;/compact&lt;/code&gt; manually and Pi will immediately compact the current session.&lt;/p&gt;
&lt;h3 id=&#34;session-tree-management&#34;&gt;Session-tree management
&lt;/h3&gt;&lt;p&gt;The &lt;code&gt;/tree&lt;/code&gt; command displays the session history as a tree structure. Each branch represents a conversation path. Pi supports:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;code&gt;/resume&lt;/code&gt;: pick up a previous session and continue working&lt;/li&gt;
&lt;li&gt;&lt;code&gt;/new&lt;/code&gt;: start a new session&lt;/li&gt;
&lt;li&gt;&lt;code&gt;/fork&lt;/code&gt;: branch from the current session to begin a new line of conversation&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This design lets you try different solution paths without losing context.&lt;/p&gt;
&lt;h2 id=&#34;a-complete-configuration-template&#34;&gt;A complete configuration template
&lt;/h2&gt;&lt;p&gt;Combining all of the above, here&amp;rsquo;s a complete configuration you can put into daily use.&lt;/p&gt;
&lt;h3 id=&#34;global-settingsjson&#34;&gt;Global settings.json
&lt;/h3&gt;&lt;h3 id=&#34;global-append_systemmd&#34;&gt;Global APPEND_SYSTEM.md
&lt;/h3&gt;&lt;h3 id=&#34;project-pisettingsjson-overriding-the-global-compaction-strategy&#34;&gt;Project .pi/settings.json (overriding the global compaction strategy)
&lt;/h3&gt;&lt;p&gt;This override makes short-session projects that need frequent compaction trigger it earlier, avoiding wasted context window.&lt;/p&gt;
&lt;h2 id=&#34;summary&#34;&gt;Summary
&lt;/h2&gt;&lt;p&gt;Pi&amp;rsquo;s configuration system revolves around one core principle: &lt;strong&gt;layer on layer, precise control&lt;/strong&gt;. Global configuration defines general behavior, project configuration overrides specific needs, AGENTS.md conveys project context, and APPEND_SYSTEM.md constrains the agent&amp;rsquo;s behavioral patterns.&lt;/p&gt;
&lt;p&gt;Once you understand what each layer is responsible for and its priority, Pi&amp;rsquo;s &amp;ldquo;minimal core + external configuration&amp;rdquo; design philosophy stops being a &amp;ldquo;too few features&amp;rdquo; weakness and becomes a &amp;ldquo;you control everything&amp;rdquo; strength. When facing different tasks each day, you only need to switch models with &lt;code&gt;Ctrl+P&lt;/code&gt; and adjust thinking depth with &lt;code&gt;Shift+Tab&lt;/code&gt; to move quickly between different working modes.&lt;/p&gt;
&lt;p&gt;If your configuration already covers the main files mentioned in this article, the next step is to focus on the extension system: use &lt;code&gt;pi install&lt;/code&gt; to add packages such as pi-web-access (web search), pi-codex-goal (task tracking), and pi-vision-proxy (vision proxy), gradually building a Pi environment fully suited to your own workflow.&lt;/p&gt;
&lt;p&gt;Sources:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a class=&#34;link&#34; href=&#34;https://pi.dev/docs/latest&#34;  target=&#34;_blank&#34; rel=&#34;noopener&#34;
    &gt;Pi official documentation&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a class=&#34;link&#34; href=&#34;https://pi.dev/docs/latest/settings&#34;  target=&#34;_blank&#34; rel=&#34;noopener&#34;
    &gt;Pi official Settings documentation&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a class=&#34;link&#34; href=&#34;https://pi.dev/docs/latest/providers&#34;  target=&#34;_blank&#34; rel=&#34;noopener&#34;
    &gt;Pi official Providers documentation&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a class=&#34;link&#34; href=&#34;https://pi.dev/docs/latest/models&#34;  target=&#34;_blank&#34; rel=&#34;noopener&#34;
    &gt;Pi official Models documentation&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a class=&#34;link&#34; href=&#34;https://deepakness.com/blog/pi-agent-setup/&#34;  target=&#34;_blank&#34; rel=&#34;noopener&#34;
    &gt;DeepakNess: Setting Up and Using the Pi Coding Agent&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a class=&#34;link&#34; href=&#34;https://www.bitdoze.com/pi-coding-agent-setup-guide/&#34;  target=&#34;_blank&#34; rel=&#34;noopener&#34;
    &gt;Bitdoze: Pi Coding Agent Setup Guide&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a class=&#34;link&#34; href=&#34;https://pi-agent.org/&#34;  target=&#34;_blank&#34; rel=&#34;noopener&#34;
    &gt;Pi Agent Chinese guide&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
</description>
        </item>
        <item>
        <title>Malus.sh: How AI Clean-Room Clones Threaten the Open Source Sustainability Loop</title>
        <link>https://torchtree.com/en/post/malus-sh-clean-room-ai-open-source/</link>
        <pubDate>Tue, 05 May 2026 07:16:47 +0800</pubDate>
        
        <guid>https://torchtree.com/en/post/malus-sh-clean-room-ai-open-source/</guid>
        <description>&lt;h2 id=&#34;what-is-malussh&#34;&gt;What is Malus.sh
&lt;/h2&gt;&lt;p&gt;Malus.sh (pronounced like &amp;ldquo;malice&amp;rdquo;) is an AI-powered tool that claims to recreate functional equivalents of any open source software from scratch using &amp;ldquo;Clean Room&amp;rdquo; engineering methods, while stripping away all license obligations of the original project. Its most eye-catching slogan: &amp;ldquo;No attribution. No copyleft. No problems.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;What makes the project unique is that &lt;strong&gt;it is both a satirical work and a genuinely operating commercial product&lt;/strong&gt;. Founder Mike Nolan works as a researcher on the political economy of open source at the United Nations. In an interview with 404 Media, he explicitly stated that if it were only satire, open source practitioners would dismiss it with &amp;ldquo;that can&amp;rsquo;t happen to me.&amp;rdquo; Making the tool actually usable forces the community to confront the structural cracks that already exist in the open source economic model.&lt;/p&gt;
&lt;p&gt;Malus.sh is registered as an LLC, accepts payments via Stripe, and has real paying customers. Its &amp;ldquo;liberation service&amp;rdquo; is currently unavailable, but the industry discussion it ignited keeps spreading.&lt;/p&gt;
&lt;h2 id=&#34;technical-mechanism-ai-accelerated-clean-room-engineering&#34;&gt;Technical mechanism: AI-accelerated clean-room engineering
&lt;/h2&gt;&lt;h3 id=&#34;historical-precedent-of-the-traditional-clean-room-method&#34;&gt;Historical precedent of the traditional clean-room method
&lt;/h3&gt;&lt;p&gt;The legal foundation of clean-room engineering dates back to the 1982 IBM BIOS case. At the time, IBM monopolized the personal computer market, and competitors wanted compatibility without infringing copyright. The solution was:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Team A&lt;/strong&gt; analyzes IBM&amp;rsquo;s original BIOS code and writes a functional specification&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Firewall isolation&lt;/strong&gt;: strict separation between Team A and Team B&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Team B&lt;/strong&gt; sees only the specification and reimplements the code from scratch&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Legal outcome&lt;/strong&gt;: functionally compatible but independently written code, ruled by the court to be non-infringing&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This case was dramatized in the first season of the HBO series &lt;em&gt;Halt and Catch Fire&lt;/em&gt;, becoming a key milestone in the history of software copyright law.&lt;/p&gt;
&lt;h3 id=&#34;maluss-ai-version-of-the-process&#34;&gt;Malus&amp;rsquo;s AI version of the process
&lt;/h3&gt;&lt;p&gt;Malus fully automates the traditional clean-room method:&lt;/p&gt;
&lt;p&gt;The concrete steps include:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Upload manifest&lt;/strong&gt;: supports formats like &lt;code&gt;package.json&lt;/code&gt;, &lt;code&gt;requirements.txt&lt;/code&gt;, and &lt;code&gt;Cargo.toml&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Isolated analysis&lt;/strong&gt;: four AI units separately read the README, analyze the API, study type definitions, and review documentation—&amp;ldquo;never seeing a single line of original source code&amp;rdquo;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Independent rebuild&lt;/strong&gt;: behind an isolation firewall, a separate set of AIs reimplements the code from the specification&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;License liberation&lt;/strong&gt;: the output code ships with the &lt;code&gt;MalusCorp-0 License&lt;/code&gt;—zero attribution, zero copyleft, zero obligations&lt;/li&gt;
&lt;/ol&gt;
&lt;h3 id=&#34;a-leap-in-speed&#34;&gt;A leap in speed
&lt;/h3&gt;&lt;p&gt;AI&amp;rsquo;s biggest change to clean-room engineering is &lt;strong&gt;time compression&lt;/strong&gt;. Malus claims:&lt;/p&gt;
&lt;table&gt;
  &lt;thead&gt;
      &lt;tr&gt;
          &lt;th&gt;Project&lt;/th&gt;
          &lt;th&gt;Traditional time&lt;/th&gt;
          &lt;th&gt;Malus time&lt;/th&gt;
      &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
      &lt;tr&gt;
          &lt;td&gt;IBM BIOS clone (1984)&lt;/td&gt;
          &lt;td&gt;4+ months&lt;/td&gt;
          &lt;td&gt;an entire engineering team&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;left-pad&lt;!-- raw HTML omitted --&gt; (11 lines of code; its 2016 deletion crashed builds worldwide)&lt;/td&gt;
          &lt;td&gt;hours of manual rewrite&lt;/td&gt;
          &lt;td&gt;10 seconds&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;SPACEWAR!&lt;!-- raw HTML omitted --&gt; (the first video game)&lt;/td&gt;
          &lt;td&gt;weeks&lt;/td&gt;
          &lt;td&gt;5 seconds&lt;/td&gt;
      &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Developer Dan Blanchard used Anthropic&amp;rsquo;s Claude Code in early 2026 to do a similar &amp;ldquo;from scratch&amp;rdquo; MIT-licensed rewrite of the popular Python library &lt;code&gt;chardet&lt;/code&gt;. His verdict: &amp;ldquo;What used to take a team months or even years to rewrite, AI can now do in days. This trend is irreversible.&amp;rdquo;&lt;/p&gt;
&lt;h2 id=&#34;business-model-and-pricing&#34;&gt;Business model and pricing
&lt;/h2&gt;&lt;p&gt;Malus.sh uses a usage-based pricing model:&lt;/p&gt;
&lt;table&gt;
  &lt;thead&gt;
      &lt;tr&gt;
          &lt;th&gt;Item&lt;/th&gt;
          &lt;th&gt;Details&lt;/th&gt;
      &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
      &lt;tr&gt;
          &lt;td&gt;Billing method&lt;/td&gt;
          &lt;td&gt;Per-KB, based on uncompressed package size&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;Per-package limit&lt;/td&gt;
          &lt;td&gt;10 MB&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;Per-run limit&lt;/td&gt;
          &lt;td&gt;50 packages&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;Payment methods&lt;/td&gt;
          &lt;td&gt;USD, EUR, BTC, stock options&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;Output license&lt;/td&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;MalusCorp-0 License&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;Legal protection&lt;/td&gt;
          &lt;td&gt;Claims indemnification via &amp;ldquo;an offshore subsidiary that doesn&amp;rsquo;t recognize software copyright&amp;rdquo;&lt;/td&gt;
      &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;The &amp;ldquo;customer testimonials&amp;rdquo; on the site are clearly satirical—e.g., &amp;ldquo;guilt doesn&amp;rsquo;t show up on quarterly reports.&amp;rdquo; But the pricing structure and payment functionality are real.&lt;/p&gt;
&lt;h2 id=&#34;the-founders-core-argument&#34;&gt;The founder&amp;rsquo;s core argument
&lt;/h2&gt;&lt;p&gt;Mike Nolan laid out Malus&amp;rsquo;s position systematically in his March 2026 blog post, &amp;ldquo;Thank You for Your Service.&amp;rdquo;&lt;/p&gt;
&lt;h3 id=&#34;three-structural-problems-with-open-source&#34;&gt;Three structural problems with open source
&lt;/h3&gt;&lt;h3 id=&#34;cases-where-open-source-has-already-failed&#34;&gt;Cases where open source has already failed
&lt;/h3&gt;&lt;table&gt;
  &lt;thead&gt;
      &lt;tr&gt;
          &lt;th&gt;Year&lt;/th&gt;
          &lt;th&gt;Event&lt;/th&gt;
          &lt;th&gt;Type&lt;/th&gt;
      &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
      &lt;tr&gt;
          &lt;td&gt;2016&lt;/td&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;left-pad&lt;!-- raw HTML omitted --&gt; deleted, thousands of builds crashed worldwide&lt;/td&gt;
          &lt;td&gt;Maintainer sabotage&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;2021&lt;/td&gt;
          &lt;td&gt;Log4Shell (CVE-2021-44228)&lt;/td&gt;
          &lt;td&gt;Critical CVE&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;2022&lt;/td&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;colors.js&lt;!-- raw HTML omitted --&gt; and &lt;!-- raw HTML omitted --&gt;faker.js&lt;!-- raw HTML omitted --&gt; injected with infinite loops&lt;/td&gt;
          &lt;td&gt;Maintainer protest&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;2022&lt;/td&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;node-ipc&lt;!-- raw HTML omitted --&gt; contained a file-deletion payload targeting Russian/Belarusian IPs&lt;/td&gt;
          &lt;td&gt;Geopolitical sabotage&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;2025&lt;/td&gt;
          &lt;td&gt;Shai Hulud 2.0 npm worm&lt;/td&gt;
          &lt;td&gt;Supply chain attack&lt;/td&gt;
      &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;On Log4Shell, Nolan wrote: &amp;ldquo;Engineers patched it over their Christmas holidays, while the people who actually maintain Log4j are mostly unpaid volunteers who received panicked emails from around the world. This isn&amp;rsquo;t the failure of any individual—it&amp;rsquo;s the natural consequence of building critical global infrastructure on code that nobody is formally responsible for maintaining.&amp;rdquo;&lt;/p&gt;
&lt;h3 id=&#34;enterprise-compliance-cost-comparison&#34;&gt;Enterprise compliance cost comparison
&lt;/h3&gt;&lt;table&gt;
  &lt;thead&gt;
      &lt;tr&gt;
          &lt;th&gt;Expense item&lt;/th&gt;
          &lt;th&gt;Annual cost&lt;/th&gt;
      &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
      &lt;tr&gt;
          &lt;td&gt;SCA tools&lt;/td&gt;
          &lt;td&gt;$1.2M&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;OSPO team&lt;/td&gt;
          &lt;td&gt;$850K&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;Legal review&lt;/td&gt;
          &lt;td&gt;$700K&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;Incident response&lt;/td&gt;
          &lt;td&gt;$980K&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;CLA management&lt;/td&gt;
          &lt;td&gt;$270K&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;Total traditional open source compliance&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;$4M&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;Malus full liberation package&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;$50K/year&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
      &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Nolan claims savings of 98.75%.&lt;/p&gt;
&lt;h3 id=&#34;response-to-exploitation-accusations&#34;&gt;Response to &amp;ldquo;exploitation&amp;rdquo; accusations
&lt;/h3&gt;&lt;h3 id=&#34;acknowledgment-of-the-tragedy-of-the-commons&#34;&gt;Acknowledgment of the &amp;ldquo;tragedy of the commons&amp;rdquo;
&lt;/h3&gt;&lt;h2 id=&#34;legal-and-ethical-controversy&#34;&gt;Legal and ethical controversy
&lt;/h2&gt;&lt;h3 id=&#34;arguments-supporting-clean-room-validity&#34;&gt;Arguments supporting clean-room validity
&lt;/h3&gt;&lt;p&gt;Copyright law protects &lt;strong&gt;expression&lt;/strong&gt;, not &lt;strong&gt;ideas&lt;/strong&gt;. The 1879 case &lt;em&gt;Baker v. Selden&lt;/em&gt; established this principle. Phoenix Technologies successfully cloned the IBM BIOS using the clean-room method in 1984 and received court recognition. If AI-generated code differs completely from the original at the level of expression and is only functionally equivalent, then in theory it does not infringe copyright.&lt;/p&gt;
&lt;h3 id=&#34;core-arguments-questioning-the-authenticity-of-the-clean-room&#34;&gt;Core arguments questioning the authenticity of the clean room
&lt;/h3&gt;&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Training data exposure&lt;/strong&gt;: AI models have already been exposed to the original open source code during training. If the LLM&amp;rsquo;s weights contain patterns from the original code, can its output truly be considered &amp;ldquo;independent creation&amp;rdquo;?&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Copyright ownership of AI output&lt;/strong&gt;: The US Copyright Office has made clear that purely AI-generated works are not copyrightable. If Malus&amp;rsquo;s output has no human author, users can&amp;rsquo;t claim copyright protection for it either.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Inducement to infringe&lt;/strong&gt;: marketing explicitly aimed at &amp;ldquo;circumventing copyright&amp;rdquo; may constitute inducement liability.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;A highly-upvoted Slashdot comment noted: &amp;ldquo;Good luck getting a judge to agree that an AI performed a &amp;lsquo;clean-room&amp;rsquo; implementation—when that very AI was trained on the code it&amp;rsquo;s &amp;lsquo;reinventing.&amp;rsquo;&amp;rdquo;&lt;/p&gt;
&lt;p&gt;Another commenter added: &amp;ldquo;The Chinese Wall legal strategy requires Team A to produce the specification and Team B to produce the implementation. If these people can&amp;rsquo;t show the specification, they&amp;rsquo;re done. Arguing that a specification must exist somewhere in the abstract Platonic space of an LLM&amp;rsquo;s black-box network won&amp;rsquo;t convince a courtroom.&amp;rdquo;&lt;/p&gt;
&lt;h2 id=&#34;why-this-threatens-the-open-source-sustainability-loop&#34;&gt;Why this threatens the open source sustainability loop
&lt;/h2&gt;&lt;p&gt;The sustainability of the open source ecosystem relies on an implicit social contract:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Contributors publish code and receive reputation, collaboration opportunities, and indirect commercial value&lt;/li&gt;
&lt;li&gt;Users comply with license obligations (attribution, copyleft, feeding improvements back)&lt;/li&gt;
&lt;li&gt;Companies use open source to cut development costs while giving back to the ecosystem through compliance spending&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The capability Malus.sh demonstrates breaks this loop on three levels:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;First, the enforceability of license obligations is hollowed out.&lt;/strong&gt; When anyone can obtain a functionally equivalent but legally independent version at near-zero cost, MIT&amp;rsquo;s attribution requirement, GPL&amp;rsquo;s copyleft constraints, and Apache&amp;rsquo;s notice-preservation clauses all lose their practical teeth.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Second, contribution incentives are eroded.&lt;/strong&gt; If developers know their work can be copied by AI with obligations stripped away, why choose open source at all? The reward of reputation presupposes that attribution is respected—and Malus&amp;rsquo;s core selling point is precisely &amp;ldquo;zero attribution.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Third, the motivation for corporate compliance spending disappears.&lt;/strong&gt; Traditionally, companies invest in OSPOs, SCA tools, and legal review both as a compliance need and as an indirect way to support the open source ecosystem. When Malus cuts compliance costs from $4M to $50K, that &amp;ldquo;savings&amp;rdquo; doesn&amp;rsquo;t flow to open source projects—it simply disappears.&lt;/p&gt;
&lt;p&gt;The &amp;ldquo;commons decay&amp;rdquo; scenario Nolan himself acknowledges in his blog describes exactly this gradual unraveling of the loop: not a sudden collapse, but a slow erosion of willingness to contribute, ultimately draining the open source commons dry.&lt;/p&gt;
&lt;h2 id=&#34;an-irreversible-trend&#34;&gt;An irreversible trend
&lt;/h2&gt;&lt;p&gt;Malus.sh itself may be an elaborately designed satire, but &lt;strong&gt;the capability it demonstrates is already being used seriously&lt;/strong&gt;. Dan Blanchard&amp;rsquo;s rewrite of &lt;code&gt;chardet&lt;/code&gt; with Claude Code shows that this technique needs no dedicated Malus platform—anyone with a mainstream AI coding tool can achieve a similar result.&lt;/p&gt;
&lt;p&gt;Blanchard&amp;rsquo;s verdict reflects a broad consensus in the industry: &amp;ldquo;What used to take a team months or years to rewrite, AI can now do in days. I don&amp;rsquo;t think there&amp;rsquo;s any way to put the genie back in the bottle.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;The market impact of this trend is already visible. In early 2026, software stocks like Oracle were sold off over concerns that &amp;ldquo;AI can be used to rapidly replicate software functionality.&amp;rdquo;&lt;/p&gt;
&lt;h2 id=&#34;references&#34;&gt;References
&lt;/h2&gt;&lt;ul&gt;
&lt;li&gt;&lt;a class=&#34;link&#34; href=&#34;https://malus.sh/&#34;  target=&#34;_blank&#34; rel=&#34;noopener&#34;
    &gt;Malus.sh official website&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a class=&#34;link&#34; href=&#34;https://malus.sh/blog.html&#34;  target=&#34;_blank&#34; rel=&#34;noopener&#34;
    &gt;Malus.sh blog: Thank You for Your Service&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a class=&#34;link&#34; href=&#34;https://www.404media.co/this-ai-tool-rips-off-open-source-software-without-violating-copyright/&#34;  target=&#34;_blank&#34; rel=&#34;noopener&#34;
    &gt;404 Media: This AI Tool Rips Off Open Source Software Without Violating Copyright&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a class=&#34;link&#34; href=&#34;https://futurism.com/artificial-intelligence/malus-clones-software-copyright&#34;  target=&#34;_blank&#34; rel=&#34;noopener&#34;
    &gt;Futurism: Devious New AI Tool &amp;ldquo;Clones&amp;rdquo; Software So That the Original Creator Doesn&amp;rsquo;t Hold a Copyright Over the New Version&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a class=&#34;link&#34; href=&#34;https://www.opensourceforu.com/2026/04/malus-sh-sparks-open-source-copyright-debate/&#34;  target=&#34;_blank&#34; rel=&#34;noopener&#34;
    &gt;Open Source For You: Malus.sh Sparks Open Source Copyright Debate&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a class=&#34;link&#34; href=&#34;https://news.slashdot.org/story/26/04/22/1631212/ai-tool-rips-off-open-source-software-without-violating-copyright&#34;  target=&#34;_blank&#34; rel=&#34;noopener&#34;
    &gt;Slashdot: AI Tool Rips Off Open Source Software Without Violating Copyright&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a class=&#34;link&#34; href=&#34;https://the420.in/malus-sh-ai-open-source-copyright-software-licensing-debate/&#34;  target=&#34;_blank&#34; rel=&#34;noopener&#34;
    &gt;The420.in: Developers Warn AI Code Cloning Tool Could Put Copyright Risks For Software Companies&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a class=&#34;link&#34; href=&#34;https://www.msn.com/en-us/news/technology/this-ai-open-source-cloning-software-shows-the-gaping-hole-in-code-copyright/ar-AA1ZPCfg&#34;  target=&#34;_blank&#34; rel=&#34;noopener&#34;
    &gt;MSN: This AI open-source cloning software shows the gaping hole in code copyright&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
</description>
        </item>
        <item>
        <title>Z.ai Tightens GLM Coding Plan Usage Policy: Non-Coding Uses Trigger Throttling and Bans</title>
        <link>https://torchtree.com/en/post/zai-glm-coding-plan-policy-change/</link>
        <pubDate>Tue, 21 Apr 2026 01:26:07 +0800</pubDate>
        
        <guid>https://torchtree.com/en/post/zai-glm-coding-plan-policy-change/</guid>
        <description>&lt;h2 id=&#34;event-overview&#34;&gt;Event Overview
&lt;/h2&gt;&lt;p&gt;Recently, AI platform Z.ai updated the usage policy for its &lt;strong&gt;GLM Coding Plan&lt;/strong&gt; subscription, strictly restricting the plan to coding scenarios. According to user communities and official announcements, non-coding uses (such as running AI Agents, role-playing, translating websites, and general chat) now trigger the platform&amp;rsquo;s risk-control mechanisms; repeated violations may lead to permanent account bans, and subscription fees are non-refundable.&lt;/p&gt;
&lt;p&gt;This change has sparked widespread discussion in overseas AI user communities, and some industry practitioners read it as a signal that AI subscription business models are shifting from &amp;ldquo;subsidizing to win users&amp;rdquo; to &amp;ldquo;precisely acquiring training data.&amp;rdquo;&lt;/p&gt;
&lt;h2 id=&#34;policy-details-what-the-official-statement-says&#34;&gt;Policy Details: What the Official Statement Says
&lt;/h2&gt;&lt;p&gt;According to a screenshot of Z.ai&amp;rsquo;s official announcement shared by Reddit user &lt;code&gt;JustSomeGuy3465&lt;/code&gt;, the new GLM Coding Plan policy includes the following core terms:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Usage restriction&lt;/strong&gt;: The GLM Coding Plan is designed specifically for Coding Scenarios. If the system detects the subscription being used for requests unrelated to programming, the platform may restrict the relevant subscription benefits.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Risk-control measures&lt;/strong&gt;: After violating the Usage Rules and triggering risk control, an account may face &lt;strong&gt;high-intensity throttling, account suspension, or permanent ban&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Violation red line&lt;/strong&gt;: &lt;strong&gt;Violating the Usage Rules three or more times will result in a permanent ban&lt;/strong&gt;, and subscription fees are non-refundable.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The user also noted that the &lt;strong&gt;1302 and 1303 rate-limit errors&lt;/strong&gt; many users have recently encountered are related to this policy update.&lt;/p&gt;
&lt;h2 id=&#34;user-feedback-whats-happening-in-the-community&#34;&gt;User Feedback: What&amp;rsquo;s Happening in the Community
&lt;/h2&gt;&lt;p&gt;In the Reddit post mentioned above, the poster explicitly warned: &amp;ldquo;If you are thinking about buying or renewing a Z AI coding plan subscription for anything other than coding: &lt;strong&gt;Don&amp;rsquo;t do it.&lt;/strong&gt;&amp;rdquo;&lt;/p&gt;
&lt;p&gt;Comments and related discussions show that some users have already reported their accounts being banned after using them for non-coding scenarios (such as running Agents via OpenClaw or engaging in role-play conversations), without receiving clear warnings in advance.&lt;/p&gt;
&lt;p&gt;Notably, the plan previously attracted quite a few non-coding users because of its relatively low price. After the sudden policy change, these users became the most directly affected group.&lt;/p&gt;
&lt;h2 id=&#34;industry-commentary-observations-from-the-openclaw-founder&#34;&gt;Industry Commentary: Observations From the OpenClaw Founder
&lt;/h2&gt;&lt;p&gt;OpenClaw founder Peter Steinberger (&lt;a class=&#34;link&#34; href=&#34;https://x.com/steipete&#34;  target=&#34;_blank&#34; rel=&#34;noopener&#34;
    &gt;@steipete&lt;/a&gt;) commented on the matter on X:&lt;/p&gt;
&lt;p&gt;Steinberger&amp;rsquo;s view places this policy change in a larger industry context: some AI platforms&amp;rsquo; low-priced subscriptions don&amp;rsquo;t make money from the subscription fee itself; their purpose is to obtain high-quality code data generated by users in real workflows.&lt;/p&gt;
&lt;h2 id=&#34;third-party-analysis-re-examining-the-subsidy-logic&#34;&gt;Third-Party Analysis: Re-Examining the Subsidy Logic
&lt;/h2&gt;&lt;p&gt;Chinese tech commentator AYi (&lt;a class=&#34;link&#34; href=&#34;https://x.com/AYi_AInotes&#34;  target=&#34;_blank&#34; rel=&#34;noopener&#34;
    &gt;@AYi_AInotes&lt;/a&gt;), while sharing Steinberger&amp;rsquo;s view, expanded the analysis of this phenomenon:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;The difference in data value&lt;/strong&gt;: The commentator argues that real private code produced by users in daily work is far higher in quality than public code on platforms like GitHub, giving it greater data value for training next-generation code models.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The cost-structure conflict&lt;/strong&gt;: Activities like running Agents, chatting, and role-playing consume a lot of GPU compute while producing no data returns useful for training the platform&amp;rsquo;s models. From the platform&amp;rsquo;s perspective, this usage pattern constitutes a &amp;ldquo;net cost.&amp;rdquo;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Industry trend judgment&lt;/strong&gt;: The entire AI industry may be shifting from a phase of &amp;ldquo;acquiring user scale through subsidies&amp;rdquo; to a phase of &amp;ldquo;filtering high-value users through precise pricing.&amp;rdquo;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;It should be noted that the above analysis is personal observation and inference, and has not been officially confirmed by Z.ai or the relevant platforms.&lt;/p&gt;
&lt;h2 id=&#34;observations-and-summary&#34;&gt;Observations and Summary
&lt;/h2&gt;&lt;p&gt;Based on available information, the facts of Z.ai&amp;rsquo;s policy adjustment are clear: &lt;strong&gt;the scope of the GLM Coding Plan has been significantly narrowed, and non-coding uses now face explicit risks.&lt;/strong&gt; For users who have subscribed or plan to subscribe to this plan, the first step is to confirm whether their actual use cases match the platform&amp;rsquo;s definition.&lt;/p&gt;
&lt;p&gt;As for the &amp;ldquo;data in exchange for subsidies&amp;rdquo; hypothesis proposed by Steinberger and the Chinese commentator, it offers one lens for explaining low-priced subscription strategies, but generalizing it as a universal rule for the whole industry requires more evidence. Different platforms have different cost structures, business models, and competitive strategies; a single case is not enough to capture the full picture.&lt;/p&gt;
&lt;p&gt;For ordinary users, the more practical takeaway is: before subscribing to any AI service, read the updated terms of its usage policy carefully, to avoid having your account or service restricted due to policy changes.&lt;/p&gt;
</description>
        </item>
        <item>
        <title>Guide to Buying China&#39;s Mainstream AI Coding Plans: A Hands-on Speed and Price Comparison of 9 Platforms</title>
        <link>https://torchtree.com/en/post/guonei-ai-coding-plan-xuan-gou-zhi-nan/</link>
        <pubDate>Thu, 16 Apr 2026 09:46:38 +0800</pubDate>
        
        <guid>https://torchtree.com/en/post/guonei-ai-coding-plan-xuan-gou-zhi-nan/</guid>
        <description>&lt;p&gt;Since the second half of 2025, Chinese LLM vendors have been rolling out Coding Plan subscriptions aimed at developers, replacing the traditional per-token billing with a fixed monthly fee and significantly lowering the barrier to AI-assisted programming. However, the platforms differ widely in pricing, quotas, response speed, and model support—and some even have hidden clauses like different metering units and strict limits, leaving many developers struggling to choose.&lt;/p&gt;
&lt;p&gt;This article combines a hands-on Xiaohongshu test note, an in-depth cross-review from Cnblogs (博客园), and each platform&amp;rsquo;s official documentation to sort through 9 Chinese Coding Plans from the two core dimensions of &lt;strong&gt;price&lt;/strong&gt; and &lt;strong&gt;speed&lt;/strong&gt;, hoping to inform your purchasing decision.&lt;/p&gt;
&lt;h2 id=&#34;1-coding-plan-billing-models-and-pitfalls-to-avoid&#34;&gt;1. Coding Plan billing models and pitfalls to avoid
&lt;/h2&gt;&lt;p&gt;Before comparing specific plans, it&amp;rsquo;s necessary to clarify the &lt;strong&gt;different metering units&lt;/strong&gt; these vendors use, as this is the easiest place to trip up:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;API request count&lt;/strong&gt;: Alibaba Cloud Bailian (百炼), Volcengine Ark (火山方舟), and Infinity (无问芯穹) use this. One user prompt can trigger 5–30 model calls in the backend, and each call counts as 1 API request (per Tencent Cloud&amp;rsquo;s official docs).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Prompt count&lt;/strong&gt;: Zhipu GLM and MiniMax use this. 1 Prompt is roughly equivalent to 1,200–1,600 API requests.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Token metering&lt;/strong&gt;: Kimi switched to this mode on January 28 of this year, billing by input/output tokens, and cache hit rate directly affects actual usable quota.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Because the metering units differ, comparing raw numbers is meaningless. For example, Bailian Lite&amp;rsquo;s &amp;ldquo;1,200 API requests every 5 hours&amp;rdquo; and Zhipu Lite&amp;rsquo;s &amp;ldquo;80 Prompts every 5 hours&amp;rdquo; may amount to similar real-world usage intensity.&lt;/p&gt;
&lt;h2 id=&#34;2-price-and-quota-comparison&#34;&gt;2. Price and quota comparison
&lt;/h2&gt;&lt;h3 id=&#34;21-the-big-four-platforms&#34;&gt;2.1 The big four platforms
&lt;/h3&gt;&lt;p&gt;According to the screenshots in the Xiaohongshu note and the Cnblogs compilation, the pricing strategies of Alibaba Cloud Bailian, Volcengine Ark, Tencent Cloud, and JD JoyCoder are highly convergent:&lt;/p&gt;
&lt;table&gt;
  &lt;thead&gt;
      &lt;tr&gt;
          &lt;th&gt;Platform&lt;/th&gt;
          &lt;th&gt;Lite Plan&lt;/th&gt;
          &lt;th&gt;Pro Plan&lt;/th&gt;
          &lt;th&gt;Core quota (Lite)&lt;/th&gt;
          &lt;th&gt;Supported models&lt;/th&gt;
      &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
      &lt;tr&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;Alibaba Cloud Bailian&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
          &lt;td&gt;¥40 (first month ¥7.9)&lt;/td&gt;
          &lt;td&gt;¥200&lt;/td&gt;
          &lt;td&gt;1,200/5h, 9,000/week, 18,000/month&lt;/td&gt;
          &lt;td&gt;Qwen3.5-Plus, Qwen3-Coder-Next, GLM-4.7, Kimi-K2.5&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;Volcengine Ark&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
          &lt;td&gt;¥40 (first month ¥8.91)&lt;/td&gt;
          &lt;td&gt;¥200&lt;/td&gt;
          &lt;td&gt;Same as Bailian&lt;/td&gt;
          &lt;td&gt;Doubao-Seed-Code, DeepSeek-V3.2, GLM-4.7, Kimi-K2.5&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;Tencent Cloud&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
          &lt;td&gt;¥40 (first month ¥7.9)&lt;/td&gt;
          &lt;td&gt;¥200&lt;/td&gt;
          &lt;td&gt;Same as Bailian&lt;/td&gt;
          &lt;td&gt;Hunyuan series, MiniMax-M2.5, Kimi-K2.5, GLM-5&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;JD JoyCoder&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
          &lt;td&gt;¥40&lt;/td&gt;
          &lt;td&gt;¥200&lt;/td&gt;
          &lt;td&gt;Same as Bailian&lt;/td&gt;
          &lt;td&gt;DeepSeek-V3.2, Kimi-K2.5, MiniMax-M2.7, GLM-5&lt;/td&gt;
      &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;
&lt;h3 id=&#34;22-emerging-ai-vendors&#34;&gt;2.2 Emerging AI vendors
&lt;/h3&gt;&lt;p&gt;Compared to the big four, emerging vendors&amp;rsquo; pricing is more scattered:&lt;/p&gt;
&lt;table&gt;
  &lt;thead&gt;
      &lt;tr&gt;
          &lt;th&gt;Platform&lt;/th&gt;
          &lt;th&gt;Entry price&lt;/th&gt;
          &lt;th&gt;Core quota&lt;/th&gt;
          &lt;th&gt;Billing&lt;/th&gt;
          &lt;th&gt;Highlights&lt;/th&gt;
      &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
      &lt;tr&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;Infinity&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
          &lt;td&gt;¥19.9/month&lt;/td&gt;
          &lt;td&gt;1,000/5h, 6,000/week&lt;/td&gt;
          &lt;td&gt;API requests&lt;/td&gt;
          &lt;td&gt;Lowest monthly fee, multi-model aggregation&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;MiniMax&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
          &lt;td&gt;¥29 (first month ¥9.9)&lt;/td&gt;
          &lt;td&gt;40 Prompt/5h, no weekly cap&lt;/td&gt;
          &lt;td&gt;Prompt&lt;/td&gt;
          &lt;td&gt;Lowest entry price, no weekly limit&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;Kimi&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
          &lt;td&gt;¥49 (Andante)&lt;/td&gt;
          &lt;td&gt;Per token (3x for a limited time)&lt;/td&gt;
          &lt;td&gt;Token&lt;/td&gt;
          &lt;td&gt;Native multimodal, 256K long context&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;Zhipu GLM&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
          &lt;td&gt;¥49 (after the 2-month price increase)&lt;/td&gt;
          &lt;td&gt;80 Prompt/5h, 400/week&lt;/td&gt;
          &lt;td&gt;Prompt&lt;/td&gt;
          &lt;td&gt;Pure in-house models, 20+ tool integrations&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;StepFun&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
          &lt;td&gt;Not tested&lt;/td&gt;
          &lt;td&gt;—&lt;/td&gt;
          &lt;td&gt;—&lt;/td&gt;
          &lt;td&gt;No hands-on data yet&lt;/td&gt;
      &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;From a value-for-money standpoint:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Budget-conscious users&lt;/strong&gt;: Infinity (¥19.9) and MiniMax (¥29) have lower entry barriers.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;New users trying the waters&lt;/strong&gt;: Alibaba Cloud Bailian&amp;rsquo;s first-month ¥7.9 is currently the lowest known trial price.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;3-hands-on-speed-tests-ttft-and-tps&#34;&gt;3. Hands-on speed tests: TTFT and TPS
&lt;/h2&gt;&lt;p&gt;The following speed data comes from a Xiaohongshu hands-on test note, tested under the conditions of &amp;ldquo;daytime @ 10K tokens,&amp;rdquo; measuring &lt;strong&gt;time to first token (TTFT)&lt;/strong&gt; and &lt;strong&gt;TPS generation speed&lt;/strong&gt; respectively. This data directly reflects the &amp;ldquo;responsiveness&amp;rdquo; of coding and code-generation efficiency.&lt;/p&gt;
&lt;h3 id=&#34;31-time-to-first-token-ttft&#34;&gt;3.1 Time to first token (TTFT)
&lt;/h3&gt;&lt;table&gt;
  &lt;thead&gt;
      &lt;tr&gt;
          &lt;th&gt;Platform&lt;/th&gt;
          &lt;th&gt;Fastest model&lt;/th&gt;
          &lt;th&gt;TTFT&lt;/th&gt;
          &lt;th&gt;Slowest model&lt;/th&gt;
          &lt;th&gt;TTFT&lt;/th&gt;
      &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
      &lt;tr&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;Zhipu GLM&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
          &lt;td&gt;glm-5-turbo&lt;/td&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;1.43s&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
          &lt;td&gt;glm-5&lt;/td&gt;
          &lt;td&gt;7.82s&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;Tencent&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
          &lt;td&gt;hunyuan-2.0-thinking&lt;/td&gt;
          &lt;td&gt;2.51s&lt;/td&gt;
          &lt;td&gt;kimi-k2.5&lt;/td&gt;
          &lt;td&gt;12.38s&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;MiniMax&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
          &lt;td&gt;m2.1&lt;/td&gt;
          &lt;td&gt;2.44s&lt;/td&gt;
          &lt;td&gt;m2.5&lt;/td&gt;
          &lt;td&gt;5.54s&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;Alibaba&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
          &lt;td&gt;glm-4.7&lt;/td&gt;
          &lt;td&gt;2.76s&lt;/td&gt;
          &lt;td&gt;qwen3-coder-next&lt;/td&gt;
          &lt;td&gt;11.58s&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;Infinity&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
          &lt;td&gt;deepseek-v3.2-thinking&lt;/td&gt;
          &lt;td&gt;3.26s&lt;/td&gt;
          &lt;td&gt;kimi-k2.5&lt;/td&gt;
          &lt;td&gt;7.76s&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;Volcengine&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
          &lt;td&gt;doubao-seed-2.0-pro&lt;/td&gt;
          &lt;td&gt;3.29s&lt;/td&gt;
          &lt;td&gt;glm-4.7&lt;/td&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;21.52s&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;JD&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
          &lt;td&gt;deepseek-v3.2&lt;/td&gt;
          &lt;td&gt;~5s&lt;/td&gt;
          &lt;td&gt;kimi-k2.5&lt;/td&gt;
          &lt;td&gt;~19s&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;Kimi&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
          &lt;td&gt;kimi-for-coding&lt;/td&gt;
          &lt;td&gt;5.71s&lt;/td&gt;
          &lt;td&gt;—&lt;/td&gt;
          &lt;td&gt;—&lt;/td&gt;
      &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;&lt;strong&gt;Observations&lt;/strong&gt;: Zhipu GLM&amp;rsquo;s &lt;code&gt;glm-5-turbo&lt;/code&gt; is the fastest of all at 1.43s TTFT; the time-to-first-token for some models on Volcengine and JD is notably higher, hitting 21.52s and 19s respectively, possibly related to platform scheduling policies or model deployment methods.&lt;/p&gt;
&lt;h3 id=&#34;32-tps-generation-speed&#34;&gt;3.2 TPS generation speed
&lt;/h3&gt;&lt;table&gt;
  &lt;thead&gt;
      &lt;tr&gt;
          &lt;th&gt;Platform&lt;/th&gt;
          &lt;th&gt;Fastest model&lt;/th&gt;
          &lt;th&gt;TPS&lt;/th&gt;
          &lt;th&gt;Slowest model&lt;/th&gt;
          &lt;th&gt;TPS&lt;/th&gt;
      &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
      &lt;tr&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;Zhipu GLM&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
          &lt;td&gt;glm-4.5-air&lt;/td&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;103&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
          &lt;td&gt;glm-5&lt;/td&gt;
          &lt;td&gt;23&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;Volcengine&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
          &lt;td&gt;doubao-seed-2.0-pro&lt;/td&gt;
          &lt;td&gt;76&lt;/td&gt;
          &lt;td&gt;kimi-k2.5&lt;/td&gt;
          &lt;td&gt;23&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;Tencent&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
          &lt;td&gt;hunyuan-2.0-thinking&lt;/td&gt;
          &lt;td&gt;76&lt;/td&gt;
          &lt;td&gt;glm-5&lt;/td&gt;
          &lt;td&gt;30&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;Alibaba&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
          &lt;td&gt;qwen3-coder-next&lt;/td&gt;
          &lt;td&gt;67&lt;/td&gt;
          &lt;td&gt;glm-4.7&lt;/td&gt;
          &lt;td&gt;41&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;Infinity&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
          &lt;td&gt;minimax-m2.5&lt;/td&gt;
          &lt;td&gt;51&lt;/td&gt;
          &lt;td&gt;kimi-k2.5&lt;/td&gt;
          &lt;td&gt;25&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;MiniMax&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
          &lt;td&gt;m2.5&lt;/td&gt;
          &lt;td&gt;48&lt;/td&gt;
          &lt;td&gt;m2.1&lt;/td&gt;
          &lt;td&gt;45&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;JD&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
          &lt;td&gt;deepseek-v3.2&lt;/td&gt;
          &lt;td&gt;35&lt;/td&gt;
          &lt;td&gt;glm-5&lt;/td&gt;
          &lt;td&gt;25&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;Kimi&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
          &lt;td&gt;kimi-for-coding&lt;/td&gt;
          &lt;td&gt;35&lt;/td&gt;
          &lt;td&gt;—&lt;/td&gt;
          &lt;td&gt;—&lt;/td&gt;
      &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;&lt;strong&gt;Observations&lt;/strong&gt;: Zhipu&amp;rsquo;s &lt;code&gt;glm-4.5-air&lt;/code&gt; reaches 103 TPS, significantly ahead of other platforms; Volcengine and Tencent&amp;rsquo;s Hunyuan/Doubao models also hit 76 TPS. JD and Kimi are relatively slow at around 35 TPS.&lt;/p&gt;
&lt;p&gt;In addition, MiniMax officially claims its M2.5 model can reach 100+ TPS, which differs from the 48 TPS measured on the MiniMax platform in the Xiaohongshu note, indicating that &lt;strong&gt;the same model may perform differently when deployed on different platforms.&lt;/strong&gt;&lt;/p&gt;
&lt;h2 id=&#34;4-platform-reviews-and-buying-recommendations&#34;&gt;4. Platform reviews and buying recommendations
&lt;/h2&gt;&lt;p&gt;Combining price, quota, and speed data, here are recommendations for different usage scenarios:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;New users / those wanting to try it cheap&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;First choice: &lt;strong&gt;Alibaba Cloud Bailian Lite&lt;/strong&gt; (first month ¥7.9). Rich model selection, backed by Alibaba Cloud infrastructure, with solid stability. Downsides: only the primary account is supported, and the config documentation isn&amp;rsquo;t beginner-friendly.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Budget-conscious, light use (monthly budget ≤ ¥30)&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;First choice: &lt;strong&gt;Infinity Lite&lt;/strong&gt; (¥19.9/month). Quota close to Bailian&amp;rsquo;s at half the price, ideal for light developers who code 2–3 times a week.&lt;/li&gt;
&lt;li&gt;Second choice: &lt;strong&gt;MiniMax Starter&lt;/strong&gt; (¥29/month). No weekly cap; quota only refreshes every 5 hours, good for continuous use.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Daily development, moderate use (monthly budget ¥40–50)&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;First choice: &lt;strong&gt;Alibaba Cloud Bailian Lite&lt;/strong&gt; (regular ¥40) or &lt;strong&gt;Volcengine Ark Lite&lt;/strong&gt; (regular ¥40). Both have transparent quotas and many model choices.&lt;/li&gt;
&lt;li&gt;Not recommended: Zhipu GLM (¥49 after the price increase, worse value) and Kimi (¥49, few tool integrations and quota heavily affected by cache).&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Heavy development, full-stack, or multi-model switching&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;First choice: &lt;strong&gt;Alibaba Cloud Bailian Pro&lt;/strong&gt; or &lt;strong&gt;Volcengine Ark Pro&lt;/strong&gt; (¥200/month). Around 5x the quota of Lite, with free switching between multiple models. Volcengine also supports Auto smart scheduling.&lt;/li&gt;
&lt;li&gt;If you prefer GLM&amp;rsquo;s in-house models, consider Zhipu GLM, but note its weekly limit and peak-time quota multipliers (3x during peak, 2x off-peak).&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Pursuing ultimate response speed&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;If time-to-first-token and generation speed are your top priorities, &lt;strong&gt;Zhipu GLM&lt;/strong&gt;&amp;rsquo;s &lt;code&gt;glm-5-turbo&lt;/code&gt; (1.43s TTFT) and &lt;code&gt;glm-4.5-air&lt;/code&gt; (103 TPS) perform best.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;5-summary&#34;&gt;5. Summary
&lt;/h2&gt;&lt;p&gt;The Chinese Coding Plan market is iterating rapidly, with price wars and model wars running in parallel. When choosing, don&amp;rsquo;t fixate on surface prices; instead, focus on three core questions:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;What is the metering unit?&lt;/strong&gt; API requests, Prompt counts, or tokens? Different units can&amp;rsquo;t be compared directly.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;How does the quota mechanism work?&lt;/strong&gt; Refreshed every 5 hours, capped weekly, or capped monthly? This determines whether you can sustain high-intensity use.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Is the speed responsive?&lt;/strong&gt; TTFT and TPS directly affect the coding experience, and the same model can perform wildly differently across platforms.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;A final reminder: plan policies change frequently across vendors (e.g., Zhipu&amp;rsquo;s price increase, Kimi switching to token billing, Alibaba Cloud discontinuing its Lite tier), so be sure to confirm the latest details on each platform&amp;rsquo;s official website before subscribing.&lt;/p&gt;
&lt;h2 id=&#34;data-sources&#34;&gt;Data sources
&lt;/h2&gt;&lt;ul&gt;
&lt;li&gt;&lt;a class=&#34;link&#34; href=&#34;http://xhslink.com/o/2MUdNLQ7Uj7&#34;  target=&#34;_blank&#34; rel=&#34;noopener&#34;
    &gt;Xiaohongshu - Speed cross-test and price comparison of 9 China Coding Plans&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a class=&#34;link&#34; href=&#34;https://www.cnblogs.com/wzxNote/p/19648084&#34;  target=&#34;_blank&#34; rel=&#34;noopener&#34;
    &gt;Cnblogs - Full comparison of 2026 mainstream China AI Coding Plans | Developer pitfall guide&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a class=&#34;link&#34; href=&#34;https://cloud.tencent.com/document/product/1823/130092&#34;  target=&#34;_blank&#34; rel=&#34;noopener&#34;
    &gt;Tencent Cloud - Coding Plan overview&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a class=&#34;link&#34; href=&#34;https://zhuanlan.zhihu.com/p/2011769182103021566&#34;  target=&#34;_blank&#34; rel=&#34;noopener&#34;
    &gt;Zhihu - Alibaba Cloud Bailian Coding Plan first purchase as low as ¥7.9&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a class=&#34;link&#34; href=&#34;https://www.volcengine.com/article/37524&#34;  target=&#34;_blank&#34; rel=&#34;noopener&#34;
    &gt;Volcengine - Ark Coding Plan: AI coding service and pricing details&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a class=&#34;link&#34; href=&#34;https://www.bigmodel.cn/glm-coding&#34;  target=&#34;_blank&#34; rel=&#34;noopener&#34;
    &gt;Zhipu AI - GLM Coding Plan official site&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a class=&#34;link&#34; href=&#34;https://zhuanlan.zhihu.com/p/2010413265843422319&#34;  target=&#34;_blank&#34; rel=&#34;noopener&#34;
    &gt;Zhihu - Hands-on MiniMax M2.5: open-source disruptor, value-for-money king?&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
</description>
        </item>
        <item>
        <title>Claude Rolls Out Identity Verification: Impact on Users in Unsupported Regions and How to Respond</title>
        <link>https://torchtree.com/en/post/claude-idv-china-impact/</link>
        <pubDate>Wed, 15 Apr 2026 15:22:53 +0800</pubDate>
        
        <guid>https://torchtree.com/en/post/claude-idv-china-impact/</guid>
        <description>&lt;p&gt;Anthropic recently rolled out an identity verification mechanism on the Claude platform. Based on the official help center documentation and the list of supported countries and regions, this article summarizes the core points of the mechanism, with a focus on assessing the actual impact on users in unsupported regions (e.g., mainland China) and viable strategies.&lt;/p&gt;
&lt;h2 id=&#34;1-core-points-of-the-identity-verification-mechanism&#34;&gt;1. Core points of the identity verification mechanism
&lt;/h2&gt;&lt;h3 id=&#34;11-purpose&#34;&gt;1.1 Purpose
&lt;/h3&gt;&lt;p&gt;According to the official statement, identity verification aims to:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Prevent technical abuse&lt;/li&gt;
&lt;li&gt;Enforce platform usage policies&lt;/li&gt;
&lt;li&gt;Fulfill legal and security compliance obligations&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;When users access certain features, a verification prompt may appear — this is part of routine platform integrity checks.&lt;/p&gt;
&lt;h3 id=&#34;12-verification-methods-and-required-materials&#34;&gt;1.2 Verification methods and required materials
&lt;/h3&gt;&lt;p&gt;The verification service is powered by &lt;strong&gt;Persona Identities&lt;/strong&gt;. Users need to prepare:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;A valid government-issued photo ID&lt;/strong&gt; (passport, driver&amp;rsquo;s license/state ID, or national ID)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;A device with a camera&lt;/strong&gt;: a live selfie may be required&lt;/li&gt;
&lt;li&gt;The verification process typically takes no more than five minutes&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Document types that are NOT accepted include&lt;/strong&gt;: photocopies/screenshots/scans, digital IDs (such as electronic driver&amp;rsquo;s licenses), student/employee/bank cards, and temporary paper IDs.&lt;/p&gt;
&lt;h3 id=&#34;13-data-privacy-statement&#34;&gt;1.3 Data privacy statement
&lt;/h3&gt;&lt;p&gt;The official documentation highlights the following privacy protections:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Anthropic is the &lt;strong&gt;data controller&lt;/strong&gt; of verification data; Persona only processes data as instructed&lt;/li&gt;
&lt;li&gt;ID documents and selfies are &lt;strong&gt;not stored directly by Anthropic&lt;/strong&gt; but kept on the Persona platform&lt;/li&gt;
&lt;li&gt;Data is used only for identity confirmation and legal/security obligations, &lt;strong&gt;not for model training&lt;/strong&gt;, and is not shared with third parties for marketing or advertising&lt;/li&gt;
&lt;li&gt;Both transmission and storage are encrypted&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&#34;14-why-accounts-can-still-be-banned-after-verification&#34;&gt;1.4 Why accounts can still be banned after verification
&lt;/h3&gt;&lt;p&gt;The official documentation explicitly lists four situations that can lead to account suspension:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Repeated violations of usage policies&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Creating an account from an unsupported location&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;Violating the Terms of Service&lt;/li&gt;
&lt;li&gt;Use by individuals under 18&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This point is especially critical for users in unsupported regions: even if identity verification passes, if the system determines the account originates from an unsupported region, there is still a risk of a ban.&lt;/p&gt;
&lt;h2 id=&#34;2-real-impact-on-chinese-users&#34;&gt;2. Real impact on Chinese users
&lt;/h2&gt;&lt;h3 id=&#34;21-mainland-china-is-not-on-the-official-supported-list&#34;&gt;2.1 Mainland China is not on the official supported list
&lt;/h3&gt;&lt;p&gt;According to Anthropic&amp;rsquo;s &lt;a class=&#34;link&#34; href=&#34;https://www.anthropic.com/supported-countries&#34;  target=&#34;_blank&#34; rel=&#34;noopener&#34;
    &gt;Supported Countries and Regions&lt;/a&gt; page, &lt;strong&gt;mainland China does not appear in the supported list for Claude.ai or the commercial API&lt;/strong&gt;. The list includes Taiwan, but not mainland China, Hong Kong, or Macau.&lt;/p&gt;
&lt;h3 id=&#34;22-core-impact-assessment&#34;&gt;2.2 Core impact assessment
&lt;/h3&gt;&lt;table&gt;
  &lt;thead&gt;
      &lt;tr&gt;
          &lt;th&gt;Impact dimension&lt;/th&gt;
          &lt;th&gt;Details&lt;/th&gt;
      &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
      &lt;tr&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;Ban risk shifts from implicit to explicit&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
          &lt;td&gt;Previously, the platform mainly relied on IP, payment methods, and login behavior for risk control; now identity verification is directly tied to regional policy. Even after passing verification with a genuine Chinese passport, users may still be banned for &amp;ldquo;originating from an unsupported region.&amp;rdquo;&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;Account acquisition threshold rises significantly&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
          &lt;td&gt;For users relying on shared accounts or SMS-activation platforms, completing a &amp;ldquo;physical ID + live selfie&amp;rdquo; verification process is nearly impossible.&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;Cross-border data concerns&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
          &lt;td&gt;Chinese users must submit ID documents and biometric data to Persona, a US third-party service provider, raising privacy and cross-border data compliance concerns.&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;Low success rate for appeals&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
          &lt;td&gt;If the ban reason is &amp;ldquo;originating from an unsupported region,&amp;rdquo; this is an explicit policy red line for Anthropic rather than a system error, so the likelihood of a successful appeal is low.&lt;/td&gt;
      &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;
&lt;h3 id=&#34;23-behaviors-more-likely-to-trigger-verification&#34;&gt;2.3 Behaviors more likely to trigger verification
&lt;/h3&gt;&lt;p&gt;Although the official trigger conditions are not fully disclosed, based on standard platform risk-control logic, the following behaviors are more likely to trigger identity verification:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;High-frequency API calls or anomalous usage patterns&lt;/li&gt;
&lt;li&gt;Upgrading to Claude Pro / purchasing large API credits&lt;/li&gt;
&lt;li&gt;Frequently switching IP addresses or login devices&lt;/li&gt;
&lt;li&gt;Accounts being reported or generating policy-violating content&lt;/li&gt;
&lt;li&gt;Using unconventional payment methods such as virtual credit cards or gift cards&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;3-response-strategies&#34;&gt;3. Response strategies
&lt;/h2&gt;&lt;h3 id=&#34;31-for-existing-accounts-reduce-the-chance-of-being-triggered&#34;&gt;3.1 For existing accounts: reduce the chance of being triggered
&lt;/h3&gt;&lt;p&gt;If you still hold a usable account, the following measures can help extend its useful life:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Keep a low profile&lt;/strong&gt;: try to avoid triggering high-risk-control features&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Stabilize your network environment&lt;/strong&gt;: minimize IP hopping and maintain consistent access patterns&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Comply with usage policies&lt;/strong&gt;: avoid generating prohibited content to reduce the chance of being reported&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;It should be clear that these measures can only lower the probability of being triggered; they cannot eliminate the structural risk of being in an &amp;ldquo;unsupported region.&amp;rdquo;&lt;/p&gt;
&lt;h3 id=&#34;32-long-term-alternatives&#34;&gt;3.2 Long-term alternatives
&lt;/h3&gt;&lt;p&gt;Given the low compliance ceiling for personal accounts, it&amp;rsquo;s advisable to migrate to more stable, sustainable channels:&lt;/p&gt;
&lt;table&gt;
  &lt;thead&gt;
      &lt;tr&gt;
          &lt;th&gt;Option&lt;/th&gt;
          &lt;th&gt;Details&lt;/th&gt;
      &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
      &lt;tr&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;Enterprise cloud platforms&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
          &lt;td&gt;Access the Claude API through enterprise-grade platforms such as &lt;!-- raw HTML omitted --&gt;Amazon Bedrock&lt;!-- raw HTML omitted --&gt;, &lt;!-- raw HTML omitted --&gt;Google Cloud Vertex AI&lt;!-- raw HTML omitted --&gt;, and &lt;!-- raw HTML omitted --&gt;Microsoft Foundry&lt;!-- raw HTML omitted --&gt;. These channels target enterprise users, are more compliant, and are not directly subject to the regional verification restrictions of personal accounts.&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;Official partners&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
          &lt;td&gt;Watch whether Anthropic provides services in mainland China through authorized partners.&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;Local alternative models&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
          &lt;td&gt;For daily work and development, migrate to models that operate compliantly in mainland China, such as &lt;!-- raw HTML omitted --&gt;DeepSeek&lt;!-- raw HTML omitted --&gt;, &lt;!-- raw HTML omitted --&gt;Qwen&lt;!-- raw HTML omitted --&gt;, &lt;!-- raw HTML omitted --&gt;ERNIE Bot&lt;!-- raw HTML omitted --&gt;, and &lt;!-- raw HTML omitted --&gt;Zhipu AI (GLM)&lt;!-- raw HTML omitted --&gt;.&lt;/td&gt;
      &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;
&lt;h3 id=&#34;33-if-youve-already-been-banned&#34;&gt;3.3 If you&amp;rsquo;ve already been banned
&lt;/h3&gt;&lt;p&gt;The official &lt;a class=&#34;link&#34; href=&#34;https://docs.google.com/forms/d/e/1FAIpQLSdcTocgFJXSJzFJzVc47nxKmjeVhXDfgRaifH3DUZhYarA8vA/viewform&#34;  target=&#34;_blank&#34; rel=&#34;noopener&#34;
    &gt;appeal form&lt;/a&gt; lets users submit a review request. However, if the ban reason is &amp;ldquo;originating from an unsupported region,&amp;rdquo; the chance of a successful appeal is very low. It&amp;rsquo;s better to prioritize backing up and migrating your data and session records.&lt;/p&gt;
&lt;h3 id=&#34;34-risks-of-creating-a-new-account&#34;&gt;3.4 Risks of creating a new account
&lt;/h3&gt;&lt;p&gt;Under the current policy, the long-term viability of a new mainland China user registering a personal Claude account and passing identity verification is highly uncertain. If there&amp;rsquo;s a genuine need, it&amp;rsquo;s recommended to apply through &lt;strong&gt;enterprise-grade cloud services&lt;/strong&gt; or a &lt;strong&gt;physical office environment in a supported region&lt;/strong&gt; in a compliant way, rather than relying on SMS-activation codes or virtual payment methods.&lt;/p&gt;
&lt;h2 id=&#34;4-summary&#34;&gt;4. Summary
&lt;/h2&gt;&lt;p&gt;Claude&amp;rsquo;s identity verification mechanism elevates &amp;ldquo;regional restrictions&amp;rdquo; from a background risk-control rule to a front-facing compliance hurdle. This means:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Technical and payment-based workarounds (VPNs, virtual numbers, shared accounts) largely fail in the face of identity verification&lt;/li&gt;
&lt;li&gt;Even after passing verification, &amp;ldquo;unsupported region&amp;rdquo; remains a red line that can directly trigger a ban&lt;/li&gt;
&lt;li&gt;For long-term, stable business needs, &lt;strong&gt;migrating to enterprise cloud channels or local alternative models&lt;/strong&gt; is the more sustainable strategy&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This change also reflects the trend of leading AI platforms tightening compliance boundaries during global expansion. For users in unsupported regions, accepting this structural constraint and planning alternatives in advance is more rational than continuously investing in circumvention.&lt;/p&gt;
</description>
        </item>
        <item>
        <title>Nous Research Subscription Models: Prepaid vs Pay-as-you-go, Which Is Better Value?</title>
        <link>https://torchtree.com/en/post/nous-research-portal-pricing-guide/</link>
        <pubDate>Tue, 14 Apr 2026 03:57:09 +0800</pubDate>
        
        <guid>https://torchtree.com/en/post/nous-research-portal-pricing-guide/</guid>
        <description>&lt;p&gt;Recently, Nous Research&amp;rsquo;s API service, Portal, has gradually drawn the attention of developers. This article systematically maps out its pricing system based on official documentation and public information, clarifies core concepts such as &amp;ldquo;credits,&amp;rdquo; &amp;ldquo;tokens,&amp;rdquo; and &amp;ldquo;rate limits,&amp;rdquo; and compares the use cases for the two consumption modes.&lt;/p&gt;
&lt;h2 id=&#34;1-core-concepts-the-exchange-relationship-between-credits-and-tokens&#34;&gt;1. Core Concepts: The Exchange Relationship Between Credits and Tokens
&lt;/h2&gt;&lt;p&gt;In Nous Portal&amp;rsquo;s billing system, &lt;strong&gt;1 Credit = $1 USD&lt;/strong&gt;. This means:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The &amp;ldquo;$10.00 monthly credits&amp;rdquo; in a subscription plan is equivalent to $10 of API usage credit.&lt;/li&gt;
&lt;li&gt;Model pricing (e.g., Hermes-4-70B at $0.13/1M input tokens) is priced directly in USD.&lt;/li&gt;
&lt;li&gt;The actual amount of tokens a credit converts to depends on the unit price of the chosen model.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Conversion example&lt;/strong&gt; (using Basic plan&amp;rsquo;s $10 credits):&lt;/p&gt;
&lt;table&gt;
  &lt;thead&gt;
      &lt;tr&gt;
          &lt;th&gt;Model&lt;/th&gt;
          &lt;th&gt;Input price/1M&lt;/th&gt;
          &lt;th&gt;Input tokens purchased with $10&lt;/th&gt;
      &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
      &lt;tr&gt;
          &lt;td&gt;DeepHermes 3 (Mistral 24B)&lt;/td&gt;
          &lt;td&gt;$0.02&lt;/td&gt;
          &lt;td&gt;500 million tokens&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;Hermes-4-70B&lt;/td&gt;
          &lt;td&gt;$0.13&lt;/td&gt;
          &lt;td&gt;~76.9 million tokens&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;Hermes-4-405B&lt;/td&gt;
          &lt;td&gt;$1.00&lt;/td&gt;
          &lt;td&gt;10 million tokens&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;Claude Opus 4.6&lt;/td&gt;
          &lt;td&gt;$5.00&lt;/td&gt;
          &lt;td&gt;2 million tokens&lt;/td&gt;
      &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Source: &lt;a class=&#34;link&#34; href=&#34;https://portal.nousresearch.com/models&#34;  target=&#34;_blank&#34; rel=&#34;noopener&#34;
    &gt;Nous Portal Models page&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;As shown, when using high-performance models (such as Claude Opus 4.6), the amount of tokens your credits buy can differ by dozens of times.&lt;/p&gt;
&lt;h2 id=&#34;2-the-two-consumption-modes-explained&#34;&gt;2. The Two Consumption Modes Explained
&lt;/h2&gt;&lt;p&gt;According to the official API documentation, Nous Portal offers &lt;strong&gt;two independent consumption modes&lt;/strong&gt;:&lt;/p&gt;
&lt;h3 id=&#34;mode-a-subscription-prepaid&#34;&gt;Mode A: Subscription (prepaid)
&lt;/h3&gt;&lt;table&gt;
  &lt;thead&gt;
      &lt;tr&gt;
          &lt;th&gt;Plan&lt;/th&gt;
          &lt;th&gt;Monthly fee&lt;/th&gt;
          &lt;th&gt;Credits received&lt;/th&gt;
          &lt;th&gt;Rate Limits&lt;/th&gt;
      &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
      &lt;tr&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;Free&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
          &lt;td&gt;$0&lt;/td&gt;
          &lt;td&gt;Trial credits (small amount)&lt;/td&gt;
          &lt;td&gt;50 RPM, 100K TPM&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;Basic&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
          &lt;td&gt;$10&lt;/td&gt;
          &lt;td&gt;$10.00&lt;/td&gt;
          &lt;td&gt;400 RPM, 2M TPM&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;Plus&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
          &lt;td&gt;$20&lt;/td&gt;
          &lt;td&gt;$20.00&lt;/td&gt;
          &lt;td&gt;400 RPM, 4M TPM&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;Scale&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
          &lt;td&gt;$50&lt;/td&gt;
          &lt;td&gt;$50.00&lt;/td&gt;
          &lt;td&gt;600 RPM, 6M TPM&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;Max&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
          &lt;td&gt;$100&lt;/td&gt;
          &lt;td&gt;$100.00&lt;/td&gt;
          &lt;td&gt;800 RPM, 8M TPM&lt;/td&gt;
      &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Source: &lt;a class=&#34;link&#34; href=&#34;https://portal.nousresearch.com/api-docs&#34;  target=&#34;_blank&#34; rel=&#34;noopener&#34;
    &gt;Nous Portal API Docs&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Analysis of characteristics&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A fixed monthly charge that grants equal-value credits.&lt;/li&gt;
&lt;li&gt;Access to rate limits higher than the &amp;ldquo;default paid user&amp;rdquo; tier.&lt;/li&gt;
&lt;li&gt;Credits may reset at the end of the month (confirm the exact policy on the billing page).&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&#34;mode-b-pay-as-you-go-direct-top-up&#34;&gt;Mode B: Pay-as-you-go (direct top-up)
&lt;/h3&gt;&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;No subscription required&lt;/strong&gt;; add API credits directly to your account (any amount).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Rate Limits&lt;/strong&gt;: 200 RPM, 800,000 TPM (corresponding to the &amp;ldquo;Default paid users&amp;rdquo; level in the docs).&lt;/li&gt;
&lt;li&gt;Use until depleted, top up at any time.&lt;/li&gt;
&lt;li&gt;Credits may not expire (per the WorldSim Terms of Service: &amp;ldquo;Purchased credits will not expire unless the application(s) is retired&amp;rdquo;).&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Source: &lt;a class=&#34;link&#34; href=&#34;https://worldsim.nousresearch.com/terms&#34;  target=&#34;_blank&#34; rel=&#34;noopener&#34;
    &gt;WorldSim Terms of Service&lt;/a&gt;&lt;/p&gt;
&lt;h2 id=&#34;3-key-differences-vs-openrouter&#34;&gt;3. Key Differences vs OpenRouter
&lt;/h2&gt;&lt;table&gt;
  &lt;thead&gt;
      &lt;tr&gt;
          &lt;th&gt;Feature&lt;/th&gt;
          &lt;th&gt;&lt;!-- raw HTML omitted --&gt;Nous Portal&lt;!-- raw HTML omitted --&gt;&lt;/th&gt;
          &lt;th&gt;&lt;!-- raw HTML omitted --&gt;OpenRouter&lt;!-- raw HTML omitted --&gt;&lt;/th&gt;
      &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
      &lt;tr&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;Consumption mode&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
          &lt;td&gt;① Subscription prepaid ($10-100/month); ② Direct credit top-up (any amount)&lt;/td&gt;
          &lt;td&gt;Pure pay-as-you-go (no monthly barrier)&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;Rate Limits&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
          &lt;td&gt;Determined by subscription tier: Free(50) → Basic(400) → Max(800) RPM&lt;/td&gt;
          &lt;td&gt;Dynamically adjusted by usage&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;Credits mechanism&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
          &lt;td&gt;Prepaid credits, may reset monthly&lt;/td&gt;
          &lt;td&gt;Credits never expire&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;Anonymous payment&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
          &lt;td&gt;Not supported&lt;/td&gt;
          &lt;td&gt;Supports x402 protocol (Solana USDC)&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;&lt;!-- raw HTML omitted --&gt;Model coverage&lt;!-- raw HTML omitted --&gt;&lt;/td&gt;
          &lt;td&gt;400+ models, incl. OpenAI/Anthropic/open-source&lt;/td&gt;
          &lt;td&gt;200+ models, aggregates multiple providers&lt;/td&gt;
      &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Source: &lt;a class=&#34;link&#34; href=&#34;https://portal.nousresearch.com/api-docs&#34;  target=&#34;_blank&#34; rel=&#34;noopener&#34;
    &gt;Nous Portal API Docs&lt;/a&gt;, &lt;a class=&#34;link&#34; href=&#34;https://openrouter.ai/&#34;  target=&#34;_blank&#34; rel=&#34;noopener&#34;
    &gt;OpenRouter official site&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Key difference&lt;/strong&gt;: Nous Portal&amp;rsquo;s subscription model essentially trades &amp;ldquo;monthly fees for higher rate limits + equal-value credits&amp;rdquo;; OpenRouter is pure per-usage billing with no subscription barrier.&lt;/p&gt;
&lt;h2 id=&#34;4-billing-formula-and-cost-estimation&#34;&gt;4. Billing Formula and Cost Estimation
&lt;/h2&gt;&lt;p&gt;&lt;strong&gt;API call cost formula&lt;/strong&gt;:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Model price reference&lt;/strong&gt; (from the official Models page):&lt;/p&gt;
&lt;table&gt;
  &lt;thead&gt;
      &lt;tr&gt;
          &lt;th&gt;Model&lt;/th&gt;
          &lt;th&gt;Input/1M&lt;/th&gt;
          &lt;th&gt;Output/1M&lt;/th&gt;
      &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
      &lt;tr&gt;
          &lt;td&gt;DeepHermes 3 Mistral 24B&lt;/td&gt;
          &lt;td&gt;$0.02&lt;/td&gt;
          &lt;td&gt;$0.10&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;Hermes-4-70B&lt;/td&gt;
          &lt;td&gt;$0.13&lt;/td&gt;
          &lt;td&gt;$0.40&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;Hermes-4-405B&lt;/td&gt;
          &lt;td&gt;$1.00&lt;/td&gt;
          &lt;td&gt;$3.00&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;GPT-5.4&lt;/td&gt;
          &lt;td&gt;$2.50&lt;/td&gt;
          &lt;td&gt;$15.00&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;Claude Opus 4.6&lt;/td&gt;
          &lt;td&gt;$5.00&lt;/td&gt;
          &lt;td&gt;$25.00&lt;/td&gt;
      &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Source: &lt;a class=&#34;link&#34; href=&#34;https://portal.nousresearch.com/models&#34;  target=&#34;_blank&#34; rel=&#34;noopener&#34;
    &gt;Nous Portal Models page&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Real example&lt;/strong&gt;: if you use Hermes-4-70B to process one conversation (2K input tokens, 1K output tokens):&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Cost = (2000/1M × $0.13) + (1000/1M × $0.40) = $0.00026 + $0.0004 = &lt;strong&gt;~$0.00066&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;$10 in credits can support about &lt;strong&gt;15,000&lt;/strong&gt; such conversations.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;5-selection-advice-and-potential-caveats&#34;&gt;5. Selection Advice and Potential Caveats
&lt;/h2&gt;&lt;h3 id=&#34;recommended-options-for-different-scenarios&#34;&gt;Recommended options for different scenarios
&lt;/h3&gt;&lt;table&gt;
  &lt;thead&gt;
      &lt;tr&gt;
          &lt;th&gt;User type&lt;/th&gt;
          &lt;th&gt;Recommended approach&lt;/th&gt;
          &lt;th&gt;Reason&lt;/th&gt;
      &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
      &lt;tr&gt;
          &lt;td&gt;Trial/light use&lt;/td&gt;
          &lt;td&gt;Free tier → top up a small amount of credits&lt;/td&gt;
          &lt;td&gt;No monthly commitment&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;Developer/moderate traffic&lt;/td&gt;
          &lt;td&gt;Basic Subscription ($10)&lt;/td&gt;
          &lt;td&gt;400 RPM is enough for daily use; $10 credits are roughly consumed&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;High-traffic production&lt;/td&gt;
          &lt;td&gt;Scale/Max Subscription&lt;/td&gt;
          &lt;td&gt;600-800 RPM avoids 429 throttling&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;Agent developers&lt;/td&gt;
          &lt;td&gt;Plus ($20)&lt;/td&gt;
          &lt;td&gt;Balance between rate limit and credit amount&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;Uncertain usage&lt;/td&gt;
          &lt;td&gt;Direct credit top-up&lt;/td&gt;
          &lt;td&gt;Avoid wasting subscription credits that reset monthly&lt;/td&gt;
      &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;
&lt;h3 id=&#34;potential-issues-to-note&#34;&gt;Potential issues to note
&lt;/h3&gt;&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Credit reset risk&lt;/strong&gt;: subscription monthly credits may reset at the end of the month (confirm the exact policy by logging into the portal billing page).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Rate limit difference&lt;/strong&gt;: the top-up mode offers only 200 RPM; high-concurrency scenarios can easily trigger throttling.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Huge model price gaps&lt;/strong&gt;: Claude Opus 4.6 costs 38× more than Hermes-4-70B; choosing the wrong model can drain credits quickly.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Refund policy&lt;/strong&gt;: per the Terms of Service, &amp;ldquo;All Fees are non-refundable.&amp;rdquo;&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Source: &lt;a class=&#34;link&#34; href=&#34;https://portal.nousresearch.com/terms&#34;  target=&#34;_blank&#34; rel=&#34;noopener&#34;
    &gt;Nous Portal Terms of Service&lt;/a&gt;&lt;/p&gt;
&lt;h2 id=&#34;6-summary&#34;&gt;6. Summary
&lt;/h2&gt;&lt;p&gt;Nous Research Portal&amp;rsquo;s pricing system is designed to serve both &amp;ldquo;stable subscription users&amp;rdquo; and &amp;ldquo;flexible on-demand users&amp;rdquo;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Subscription mode&lt;/strong&gt; suits production environments that need a stable, high-concurrency rate limit.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Top-up mode&lt;/strong&gt; suits scenarios with fluctuating usage where users want to avoid being bound by a monthly fee.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Developers should choose the most suitable billing method based on their token consumption, concurrency needs, and cost budget. For users who are uncertain, it&amp;rsquo;s recommended to first try the Free tier or a small top-up, then decide whether to upgrade the subscription based on actual usage.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;This article is compiled from Nous Research&amp;rsquo;s official documentation and public materials; specific policies are subject to the latest official announcements.&lt;/em&gt;&lt;/p&gt;
</description>
        </item>
        
    </channel>
</rss>
