服务尚未上线Not live yet · 现在只登记意向,不收钱we are only collecting interest — nobody pays anything yet
正在筹备 · 满 80 人开机 In the works · 80 people and we launch

人人用得起的 AI AI that everyone can afford

用的是开源 Qwen3.8-27B——它和 Claude Opus 4.6 在编码与 agent 类基准上互有胜负,价格却只有 1%:计划价 ¥4.9/月起(≈2 亿 token)。没有 5 小时窗口,没有每周限额。服务还没上线:满 80 个想用的人,我们就开机。 Built on open-weight Qwen3.8-27B — neck and neck with Claude Opus 4.6 on coding and agent benchmarks, at 1% of the price: planned from $0.73/month (≈200M tokens). No 5-hour window, no weekly cap. Not live yet: 80 people who want it and we turn it on.

open-weight Qwen3.8-27B · ~1% of Opus 4.6 pricing · plans from $0.73, or pay as you go

现在登记不收钱 · 开机后第一批邀请发给你 · 不满意随时退出 Signing up is free · first invites go to the list · leave anytime

POST https://api.everyoneai.cc/v1/chat/completions开机后可用available at launch
# 换掉 base_url 即可,其余代码不用动
curl https://api.everyoneai.cc/v1/chat/completions \
  -H "Authorization: Bearer $EVERYONEAI_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen3.8-27b",
    "messages": [
      {"role": "system", "content": "<20 万字的系统提示词 / 知识库>"},
      {"role": "user",   "content": "总结第三章的结论"}
    ],
    "stream": true,
    "max_tokens": 2048
  }'

# 计费口径(每百万 token / per 1M tokens)
prompt_cache_hit_tokens   ¥0.01$0.002
prompt_cache_miss_tokens  ¥0.20$0.03
completion_tokens          ¥1.15$0.17
开机进度Launch progress 目标 80 人target 80
0 / 80
目标Goal
80 人80 people
已登记Signed up
0
还差Still needed
80 人80
达成后Once reached
1 周内开机live within 1 week
现在登记不收钱。满 80 个真的想用的人,我们就去买机器、开机,第一批邀请发给登记的人。不满 80,说明这件事现在还不该做,我们继续等——也不会收任何人的钱。 Signing up costs nothing. At 80 people who actually want this, we buy the hardware and turn it on, and the first invites go to everyone on the list. Below 80, we keep waiting — and we never take anyone's money in the meantime.
套餐Token plans

一个月一杯奶茶的钱,用到够Less than a coffee a month — and it is enough

开机后的计划价格:¥4.9 / ¥9.9 / ¥19.9。额度用完后自动转按量付费——不降速、不断服。Planned pricing at launch: $0.73, $1.47 and $2.95. When the quota runs out you roll onto pay-as-you-go — no throttling, no cut-off.

尝鲜包Starter
¥4.9$0.73 / 月/ month

先试试好不好用。一顿早餐的钱。Try it out. The price of a coffee.

  • ≈ Token 总量≈ Total tokens2 亿200M
  • 计费额度Metered quota700 万单位7M units
  • ≈ Agent 调用≈ Agent calls1,800 次1,800
  • 折合Unit price¥0.700 / M$0.104 / M
  • 超出后Overflow转按量付费pay-as-you-go
标准包 · 最受欢迎Standard · most popular
¥9.9$1.47 / 月/ month

日常主力。够一个 Agent 开发者重度使用。The daily driver. Enough for an agent developer running heavy.

  • ≈ Token 总量≈ Total tokens4 亿400M
  • 计费额度Metered quota1,400 万单位14M units
  • ≈ Agent 调用≈ Agent calls3,500 次3,500
  • 折合Unit price¥0.707 / M$0.105 / M
  • 超出后Overflow转按量付费pay-as-you-go
生产力包Pro
¥19.9$2.95 / 月/ month

我们的价格上限。多项目并行、团队共享都够。Our ceiling price. Several projects or a small team.

  • ≈ Token 总量≈ Total tokens8 亿800M
  • 计费额度Metered quota2,800 万单位28M units
  • ≈ Agent 调用≈ Agent calls7,000 次7,000
  • 折合Unit price¥0.711 / M$0.105 / M
  • 超出后Overflow转按量付费pay-as-you-go
这些价格要 80 个人才能开起来These prices need 80 people to happen
现在登记不收钱,满 80 人我们开机Signing up is free — at 80 people we turn it on
登记意向Sign up

Token 量按重度 Agent 用量估算:输入:输出 ≈ 110:1,其中 97.5% 的输入命中缓存,单次调用约 11 万 token 上下文 + 1,000 token 输出。此时 1 单位 ≈ 28 个 token,一次调用约 3,970 单位。 Token estimates assume heavy agent usage: input:output ≈ 110:1, with 97.5% of input served from cache, roughly 110K tokens of context per call plus 1,000 output tokens. That is ~28 tokens per unit, or ~3,970 units per call.

对照普通对话(1.5K 输入、30% 缓存、600 输出):1 次对话约 777 单位,同样额度能跑的对话次数是上面 Agent 调用次数的 5 倍以上。缓存命中率越高、额度越耐用。 For plain chat (1.5K in, 30% cached, 600 out) a turn is ~777 units — the same quota covers more than 5× as many of those. The higher your cache-hit rate, the further your quota goes.

每月跑 1.5 万次重度 Agent 调用的用户,约需 6,000 万单位:¥19.9 档覆盖约一半,其余自动转按量付费,不会停服。 A user running 15,000 heavy agent calls a month needs roughly 60M units: the $2.95 plan covers about half, and the rest rolls onto metered billing. Nothing stops.

额度用完了怎么办?不会停服——自动转成按量付费,不降速、不排队降级。你也可以在控制台设一个月度封顶,到点自动停。 What happens when the quota runs out? Nothing stops. You roll onto pay-as-you-go — no throttling, no deprioritized queue. You can also set a monthly spend cap in the console and we stop there.

兜底费率(按量付费,每百万 token)Overflow rates (pay-as-you-go, per 1M tokens)
  • 输入 ≤64K(未命中)Input ≤64K (miss)¥0.20$0.03
  • 输入 >64K(未命中)Input >64K (miss)¥0.35$0.05
  • 输入(缓存命中)Input (cache hit)¥0.01$0.002
  • 输出(标准档)Output (Standard)¥1.15$0.17
  • 输出(极速档)Output (Turbo)¥1.90$0.28
  • 输出(异步批处理)Output (Async batch)¥0.55$0.08
对比 · Coding Planvs Coding Plans

没有 5 小时窗口,没有每周限额No 5-hour window. No weekly cap.

你不需要换客户端——Claude Code、Cline、Cursor 都能继续用,只把 base_url 换成我们的。计费方式换掉之后,那些"撞墙"就消失了。You do not have to change your client — keep using Claude Code, Cline or Cursor and just point base_url at us. Change the billing model and the walls disappear.

对比项Item 传统 Coding PlanTypical coding plan everyoneai 套餐everyoneai plans
计费窗口Billing window 5 小时滚动窗口,窗口内超额直接限流5-hour rolling window; you get throttled inside it 无窗口none
周期上限Period cap 每周限额,还分「全模型周限」和「单模型周限」两档Weekly caps — often an all-models cap plus a single-model cap 无周限,只有月度额度no weekly cap, monthly quota only
撞限之后When you hit it 窗口内只能等重置;周限撞了要等满 7 天,或者升级、转 APIWait for the session reset — or, on the weekly cap, wait 7 days, upgrade, or move to API billing 自动转按量付费,不降速不断服auto-roll to pay-as-you-go, no throttle, no cut-off
窗口大小Window size 各家不同,有的更大varies; some are larger 262K,够读完一份合同/一本书
我们没去堆更大的窗口——窗口越大注意力越散、单位成本越高
262K — a full contract or a book
we did not chase a bigger window: bigger means more diffuse attention and a higher unit cost
客户端Client 各家绑定自家客户端tied to the vendor's own client 你现有的都能用(改 base_url)keep yours — change base_url
月费对比Monthly price 对方Theirs 我们Ours 便宜Cheaper by
GitHub Copilot Pro ¥67$10 ¥4.9$0.73 13.8×
Cursor Pro ¥135$20 ¥9.9$1.47 13.6×
Claude Pro ¥135$20 ¥9.9$1.47 13.6×
Claude Max 5× ¥674$100 ¥19.9$2.95 33.9×
Claude Max 20× ¥1,348$200 ¥19.9$2.95 67.7×

对比的是计费方式与价格,不是模型能力——这些是前沿模型的订阅,我们跑的是开源 Qwen3.8-27B(能力对比见上面一节)。我们做的是把 Anthropic 协议一比一兼容,让你现有的客户端可以直接接过来。英文界面显示美元价(按 1 USD = 6.74 CNY 折算),对方价格以各厂商最新公告为准。This compares billing mechanics and price, not model capability — those are frontier-model subscriptions and we serve open-weight Qwen3.8-27B (capability comparison is in the section above). What we did is make the Anthropic protocol drop-in compatible so your existing client can point at us. USD shown in the English interface, converted at 1 USD = 6.74 CNY; vendor prices subject to their latest announcements.

~1%
相对 Opus 4.6 的输出价格of Opus 4.6 output pricing
Qwen3.8-27B
开源模型,Apache 2.0open weights, Apache 2.0
6 / 6
与 Opus 4.6 的能力项对比(赢 3 输 3)capability rows vs Opus 4.6 — 3 wins, 3 losses
80
人登记就开机people and we launch
能力Capabilities

不是"便宜的小模型",是"够用而且不给你设墙的模型"Not a cheap small model — a capable one that does not put walls around you

开源 Qwen3.8-27B:编码与 agent 类基准和 Opus 4.6 互有胜负,权重以 Apache 2.0 开源,你可以自己部署、也可以直接用我们的。为 Agent、长文档、代码库问答和批量处理而设计。Open-weight Qwen3.8-27B: level with Opus 4.6 on coding and agent benchmarks, Apache 2.0, self-hostable or use ours. Built for agents, long documents, codebase Q&A and batch processing.

◧

262K 为什么够用Why 262K is enough

够读约 20 万字——一整份合同、一本书、一整个中型代码库。我们没有去堆"更大的窗口":窗口越大,注意力越分散、单位成本也越高,而日常真正吃 token 的是 Agent 的循环调用和对同一份文档的反复追问,那些靠前缀缓存解决,不靠把上下文拉长。窗口内的每一段都算得准、算得便宜,比一个用不满的大数字更值。Enough for roughly 200k Chinese characters — a full contract, a book, a mid-sized codebase. We deliberately did not chase a bigger window: the larger it gets, the more diffuse the attention and the higher the unit cost, while everyday token burn comes from agent loops and repeated questions over the same document — solved by prefix caching, not by stretching the context. Every token inside the window is accurate and cheap, which beats a big number you never fill.

◈

Qwen3.8-27B 的血缘Qwen3.8-27B lineage

底座是开源的 Qwen3.8-27B(Apache 2.0,可自托管)。我们跑的是它的省显存推理版:保留全精度模型 98.2% 的综合基准分,服务成本降到约十分之一。复杂推理、长链路难题仍然建议用前沿模型——我们做的是"够用 + 便宜 + 不设墙"。The base is open-weight Qwen3.8-27B (Apache 2.0, self-hostable). We serve a memory-lean inference build of it: 98.2% of the full-precision model's aggregate benchmark score at about a tenth of the serving cost. For the hardest reasoning and long-horizon problems, use a frontier model — ours is "good enough, cheap, and without walls".

◱

前缀缓存按命中计费Prefix caching, billed on hits

系统提示词、知识库、多轮历史只算一次。缓存命中 ¥0.01/M,只有未命中的二十分之一。System prompts, knowledge bases and conversation history are charged once. Cache hits cost $0.002/M — one twentieth of a miss.

⇄

双协议零改造Drop-in for two protocols

同时提供 OpenAI Chat Completions 与 Anthropic Messages 接口。现有 SDK 与 Agent 框架改 base_url 就能用。Both OpenAI Chat Completions and Anthropic Messages endpoints. Point your existing SDK or agent framework at a new base_url and you are done.

◐

视觉 · 工具调用 · 思考模式Vision · tools · thinking mode

支持图片与多图输入、function calling、流式输出;思考过程可保留,也可关闭以降低延迟与成本。Image and multi-image input, function calling, streaming. Keep or disable the reasoning trace to trade latency and cost.

◍

三档服务,按场景选Three tiers, pick your trade-off

标准档走批量通道;极速档独占低并发通道,所以更快;异步批处理最便宜,¥0.55/M,24 小时内交付。Standard runs on the batched lane. Turbo gets a dedicated low-concurrency lane, so it is faster. Async batch is cheapest at $0.08/M, delivered within 24 hours.

对比Comparison

一个 27B 的开源模型,追到了 Opus 4.6 的同一梯队A 27B open model that caught up with Opus 4.6

左边是 Anthropic 的前沿模型 Claude Opus 4.6(2026 年 2 月发布),右边是开源 Qwen3.8-27B(2026 年 8 月发布)——也就是我们跑的那个模型。能力上互有胜负,价格差约 100 倍;哪个更值,你自己判断。On the left, Anthropic's frontier Claude Opus 4.6 (released February 2026). On the right, open-weight Qwen3.8-27B (August 2026) — the model we serve. They trade wins by task, and the price differs by about 100×. You decide which one your job needs.

每百万 tokenPer 1M tokens Claude Opus 4.6 everyoneai.cc 便宜倍数Cheaper by
输入(未命中)Input (cache miss) ¥33.70$5.00 ¥0.20$0.03 135×
输入(缓存命中)Input (cache hit) ¥3.37$0.50 ¥0.01$0.002 112×
输出Output ¥168.50$25.00 ¥1.15$0.17 105×
模型Model Opus 4.6
前沿闭源
Opus 4.6
frontier, closed
Qwen3.8-27B
27B · Apache 2.0 开源
Qwen3.8-27B
27B · Apache 2.0
—
能力对比(越高越好)Capability (higher is better) Claude Opus 4.6 Max Qwen3.8-27B(我们)Qwen3.8-27B (ours) 谁领先Leader
编码 · SWE-bench ProCoding · SWE-bench Pro 53.4 61.7 我们 +8.3ours +8.3
编码 · LiveCodeBench v6Coding · LiveCodeBench v6 88.8 90.3 我们 +1.5ours +1.5
电脑操作 · OSWorld-VerifiedComputer use · OSWorld-Verified 72.7 84.3 我们 +11.6ours +11.6
终端任务 · Terminal-Bench 2.1Terminal · Terminal-Bench 2.1 78.2 73.0 Opus +5.2Opus +5.2
知识推理 · GPQA DiamondKnowledge · GPQA Diamond 91.3 89.2 Opus +2.1Opus +2.1
最难题 · Humanity's Last ExamHardest set · Humanity's Last Exam 40.0 30.8 Opus +9.2Opus +9.2

怎么读这张表:编码、电脑操作这类"手上有活"的任务,27B 的开源模型已经反超;纯知识与最难的推理题,Opus 4.6 仍然领先。所以我们的建议很直接——难题、长链路任务请去用 Opus;量大的、重复的、要跑一整天的活,用我们。两者的价格差约 100 倍,但能力差远没有 100 倍。How to read this: on hands-on work — coding, computer use — the 27B open model already leads; on pure knowledge and the hardest reasoning, Opus 4.6 still wins. So our advice is blunt: take the hardest, longest-horizon work to Opus; run the high-volume, repetitive, all-day work on us. The price gap is about 100×; the capability gap is nothing like it.

数据来源与口径:分数取自 Qwen3.8-27B 官方模型卡(2026-08)。其中 Opus 4.6 Max 为 Anthropic 官方公布分,Qwen3.8-27B 及同表基线是在 Claude Code harness 下重跑的(temperature 1.0、top_p 0.95),两边口径不完全一致,几分的差距不能当铁证。另外 Opus 4.6 自身有 Max / 标准等多个变体,与其他来源的分数不可混比——表内比较一律限定在同一张官方表格之内。另外说明:这一行比较的是 Qwen3.8-27B 这个模型本身,我们实际服务的是它的省显存推理版本,综合基准保留全精度模型的 98.2%,换来约 1/10 的服务成本。Source and caveats: scores are from the official Qwen3.8-27B model card (Aug 2026). Opus 4.6 Max figures are Anthropic's own published numbers, while Qwen3.8-27B and the baseline models were re-run under the Claude Code harness (temperature 1.0, top_p 0.95) — the setups are not identical, so a few points either way is not proof of anything. Opus 4.6 also ships in several variants, so its scores from other sources must not be mixed in; every comparison here stays inside that one official table. One more note: this row compares the Qwen3.8-27B model itself, while what we serve is a memory-lean inference build of it, retaining 98.2% of the full-precision model's aggregate benchmark score at roughly a tenth of the serving cost.

买套餐的话,输出成本可低至 ¥0.70 / 百万单位(¥19.9 生产力包),相当于 Opus 4.6 的 0.42%。On the $2.95 Pro plan, output-equivalent cost drops to $0.10 per 1M units — about 0.42% of Opus 4.6.

Opus 4.6 价格取自 Anthropic 2026-02 公开发布价,汇率按 1 USD = 6.74 CNY。以对方最新公告为准。Opus 4.6 pricing as published by Anthropic (Feb 2026); converted at 1 USD = 6.74 CNY. Subject to their latest published rates.

按量付费Pay as you go

套餐额度用完之后的兜底价The rate you roll onto after the quota

不买套餐也能直接用。三档服务按速度定价,单路 70–200 tok/s,输出 ¥0.55–1.90 / 百万 token。买套餐的话,同样的量再便宜 12–22%。No plan required. Three tiers priced by speed — 70 to 200 tok/s per stream, $0.08–0.28 per 1M output tokens. With a plan the same volume costs another 12–22% less.

服务档Tier 并发额度Concurrency 单路输出速度Speed per stream 输出价格 ¥/MOutput $/M 适用场景Best for
极速档Turbo 独占 ≤2 路dedicated ≤2 130–200 tok/s ¥1.90$0.28 交互式编程、实时客服live coding, real-time support
标准档Standard 共享 16 路池shared 16-lane pool 70–100 tok/s ¥1.15$0.17 对话、RAG、内容生成chat, RAG, content generation
异步批处理Async batch 队列调度queued 不承诺延迟no latency SLA ¥0.55$0.08 离线清洗、批量摘要、评测offline cleaning, bulk summarization, evals

速度为开机前的实测值(单路、短上下文)。长上下文(>128K)下单路输出约 90 tok/s;总吞吐随并发提升,满载时会按提交顺序排队,不会丢请求。Speeds are pre-launch measurements, single stream on short context. Beyond 128K context expect ~90 tok/s per stream. Total throughput rises with concurrency; when saturated, requests queue in submission order and are never dropped.

计费项Line item 换算成单位Units 说明Notes
输出 token(标准档)Output tokens (Standard)1 token = 1 单位基准单位the base unit
输入 token ≤64K(未命中)Input ≤64K (miss)1 token = 0.16 单位首次读入的提示词first read of a prompt
输入 token >64K(未命中)Input >64K (miss)1 token = 0.35 单位长文档一次性读入one-shot long-document read
输入 token(缓存命中)Input (cache hit)1 token = 0.02 单位相同前缀再次出现,几乎免费same prefix again — nearly free
输出 token(异步批处理)Output (async batch)1 token = 0.5 单位24 小时内交付,半价delivered within 24h, half price
按量价目Metered rates 人民币 / 百万 tokenUSD / 1M tokens 说明Notes
输入(≤64K,未命中)Input (≤64K, miss)¥0.20$0.03首次读入first read
输入(>64K,未命中)Input (>64K, miss)¥0.35$0.05长文档一次性读入long-document read
输入(缓存命中)Input (cache hit)¥0.01$0.002相同前缀再次出现same prefix again
输出(标准档)Output (Standard)¥1.15$0.17批量并发batched
输出(极速档)Output (Turbo)¥1.90$0.28独占低并发通道,单路 200+ tok/sdedicated lane, 200+ tok/s per stream
输出(异步批处理)Output (Async batch)¥0.55$0.0824 小时内交付delivered within 24h

英文界面显示美元价,按 1 USD = 6.74 CNY 折算并取整;人民币价格为准。USD prices are shown in the English interface, converted at 1 USD = 6.74 CNY and rounded; CNY is the reference price.

接入Get started

三行代码迁移Three lines to migrate

兼容 OpenAI 与 Anthropic 两套协议。已有代码只需替换 base_url 与 API Key。Compatible with both protocols. Replace the base_url and the API key — that is the whole migration.

Python · OpenAI SDK
from openai import OpenAI

client = OpenAI(
    base_url="https://api.everyoneai.cc/v1",
    api_key="$EVERYONEAI_KEY",
)

resp = client.chat.completions.create(
    model="qwen3.8-27b",
    messages=[{"role": "user",
               "content": "总结这份合同的风险点"}],
    stream=True,
)
Python · Anthropic SDK
from anthropic import Anthropic

client = Anthropic(
    base_url="https://api.everyoneai.cc",
    api_key="$EVERYONEAI_KEY",
)

msg = client.messages.create(
    model="qwen3.8-27b",
    max_tokens=2048,
    messages=[{"role": "user",
               "content": "总结这份合同的风险点"}],
)
常见问题FAQ

你可能想知道的Things you probably want to know

现在能用吗?Can I use it today?

还不能。机器一台没买,现在只登记意向,不收任何人的钱。满 80 个真的想用的人,我们就去买机器、开机,第一批邀请发给登记的人。不满 80,说明这件事现在还不该做,我们继续等。Not yet. No hardware has been bought and we take no money at this stage — we are only collecting interest. At 80 people who actually want this, we buy the hardware and turn it on, and the first invites go to the list. Below 80, we keep waiting.

现在能看到的是落地页和登记入口。服务本身会跑在开源 Qwen3.8-27B 上(Apache 2.0,你也可以自己部署一份),但我们不会等到"什么都完美"才开——先把 80 个真的想用的人凑齐。What exists today is this page and the sign-up. The service itself will run open-weight Qwen3.8-27B (Apache 2.0 — you could self-host it today), but we are not waiting for perfect before starting: first, 80 people who actually want it.

为什么是 80?因为这个规模能让我们把机器开起来、并撑住前三个月的运营。达不到就是白烧钱,我不想那样开始。Why 80? Because that is the scale at which the machine can be switched on and kept running for the first three months. Below it, we would just be burning money — and I would rather not start that way.

为什么是 262K?别人都 1M 了,够用吗?Why 262K when others offer 1M? Is that enough?

够用,而且是刻意选的。262K 约等于 20 万字,一次装得下一整份合同、一本书、一整个中型代码库。我们没去堆更大的窗口,原因有两个:① 便宜——窗口越大,同一张卡能同时服务的人越少,摊到你头上的成本越高;② 准——上下文越长,注意力越分散,中间部分容易被忽略。我们宁可把这 262K 里的每一段都答得准、算得便宜,也不给一个大部分人用不满的大数字。It is enough, and the size is deliberate. 262K is roughly 200k Chinese characters — a full contract, a book, a mid-sized codebase in one request. We did not chase a bigger window, for two reasons: (1) cost — the larger the window, the fewer people one card can serve at once, and the more it costs you; (2) accuracy — the longer the context, the more diffuse the attention, and the middle tends to get ignored. We would rather have every token inside 262K answered accurately and cheaply than ship a big number most people never fill.

真要读更长的东西也有解:前缀缓存让同一份长文档的反复追问只算一次钱;再长的资料可以切成多轮,或者走异步批处理。我们实测过 262,144 token 的整篇文档检索(信息埋在全文 33%、66%、90% 三个位置,全部按序取回),窗口内的可靠性是验证过的。For genuinely longer inputs there are answers: prefix caching bills repeated questions over the same long document once; anything longer can be split across turns or pushed through async batch. We did verify retrieval across a full 262,144-token document — information planted at 33%, 66% and 90% came back in order — so reliability inside the window is tested, not assumed.

用的是什么模型?是完整的 Qwen3.8-27B 吗?Which model do you serve — is it the full Qwen3.8-27B?

对外就叫 Qwen3.8-27B:开源、Apache 2.0、27B 稠密参数,你可以自己下载权重、自己部署。我们实际跑的是它的省显存推理版(压缩了权重表示,官方基准保留全精度模型 98.2% 的综合能力),换来的是单卡能同时服务更多人、价格压到约 1/10。所以如果你要的是"绝对满血",自托管原版即可;如果你要的是"便宜、不设墙、够用",用我们的。We call it Qwen3.8-27B: open weights, Apache 2.0, 27B dense, downloadable and self-hostable. What we actually serve is a memory-lean inference build of it (a compressed weight representation that keeps 98.2% of the full-precision model's aggregate benchmark score), which is what lets one card serve more people at roughly a tenth of the cost. Want the absolute full-fat version? Self-host the original. Want cheap, wall-free and good enough? Use ours.

为什么能比 Opus 4.6 便宜 100 倍?质量差多少?How can this be 100× cheaper than Opus 4.6? How much worse is it?

便宜来自三件事:这个 27B 模型的显存占用只有前沿模型的零头(一张消费级显卡就能跑)、我们把并发利用率拉满、不养任何品牌溢价。质量上,编码、电脑操作这些"手上有活"的任务它已经反超 Opus 4.6,知识推理和最难的那批题仍然落后(对比表在上面)。但要说清楚一点:连跑同一个开源模型的第三方 API,单价都比我们贵 5 倍左右——便宜不是因为模型小,是因为我们没有把中间利润堆上去。Three things: this 27B model's memory footprint is a fraction of a frontier model's (it runs on one consumer GPU), we keep concurrency utilization high, and we carry no brand premium. On quality it already leads Opus 4.6 on hands-on work like coding and computer use, and still trails on knowledge and the hardest reasoning (table above). One thing worth stating plainly: even third-party APIs serving the very same open model charge about 5× our price — the price is low because we did not stack margin on top, not because the model is small.

缓存命中怎么判定、怎么计费?How are cache hits determined and billed?

服务端按前缀精确匹配复用 KV 缓存。响应 usage 中区分 prompt_cache_hit_tokens 与 prompt_cache_miss_tokens,分别按 ¥0.01/M 与 ¥0.20/M 计费。固定系统提示词、同一份长文档的多轮追问、Agent 的重复上下文都能命中。We do exact prefix matching on the server and reuse the KV cache. The usage block separates prompt_cache_hit_tokens from prompt_cache_miss_tokens, billed at $0.002/M and $0.20/M respectively. Fixed system prompts, follow-up questions on the same long document and repeated agent context all hit.

容量用完了会怎样?扩容要多久?What happens when capacity runs out, and how long does scaling take?

我们会按需增加机器,但新机器需要「开机 → 拉取模型 → 预热」,约 5–10 分钟。满载时新请求会进入队列并按提交顺序处理,不会丢失。如果你有可预期的峰值,提前告知我们会事先扩容;紧急情况也可以通过控制台申请优先通道。We add machines on demand, but a new one has to boot, pull the model and warm up — roughly 5–10 minutes. When saturated, new requests queue in submission order and are never dropped. Tell us about expected peaks and we scale ahead of time; you can also request a priority lane from the console.

套餐额度怎么算?输入 token 也算吗?How is plan quota counted — do input tokens count too?

算,但按成本折算成"单位"。1 个输出 token = 1 单位;输入 token(≤64K)只算 0.16 单位;长文档输入(>64K)算 0.35 单位;缓存命中只要 0.02 单位。所以把系统提示词和长文档交给缓存,额度几乎不掉。完整换算表在上面的"按量付费"一节。Yes, but converted into units by cost. 1 output token = 1 unit; input tokens under 64K cost 0.16 units; long-document input above 64K costs 0.35 units; cache hits cost just 0.02 units. Reuse your system prompt and long documents through the cache and your quota barely moves. Full table is in the pay-as-you-go section above.

可以随时退订吗?用不完的额度能留到下个月吗?Can I cancel? Does unused quota roll over?

随时可退,按未使用天数比例退款。当月未用完的额度可结转一个月(第二个月底清零)。额度用完后不会停服,会自动转按量付费,你也可以在控制台设一个月度封顶自动停。Cancel anytime with a pro-rata refund for unused days. Unused quota rolls over for one month and expires at the end of the following month. When the quota runs out we do not stop you — you roll onto pay-as-you-go, and you can set a monthly spend cap in the console.

和 Claude Code 那种订阅有什么区别?How is this different from a Claude Code style subscription?

三点:没有 5 小时滚动窗口,没有每周限额,价格便宜一个数量级。传统订阅撞到周限后只能等满 7 天、升级或转 API;我们额度用完后自动转按量付费,不降速、不断服。而且我们不绑定客户端——Claude Code、Cline、Cursor 改个 base_url 就能接到我们这里。Three things: no 5-hour rolling window, no weekly cap, and an order of magnitude cheaper. On a classic subscription, hitting the weekly cap means waiting seven days, upgrading, or switching to API billing; with us the quota simply rolls into pay-as-you-go with no throttle and no cut-off. And we are not tied to a client — point Claude Code, Cline or Cursor at our base_url and you are done.

捐款真的会让我的额度变多吗?Will donating actually raise my quota?

会。捐款的 95% 进入额度池,用于增加算力和统一上调三档额度——上调对全体用户生效,包括没有捐款的用户,也包括你。剩下的 5% 用于支付通道手续费与账务成本。每月 1 号公示上月收到的金额、进入额度池的金额,以及额度实际提升了多少。额度池按累积制计算,累计达标后 30 天内完成上调。Yes. 95% of every donation goes into the quota pool, which is spent on adding capacity and raising all three plan quotas — the increase applies to every user, including those who never donated, and including you. The remaining 5% covers payment processing and accounting. On the 1st of each month we publish the total received, the amount that went into the pool, and how much the quotas actually rose. The pool is cumulative: increases land within 30 days of crossing a threshold.

捐款不产生任何额外权益:不会给你更多额度、不会插队、不会解锁功能。额度池的余额和流水在控制台随时可查。这不是慈善募捐,是用户自愿多付、用来把价格压给所有人的赞助。Donating grants no extra privileges: no bonus quota, no queue priority, no unlocked features. The pool balance and full ledger are visible in the console at any time. Not a charity appeal — it is users voluntarily paying a little extra, spent on keeping the price low for everyone.

我的数据会被用于训练吗?Is my data used for training?

不会。请求内容仅用于本次推理,不用于训练、不对外共享。企业版可提供数据不出境与完整私有化部署。No. Request content is used only to serve that request — never for training, never shared. An enterprise tier offers data residency and fully private deployment.

人人用得起的 AI,需要 80 个人一起开始AI everyone can afford — starting with 80 people

现在登记不收钱,也不绑卡。满 80 人我们开机,第一批邀请发给你。Signing up costs nothing and needs no card. At 80 people we launch, and the first invites go out.

登记意向Sign up 看计划价格See planned pricing