现在能用吗?Can I use it today?
还不能。机器一台没买,现在只登记意向,不收任何人的钱。满 80 个真的想用的人,我们就去买机器、开机,第一批邀请发给登记的人。不满 80,说明这件事现在还不该做,我们继续等。Not yet. No hardware has been bought and we take no money at this stage — we are only collecting interest. At 80 people who actually want this, we buy the hardware and turn it on, and the first invites go to the list. Below 80, we keep waiting.
现在能看到的是落地页和登记入口。服务本身会跑在开源 Qwen3.8-27B 上(Apache 2.0,你也可以自己部署一份),但我们不会等到"什么都完美"才开——先把 80 个真的想用的人凑齐。What exists today is this page and the sign-up. The service itself will run open-weight Qwen3.8-27B (Apache 2.0 — you could self-host it today), but we are not waiting for perfect before starting: first, 80 people who actually want it.
为什么是 80?因为这个规模能让我们把机器开起来、并撑住前三个月的运营。达不到就是白烧钱,我不想那样开始。Why 80? Because that is the scale at which the machine can be switched on and kept running for the first three months. Below it, we would just be burning money — and I would rather not start that way.
为什么是 262K?别人都 1M 了,够用吗?Why 262K when others offer 1M? Is that enough?
够用,而且是刻意选的。262K 约等于 20 万字,一次装得下一整份合同、一本书、一整个中型代码库。我们没去堆更大的窗口,原因有两个:① 便宜——窗口越大,同一张卡能同时服务的人越少,摊到你头上的成本越高;② 准——上下文越长,注意力越分散,中间部分容易被忽略。我们宁可把这 262K 里的每一段都答得准、算得便宜,也不给一个大部分人用不满的大数字。It is enough, and the size is deliberate. 262K is roughly 200k Chinese characters — a full contract, a book, a mid-sized codebase in one request. We did not chase a bigger window, for two reasons: (1) cost — the larger the window, the fewer people one card can serve at once, and the more it costs you; (2) accuracy — the longer the context, the more diffuse the attention, and the middle tends to get ignored. We would rather have every token inside 262K answered accurately and cheaply than ship a big number most people never fill.
真要读更长的东西也有解:前缀缓存让同一份长文档的反复追问只算一次钱;再长的资料可以切成多轮,或者走异步批处理。我们实测过 262,144 token 的整篇文档检索(信息埋在全文 33%、66%、90% 三个位置,全部按序取回),窗口内的可靠性是验证过的。For genuinely longer inputs there are answers: prefix caching bills repeated questions over the same long document once; anything longer can be split across turns or pushed through async batch. We did verify retrieval across a full 262,144-token document — information planted at 33%, 66% and 90% came back in order — so reliability inside the window is tested, not assumed.
用的是什么模型?是完整的 Qwen3.8-27B 吗?Which model do you serve — is it the full Qwen3.8-27B?
对外就叫 Qwen3.8-27B:开源、Apache 2.0、27B 稠密参数,你可以自己下载权重、自己部署。我们实际跑的是它的省显存推理版(压缩了权重表示,官方基准保留全精度模型 98.2% 的综合能力),换来的是单卡能同时服务更多人、价格压到约 1/10。所以如果你要的是"绝对满血",自托管原版即可;如果你要的是"便宜、不设墙、够用",用我们的。We call it Qwen3.8-27B: open weights, Apache 2.0, 27B dense, downloadable and self-hostable. What we actually serve is a memory-lean inference build of it (a compressed weight representation that keeps 98.2% of the full-precision model's aggregate benchmark score), which is what lets one card serve more people at roughly a tenth of the cost. Want the absolute full-fat version? Self-host the original. Want cheap, wall-free and good enough? Use ours.
为什么能比 Opus 4.6 便宜 100 倍?质量差多少?How can this be 100× cheaper than Opus 4.6? How much worse is it?
便宜来自三件事:这个 27B 模型的显存占用只有前沿模型的零头(一张消费级显卡就能跑)、我们把并发利用率拉满、不养任何品牌溢价。质量上,编码、电脑操作这些"手上有活"的任务它已经反超 Opus 4.6,知识推理和最难的那批题仍然落后(对比表在上面)。但要说清楚一点:连跑同一个开源模型的第三方 API,单价都比我们贵 5 倍左右——便宜不是因为模型小,是因为我们没有把中间利润堆上去。Three things: this 27B model's memory footprint is a fraction of a frontier model's (it runs on one consumer GPU), we keep concurrency utilization high, and we carry no brand premium. On quality it already leads Opus 4.6 on hands-on work like coding and computer use, and still trails on knowledge and the hardest reasoning (table above). One thing worth stating plainly: even third-party APIs serving the very same open model charge about 5× our price — the price is low because we did not stack margin on top, not because the model is small.
缓存命中怎么判定、怎么计费?How are cache hits determined and billed?
服务端按前缀精确匹配复用 KV 缓存。响应 usage 中区分 prompt_cache_hit_tokens 与 prompt_cache_miss_tokens,分别按 ¥0.01/M 与 ¥0.20/M 计费。固定系统提示词、同一份长文档的多轮追问、Agent 的重复上下文都能命中。We do exact prefix matching on the server and reuse the KV cache. The usage block separates prompt_cache_hit_tokens from prompt_cache_miss_tokens, billed at $0.002/M and $0.20/M respectively. Fixed system prompts, follow-up questions on the same long document and repeated agent context all hit.
容量用完了会怎样?扩容要多久?What happens when capacity runs out, and how long does scaling take?
我们会按需增加机器,但新机器需要「开机 → 拉取模型 → 预热」,约 5–10 分钟。满载时新请求会进入队列并按提交顺序处理,不会丢失。如果你有可预期的峰值,提前告知我们会事先扩容;紧急情况也可以通过控制台申请优先通道。We add machines on demand, but a new one has to boot, pull the model and warm up — roughly 5–10 minutes. When saturated, new requests queue in submission order and are never dropped. Tell us about expected peaks and we scale ahead of time; you can also request a priority lane from the console.
套餐额度怎么算?输入 token 也算吗?How is plan quota counted — do input tokens count too?
算,但按成本折算成"单位"。1 个输出 token = 1 单位;输入 token(≤64K)只算 0.16 单位;长文档输入(>64K)算 0.35 单位;缓存命中只要 0.02 单位。所以把系统提示词和长文档交给缓存,额度几乎不掉。完整换算表在上面的"按量付费"一节。Yes, but converted into units by cost. 1 output token = 1 unit; input tokens under 64K cost 0.16 units; long-document input above 64K costs 0.35 units; cache hits cost just 0.02 units. Reuse your system prompt and long documents through the cache and your quota barely moves. Full table is in the pay-as-you-go section above.
可以随时退订吗?用不完的额度能留到下个月吗?Can I cancel? Does unused quota roll over?
随时可退,按未使用天数比例退款。当月未用完的额度可结转一个月(第二个月底清零)。额度用完后不会停服,会自动转按量付费,你也可以在控制台设一个月度封顶自动停。Cancel anytime with a pro-rata refund for unused days. Unused quota rolls over for one month and expires at the end of the following month. When the quota runs out we do not stop you — you roll onto pay-as-you-go, and you can set a monthly spend cap in the console.
和 Claude Code 那种订阅有什么区别?How is this different from a Claude Code style subscription?
三点:没有 5 小时滚动窗口,没有每周限额,价格便宜一个数量级。传统订阅撞到周限后只能等满 7 天、升级或转 API;我们额度用完后自动转按量付费,不降速、不断服。而且我们不绑定客户端——Claude Code、Cline、Cursor 改个 base_url 就能接到我们这里。Three things: no 5-hour rolling window, no weekly cap, and an order of magnitude cheaper. On a classic subscription, hitting the weekly cap means waiting seven days, upgrading, or switching to API billing; with us the quota simply rolls into pay-as-you-go with no throttle and no cut-off. And we are not tied to a client — point Claude Code, Cline or Cursor at our base_url and you are done.
捐款真的会让我的额度变多吗?Will donating actually raise my quota?
会。捐款的 95% 进入额度池,用于增加算力和统一上调三档额度——上调对全体用户生效,包括没有捐款的用户,也包括你。剩下的 5% 用于支付通道手续费与账务成本。每月 1 号公示上月收到的金额、进入额度池的金额,以及额度实际提升了多少。额度池按累积制计算,累计达标后 30 天内完成上调。Yes. 95% of every donation goes into the quota pool, which is spent on adding capacity and raising all three plan quotas — the increase applies to every user, including those who never donated, and including you. The remaining 5% covers payment processing and accounting. On the 1st of each month we publish the total received, the amount that went into the pool, and how much the quotas actually rose. The pool is cumulative: increases land within 30 days of crossing a threshold.
捐款不产生任何额外权益:不会给你更多额度、不会插队、不会解锁功能。额度池的余额和流水在控制台随时可查。这不是慈善募捐,是用户自愿多付、用来把价格压给所有人的赞助。Donating grants no extra privileges: no bonus quota, no queue priority, no unlocked features. The pool balance and full ledger are visible in the console at any time. Not a charity appeal — it is users voluntarily paying a little extra, spent on keeping the price low for everyone.
我的数据会被用于训练吗?Is my data used for training?
不会。请求内容仅用于本次推理,不用于训练、不对外共享。企业版可提供数据不出境与完整私有化部署。No. Request content is used only to serve that request — never for training, never shared. An enterprise tier offers data residency and fully private deployment.