Kimi K3 Now Available via Telnyx Inference API
Posted by fionaattelnyx 6 days ago
Moonshot AI released open weights for Kimi K3 today and it's live on Telnyx Inference. The architecture and Moonshot's own benchmarks are in their technical blog. This post is about running it on Telnyx.
What we are adding: K3 is now available via the Telnyx Inference API, hosted on GPUs that we own and operate.
This matters due to the size of this model. A 2.8T model needs dedicated infra to serve well. Because we own and operate the GPUs, we can contorl throughput. with no inter-provider hops and no cloud tenant, latency is minimized. We have GPUs in each of the US, EU, APAC, and MENA. Inference runs in the region you pick, with zero data retention. We do not store prompts or completions after the response returns. Because we own the infra, the per-token price reflects the cost of running the model, not the cost of renting someone else's plus their margin.
We have not benchmarked K3 ourselves yet. Moonshot's own numbers and early third-party evaluations put it at frontier level for coding and agentic work, trailing only Claude Fable 5 and GPT 5.6 Sol on aggregate. Full breakdown in their blog.
Pricing on Telnyx: $2.70/1M input tokens, $13.50/1M output tokens, $0.27/1M cached input tokens. Prompt caching enabled by default. Served via an OpenAI-compatible endpoint so you can test easily.
Model ID: [MODEL_ID] API: https://api.telnyx.com/v2/ai/chat/completions Docs: https://developers.telnyx.com/docs/inference Technical blog (Moonshot): https://www.kimi.com/blog/kimi-k3
Comments
Comment by apexalpha 6 days ago
Why is this provider specifically on the front page?
Also why is this not available over openrouter?
Comment by contactdq 6 days ago
We've been trying to get on OpenRouter. Just reached out to their CEO. Would be great if people could request it as well.
We're unique in that we have B300 in USA, EU, UAE, and AUS.
Also, building out other agentic primitives like stateful actors (in beta, but will be rolled out globally in the coming weeks):
https://developers.telnyx.com/docs/edge-compute/stateful-act...
Also, if you're doing things agentically, check out AI repo:
Comment by ivanvanderbyl 5 days ago
Comment by luizanao 5 days ago
Comment by blackoil 6 days ago
Comment by heisgone 6 days ago
Comment by contactdq 6 days ago
Comment by ljlolel 6 days ago
Comment by dan_gee 6 days ago
Comment by accountforih 6 days ago
Comment by Mossy9 6 days ago
€2.693/M input €13.464/M output Surprisingly, cache is not mentioned
Upd: Tensorix joined the fray, with the same prices, with cache at €0.673/M read
Comment by zius 6 days ago
> Hi there! Quick note — I'm actually Claude, made by Anthropic, not Kimi. But no worries!
and
> Just a quick heads-up — I'm Claude, made by Anthropic, not Kimi! Kimi is a different AI assistant (made by Moonshot AI), so it looks like there might be a little mix-up.
Comment by tomashubelbauer 6 days ago
Comment by victorbjorklund 6 days ago
Comment by scotty79 6 days ago
Comment by embedding-shape 6 days ago
This approach is basically how all "models know who they are" (they don't actually typically "know" that at all), it's just a system prompt instruction in the platform you use.
Comment by brookst 6 days ago
Comment by dannyw 5 days ago
Comment by scotty79 5 days ago
Comment by XenophileJKO 6 days ago
Comment by scotty79 6 days ago
I think Qwen teaches its models that they are Qwen. Most others don't bother.
Comment by irthomasthomas 6 days ago
Comment by cleaning 6 days ago
Comment by dan_gee 6 days ago
Comment by FooBarWidget 6 days ago
Comment by nojs 6 days ago
Comment by w4yai 6 days ago
Comment by Razengan 6 days ago
Comment by jingpostmedia 6 days ago
Comment by croes 6 days ago
Comment by TZubiri 6 days ago
Comment by victorbjorklund 6 days ago
Comment by TZubiri 5 days ago
But the answer is no, Claude didn't train itself on ChatGPT, but Kimi K3 did train itself on Claude.
There has been no accusation that I am aware of by OpenAI against Anthropic, (which are in the same jurisdiction, so it would amount to legal action). On the other hand there have been accusations with extensive detail by Anthropic on how chinese models are attacking Anthropic to reverse engineer and copy their product.
There's nuance if you care to see it, but maybe it's easier to pretend that everyone is ripping each other off so that you can consume a ripoff without seeing yourself as at fault.
Comment by victorbjorklund 5 days ago
Comment by TZubiri 4 days ago
But the reason why Kimi thinks it is Claude is different, it's because chinese companies make thousands of fake accounts, fire prompts at Claude and train their models to imitate Claude.
Comment by croes 6 days ago
Comment by lukewarm707 6 days ago
(joke and walter mitty, i know the government reads my messages)
Comment by Grimblewald 6 days ago
Comment by theredsix 6 days ago
Comment by norbert515 6 days ago
Comment by mesmertech 6 days ago
Or does openrouter have like a specific contract you have to sign with them and requirements or smth? https://openrouter.ai/moonshotai/kimi-k3#providers
Comment by ljlolel 6 days ago
Comment by addandsubtract 6 days ago
Comment by hmokiguess 6 days ago
Comment by DaSHacka 6 days ago
Comment by hmokiguess 6 days ago
Comment by hmokiguess 6 days ago
Comment by fionaattelnyx 5 days ago
Comment by TZubiri 6 days ago
But I don't feel like providing inference is a professional move, it feels like out of scope for a telephony IaaS company. Feels like a FOMO moment where a reputation of years is crashed in a couple of weekends of being drawn into a fad.
And the fact that it's a chinese model doesn't quite help? I guess it's on brand with the 'cheap' pay as you go brand telnyx might already be associated to.
But more so it reads like Telnyx is trying to 'jump' into the trend of the 'open weights' discussion to compete with closed source incumbents. But we are at the tail end of the boom, anti ai sentiment is ever growing, customers now despise AI, especially in support channels, which is presumably the hook that Telnyx would have into 'AI'(LLMs). At this stage any company or individual that tries to join into the buzzword fueled cycle will pay the full fixed cost reputational price, but only reap the leftover hay from when the sun shone.
AI(LLM) on support channels is essentially a decapitalization of a company/brand, the company has a reputation that customers value, and might be worth good money in the market, and by implementing AI (LLMs) on support, a lot of costs can be cut, while the brand loses value, not sustainable. And by Telnyx (or any B2B company)asking their clients to participate in this decapitalization move, they essentially gamble their reputation as well, albeit with better odds as shovel sellers.
Comment by contactdq 6 days ago
We're moving past "telephony" and focusing on building more agentic primitives at the telecom edge.
This isn't a recent jump. We've been on this path for some time.
Disagree with you on the customer support side and re: decapitalization in general. People will want to talk to competent bots. We have people literally prompting our AI agents (which are capable of doing troubleshooting) in our shared Slack channels. You simply cannot beat the speed with which they can provide a quality response.
We remain committed to delivering the highest quality product at the lowest possible price point - across all of our product lines.
Comment by jbstack 6 days ago
Comment by inigyou 6 days ago
Comment by trollbridge 6 days ago
Comment by maelito 6 days ago
Comment by mkl 6 days ago
The top text says MENA as well, but the regions page doesn't mention it.
Comment by maelito 5 days ago
Comment by smallerize 6 days ago
Comment by htrp 6 days ago
Looks like they are first vendor to undercut in price.
Comment by bedros 6 days ago
Comment by morpheuskafka 6 days ago
Comment by contactdq 6 days ago
Comment by trollbridge 6 days ago
Comment by tokai 6 days ago
Comment by rvz 6 days ago
Comment by inigyou 6 days ago
Comment by philjohn 6 days ago
Comment by neuroticnews25 6 days ago
Comment by LoganDark 6 days ago
OVH same thing. Tried to buy a VPS from them some years back and they said no VPS unless I provided ID. Would not refund me. Tried to dispute but my bank just gave me a credit instead.
Comment by contactdq 6 days ago
Comment by nujabe 6 days ago
Comment by oxidant 6 days ago
Comment by TZubiri 6 days ago
Such use would be incompatible because it would lure in fraud, which would ruin the reputation of shared comms resources like ip blocks, phone number blocks, etc..
If you are doing real business, you are providing KYC as a daily matter in procurement, thereby protecting consumers from Sybil scum.
Comment by inigyou 6 days ago
No? Firstly you can't do a Sybil attack when each identity costs actual resources, and secondly you can still have lots of phone numbers, they just know who each one belongs to.
Comment by illliillll 6 days ago
Comment by tomr75 6 days ago
Comment by jakswa 6 days ago
Comment by NitpickLawyer 6 days ago
Comment by trollbridge 6 days ago
Comment by BVHauge 4 days ago
Comment by madhu_ghalame 6 days ago
Comment by fionaattelnyx 5 days ago
Comment by nttylock 6 days ago
Comment by marsven_422 6 days ago
Comment by hiherer 6 days ago
Comment by buffer_overlord 6 days ago
Comment by teravor 6 days ago
> Because we own the infra, the per-token price reflects the cost of running the model, not the cost of renting someone else's plus their margin.
is not compatible with > Pricing on Telnyx: $2.70/1M input tokens, $13.50/1M output tokens, $0.27/1M cached input tokens.
since you asserted something false and bizarre, how about telling us what is your markup?Comment by alexeldeib 6 days ago
Comment by gruez 6 days ago
Comment by rubslopes 6 days ago
Comment by teravor 6 days ago
We estimate that the true blended price per million tokens for running Opus 4.7 on agentic tasks at $0.99 despite the sticker price being $5/$25 per MTok.
https://newsletter.semianalysis.com/p/ai-value-capture-the-s...according to Semianalysis those prices would be far above actual costs.
Comment by petu 6 days ago
> Agentic workloads have extremely high input-to-output ratios (our Claude Code usage has a ratio of about 300:1) and high cache hit rates (90%+). Because cached input tokens only cost $0.50/MTok, most of the tokens end up in the cheapest tier.
90% cache hit input blend: 0.9 * $0.5 + 0.1 * $5 = $0.95 per MTok.
300:1 input/output blend: (300 * $0.95 + $25) / 301 = $1.03 per MTok.
They don't say exact cache hit rate they calculated for ("90%+"), so close enough.
Comment by Grimblewald 6 days ago
I largely agree things are overpriced, I just don't think that article is the right basis for saying it about this one.
Comment by tpm 6 days ago
Comment by FergusArgyll 6 days ago
Not the **literal** cost