Source record: Moonshot AI’s Kimi K3 official documentation was checked on 22 July 2026. At that check, the page documented API access and said full model weights would be released by 27 July. Open the official page.
Time boundary: This issue preserves what the page established when checked on 22 July 2026. It is not a claim about availability after that date. Later official records report full weights on 27 July; Issue 09 reads that later state.

The 22 July record supported one direct conclusion: teams could evaluate K3 through Moonshot’s hosted API. It did not then support a self-hosting conclusion. Hosted access, downloadable artifacts, licensing terms, and reproducible deployment were separate operational states.

One launch, two operational states

In the official record checked on 22 July, K3 is documented through a live API quickstart. Moonshot’s product claims also describe a 2.8-trillion-parameter architecture, native visual understanding, and a one-million-token context window. source ↓

That was enough to put response behavior, tool use, latency, pricing, and integration friction into an API evaluation queue. It did not establish the artifacts, license, memory profile, runtime support, or economics of a self-hosted deployment.

API access is operational, but it is not frictionless

The official page says K3 unlocks after a successful top-up of at least one US dollar, with account tier determining concurrency and request and token limits. Its quickstart uses the model ID kimi-k3 through Moonshot’s OpenAI-compatible API. source ↓

This is enough to establish a real evaluation path. It also means “available” should not be read as anonymous, unrestricted, or infrastructure-independent access. Account state, tier limits, pricing, and the hosted service remain part of the operational product.

At the 22 July checkpoint, the page said the full model weights would be released by 27 July 2026. At that moment, “will be released” described a future state; it did not establish that the artifacts and license were already available. source ↓

Scale is a company claim before it is independent evidence

The page describes K3 as a 2.8-trillion-parameter mixture-of-experts model built on Kimi Delta Attention and Attention Residuals, with native visual understanding and a one-million-token context window. It says 16 of 896 experts are active and claims roughly 2.5 times the scaling efficiency of K2.

Those details are useful for understanding Moonshot’s design thesis. They do not yet answer the production questions that matter most: sustained throughput, memory placement, interconnect demand, quantization tolerance, long-context cost, tool-use reliability, or the gap between official service behavior and third-party deployments. Architecture description, company benchmark, and independent reproduction should remain three separate evidence layers.

What application teams can measure now

A hosted endpoint creates a legitimate test surface. Teams can build a bounded evaluation that separates model quality from provider operations:

  1. 01
    Task behavior.

    Test coding, visual reasoning, long-context retrieval, tool use, and multi-turn consistency against a stable task set.

  2. 02
    Service behavior.

    Record latency, streaming behavior, rate limits, failure recovery, caching, and any difference between product and API modes.

  3. 03
    Commercial fit.

    Measure token consumption, retry cost, observability, data-handling requirements, and integration work rather than relying on headline model scale.

  4. 04
    Exit conditions.

    Define what result would justify continued API use, waiting for weights, or abandoning the evaluation.

A million-token window changes the test, not the standard

The official documentation advertises a one-million-token context window and automatic prefix caching. It also explains that cache reuse depends on prompt structure and that long prefixes must remain unchanged for later requests to attempt a hit.

That creates a larger evaluation surface rather than a shortcut to a verdict. Teams should test retrieval accuracy at different positions, instruction persistence, tool-call stability, latency, cache behavior, and cost across realistic document sets. The ability to accept a long prompt does not establish that every part of the prompt influences the answer reliably, or that a million-token workflow is economical in production.

Long context also changes failure handling. A useful integration needs observable truncation, retry, and caching behavior so that an apparently successful response does not conceal a dropped document, stale prefix, or shifted instruction hierarchy.

What self-hosting teams still cannot decide

At the 22 July checkpoint, the official page did not provide a complete deployable artifact set or enough evidence to choose hardware and serving architecture. A serious self-hosting decision still needed the exact weight files, license, precision options, tokenizer and configuration, memory requirements, supported runtimes, reference inference code, and evidence that third-party serving reproduced expected behavior.

The model’s scale makes this distinction unusually important. Even if the announced weights arrive on time, availability of a download will not by itself establish economical deployment. The ecosystem must show whether the model can be partitioned, quantized, cached, routed, and monitored in a way that matches real workloads.

Why the distinction matters

At that checkpoint, product access could support experiments and procurement comparison. A promised weight release could not support hardware sizing, serving-stack selection, security review, or production-readiness claims until the artifacts and terms existed and could be examined.

The editorial discipline is simple: do not let one word—“released”—collapse a sequence of different states. An announcement can be credible, an API can be useful, and a self-hosting claim can still be premature at the same time.

The release should be judged as a sequence

At 22 JulyHosted product and API

Usable for controlled evaluation under Moonshot’s service contract.

Announced nextFull model weights by 27 July

A dated company commitment, still future-facing at the 22 July check.

Decision gateLicense, artifacts, runtimes, reproduction

The evidence needed to turn downloadable files into a self-hosting decision.

A missed date would matter, but an on-time file upload would not finish the story. The decisive transition is from a promised artifact to a usable, licensed, documented, and independently examined deployment surface.

Limit

The official page was checked on 22 July 2026. This issue preserves that checkpoint; later events do not retroactively change what the page established that day. Company architecture and performance statements were not independently verified.

What would change this assessment

A newly checked official release containing the exact artifacts and license; credible memory and inference requirements; supported runtimes; third-party hosting; independent reproduction; or a changed release date.

Time-bounded editorial reading of a checked company source link. Architecture and release statements remain company claims until independently verified.