> ## Documentation Index
> Fetch the complete documentation index at: https://docs.tensormesh.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# CLI

> The Tensormesh Platform CLI for observing and operating LMCacheEngine fleets on Kubernetes.

The **Tensormesh Platform CLI** is the administrator command line for Tensormesh
Platform. It is built for the platform and SRE engineers who run the cache: it
reads LMCacheEngine resources and talks to the KV-cache server pod on each GPU
node.

The command itself is `tmo-cli`. It covers two surfaces:

* the Kubernetes CRD (`lmcacheengines.lmcache.lmcache.ai`), for engine status,
  conditions, and endpoints
* each server pod's HTTP control API, reached through the Kubernetes API-server
  proxy, for cache state, quotas, and diagnostics

<Note>
  This is the cluster-side administrator CLI. It is not the **Tensormesh CLI**
  (`tm`), which manages accounts, billing, and serverless inference against the
  hosted API.
</Note>

## Install

The CLI ships with the gated Tensormesh Platform distribution rather than through
public Homebrew or krew. It installs from the Tensormesh registry
using the same access token as the Platform chart and images. That token is scoped
to artifact pull and grants no access to Tensormesh source code; if you don't have
one, request it from the Tensormesh team. It is the same token used in the
[install guide](/v1.0.0/installation/helm).

The install script needs `curl`, `tar`, and either `shasum` or `sha256sum`.

### One-line install

```bash theme={null}
curl -sSfL https://operator.tensormesh.ai/install-cli.sh | TMO_TOKEN=<TOKEN_FROM_TENSORMESH> sh
```

The script detects your OS and architecture, pulls the matching binary from the Tensormesh registry,
checks the archive against the digest recorded in the OCI manifest, and installs to
`/usr/local/bin`, falling back to `~/.local/bin` when that isn't writable so no
`sudo` is required. To pin a version or change the target, pass flags after `-s --`:

```bash theme={null}
curl -sSfL https://operator.tensormesh.ai/install-cli.sh | sh -s -- \
  --token <TOKEN_FROM_TENSORMESH> --version <VERSION> --install-dir ~/bin
```

`--version` takes a release tag including the leading `v`, and defaults to `latest`.
Each flag also has an environment-variable form: `TMO_TOKEN`, `TMO_VERSION`,
`TMO_INSTALL_DIR`.

### Manual install (air-gapped or CI)

To pull the artifact yourself, authenticate to the Tensormesh registry and use [`oras`](https://oras.land):

```bash theme={null}
echo '<TOKEN_FROM_TENSORMESH>' | oras login artifacts.tensormesh.ai -u tensormesh --password-stdin

oras pull artifacts.tensormesh.ai/tensormesh-production/cli/tmo-cli:latest   # or a :<VERSION> tag
tar -xzf tmo-cli_*_linux_amd64.tar.gz tmo-cli                # match your OS/arch
install -m 0755 tmo-cli ~/.local/bin/tmo-cli
```

Archives are named `tmo-cli_<version>_<os>_<arch>.tar.gz`, where `<version>` drops
the leading `v` of the release tag. Builds are published for linux and darwin, on
amd64 and arm64.

### Verify

```bash theme={null}
tmo-cli version
tmo-cli diag preflight       # checks your RBAC against what the admin commands need
```

### Shell completion

```bash theme={null}
source <(tmo-cli completion zsh)     # this session; also bash, fish, powershell
```

To install permanently, run `tmo-cli completion <shell> --help` and follow the
instructions for your shell and platform.

## How it connects

`tmo-cli` uses your existing kubeconfig. If `kubectl` works, `tmo-cli` works.
Resolution order is `--kubeconfig`, then `$KUBECONFIG`, then the `tmo-cli` config
default, then `~/.kube/config`.

```bash theme={null}
tmo-cli config use-context prod-us-east
tmo-cli config set-default --output table
tmo-cli config view
```

Commands that talk to engine pods (`node`, `cache`, `quota`, `diag`) go through the
Kubernetes API-server proxy by default, so they need no network path beyond the one
`kubectl` already uses and RBAC on `pods/proxy` gates them. Pass `--direct` to reach
pod host IPs instead, which is useful when running inside the cluster.

## Flags

Four flags are accepted by every command:

| Flag                             | Meaning                                                                           |
| -------------------------------- | --------------------------------------------------------------------------------- |
| `--kubeconfig`                   | path to kubeconfig, overriding `$KUBECONFIG`                                      |
| `--context`                      | kubeconfig context to use                                                         |
| `-o, --output table\|json\|yaml` | output format; JSON and YAML carry a versioned envelope (`tmoOutputVersion: v1`)  |
| `--no-color`                     | disable ANSI color; also honors `NO_COLOR`, and auto-off when stdout is not a TTY |

The rest are command-scoped, and the scoping matters when you script:

| Flag                         | Where it applies                                                                                                                                                                   |
| ---------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `-n, --namespace`            | every command except `config` and `version` (default `default`)                                                                                                                    |
| `-e, --engine`               | commands that reach engine pods: `node list`, `cache *`, `quota *`, `diag health\|periodic\|threads\|bundle\|loglevel set`. `engine` commands take the name as an argument instead |
| `-A, --all-namespaces`       | `engine list`, plus the pod fan-out commands. Not available on `quota *` or `diag preflight`                                                                                       |
| `--node`                     | `cache status`, `cache clear`, `quota *`, `diag threads`, `diag loglevel set`, `engine logs`                                                                                       |
| `--direct`                   | all pod-reaching commands                                                                                                                                                          |
| `--workers`, `--pod-timeout` | fan-out commands only (defaults 8 and 5s). Not on `quota *`, which queries a single server                                                                                         |

## Exit codes

Stable and scriptable. Changing one is a breaking change.

| Code | Meaning                                                                  |
| ---- | ------------------------------------------------------------------------ |
| `0`  | success                                                                  |
| `1`  | generic or unexpected error                                              |
| `2`  | usage error: bad flag, unknown command, or a removed `engine` write verb |
| `3`  | not found (engine, server, cache\_salt)                                  |
| `4`  | validation or guardrail rejection                                        |
| `5`  | timed out, for example `engine wait`                                     |
| `6`  | confirmation declined, or `--yes` required and absent                    |

`diag health` overloads `1`: it exits `0` when every pod in scope is reachable and
healthy, and `1` when any pod is unhealthy or unreachable. That makes it usable
directly as a health gate in CI.

## Common tasks

### See the fleet

```bash theme={null}
tmo-cli engine list -A                # all engines, all namespaces
tmo-cli engine list -A -w             # watch for changes (table output only)
tmo-cli engine describe my-cache      # spec, status, conditions, endpoints, events
```

`engine describe` looks up recent events by default; pass `--no-events` to skip that
call on a busy cluster.

### Deploy or change engines (owned by Helm)

<Note>
  Fleet membership and engine spec belong to the Helm release. `tmo-cli` cannot
  create, apply, edit, or delete engines; those verbs exit `2` and point you at
  Helm. To add, change, or remove an engine, edit the chart values and run
  `helm upgrade`.
</Note>

The CLI's job in a rollout is the readiness gate afterwards:

```bash theme={null}
helm upgrade tensormesh-operator oci://artifacts.tensormesh.ai/tensormesh-production/charts/tensormesh-operator -f values.yaml
tmo-cli engine wait my-cache --for=AllInstancesReady --timeout=5m
tmo-cli engine describe my-cache
```

`--for` accepts `Available`, `AllInstancesReady`, or `ConfigValid`, and defaults to
`AllInstancesReady`. `--timeout` defaults to `5m`. A timeout exits `5`.

### Inspect servers and cache state

```bash theme={null}
tmo-cli node list -e my-cache                       # per-node servers with live health
tmo-cli cache status -e my-cache                    # aggregated /status, per-node and TOTAL
tmo-cli cache status -e my-cache --node gpu-node-07 # raw /status passthrough for one server
tmo-cli engine logs my-cache --node gpu-node-07 -f  # --all multiplexes every pod
```

`engine logs` shows the last 100 lines per pod; change that with `--tail` (`-1` for
the whole log).

### Operate (destructive, gated)

```bash theme={null}
tmo-cli cache clear -e my-cache --node gpu-node-07  # one server; prompts
tmo-cli cache clear -e my-cache --all --yes         # whole fleet, no prompt

tmo-cli quota list -e my-cache                      # per-cache_salt L2 budgets
tmo-cli quota set acme-prod --limit 48 -e my-cache  # --limit is in GB
tmo-cli quota rm bench-harness -e my-cache --yes    # data evicts next cycle
```

<Warning>
  `cache clear` bypasses read and write locks and drops L1 contents. It refuses
  while a server still has active sessions; add `--force` to clear anyway. It and
  `quota rm` print a blast-radius preamble and prompt for confirmation, which
  `-y/--yes` skips. In a non-TTY without `--yes` they exit `6` rather than guess.
</Warning>

<Note>
  Quotas are enforced only on engines running the `IsolatedLRU` eviction policy.
  Against any other policy the CLI accepts the command and warns, but the budget
  has no effect. Quota commands query one server (the first ready endpoint, or
  `--node`), since every server of an engine shares the same quota config.
</Note>

### Diagnose and file a ticket

```bash theme={null}
tmo-cli diag health -e my-cache       # healthcheck + periodic-thread sweep
tmo-cli diag periodic -e my-cache     # per-server periodic-thread detail
tmo-cli diag threads --node gpu-node-07 -e my-cache
tmo-cli diag loglevel set lmcache DEBUG --node gpu-node-07 -e my-cache
tmo-cli diag bundle -e my-cache --output-file support.tgz
```

`diag loglevel set` takes a logger and one of `DEBUG`, `INFO`, `WARNING`, `ERROR`,
and applies to one server with `--node` or the whole engine with `--all`.

`diag bundle` collects the CR, recent events, and each server's `/conf`, `/status`,
`/threads`, and `/periodic-threads-health` into a gzip tarball, writing to
`tmo-bundle-<engine>-<timestamp>.tgz` unless you pass `--output-file`. Secret-looking
values are redacted, and `/env` is excluded entirely unless you pass `--include-env`
(it is still redacted). Attach the tarball to a support ticket.

## Output and scripting

Table output is for reading. `-o json` and `-o yaml` are the automation contract,
wrapped in an envelope carrying `tmoOutputVersion: v1`:

```bash theme={null}
tmo-cli node list -e my-cache -o json | jq '.items[] | select(.healthy != true)'
tmo-cli engine get my-cache -o json | jq '.item.status.readyInstances'
```

<Warning>
  Use `.healthy != true`, not `.healthy == false`. A pod the CLI could not reach has
  no `healthy` field at all, only `error`, so an equality test silently drops the
  unreachable pods you are looking for.
</Warning>

Single-item commands render under `.item`, lists under `.items`. Color is disabled
automatically in pipes and non-terminals, and never appears in JSON or YAML.

## Troubleshooting

| Symptom                                                 | Cause and fix                                                                              |
| ------------------------------------------------------- | ------------------------------------------------------------------------------------------ |
| `... forbidden ... ask for the 'tensormesh-admin' role` | your RBAC lacks a verb the command needs; run `tmo-cli diag preflight` to see which        |
| `node list` or `cache status` returns nothing           | the engine has no ready endpoints; check that a GPU node is up and the engine is `Running` |
| `unreachable` in a health table                         | that server's HTTP API did not answer; check `tmo-cli engine describe` and pod status      |
| `unknown shorthand flag: 'A'`                           | that command has no `-A`; `quota` and `diag preflight` are namespace-scoped                |
| `engine wait` exits `5`                                 | the condition did not become true within `--timeout`                                       |
| `cache clear` refuses                                   | the server has active sessions; re-run with `--force` once you are sure                    |
| fan-out commands hang on a large fleet                  | raise `--workers` or `--pod-timeout`; partial results still render                         |

<Note>
  `tmo-cli` does not phone home. There is no active telemetry in the shipped build.
</Note>
