Arative Token Stats

relay loading…
Intelligence · Speed · Cost per Task

Live charts from Artificial Analysis. Click a bar to open that model.

Today's Winner Speed / Intelligence Balance: Muse Spark 1.3
Which model to pick? Guidelines
Integrate your system with Typesafe Jev as the automatic router decision engine based on the following guidelines:
Typesafe Jev Router Decision Guidelines
Automatic Policy
Maximum Quality & Precision Private
Route sensitive or BancoliKnowledge-adjacent tasks requiring deepest reasoning and highest accuracy.
Claude Opus 5.5
Default Speed & High Quality
Default driver for daily development, fast execution, 1M context, and standard coding workflows.
Muse Spark 1.3 Contributor
Speed & Quality Private
Route fast 1M context tasks when zero-training privacy and strict confidentiality are required.
Muse Spark 1.3 (Cursor Subscription)
Rapid Response & Live Search Private
Route for fast turnarounds, live web research, and rapid private queries.
Grok 4.7 Fast
Complex Problem Solving
Route for multi-step logic, code review audits, and high-complexity analytical tasks.
GPT Sol 6.1 & Astra
High volume work at extremely low cost
Route bulk, repetitive, high-throughput jobs where price per task matters more than peak quality.
GPT-6 Luna
What needs “no training”:

Very little. We only require private models when interacting directly with BancoliKnowledge. Everything else is fair game. In fact, for certain workflows, we explicitly want AI models to train on and understand Bancoli’s offerings. This includes most of the core application itself and all public-facing materials.

When developing or building directly from BancoliKnowledge, always ensure the active model is set to one of the private options.
Tribe Seats for GPT Models
Tribe Seats
Dedicated Seats
Hosted by NVIDIA in the US 🇺🇸

Kimi K3

live
Model moonshotai/kimi-k3Context 1M

DeepSeek V4 Flash

live
Model deepseek-ai/deepseek-v4-flash-0731Context 1M

Nemotron 3.5 Lightning 30B

live
Model nvidia/nemotron-3.5-lightning-30b-a3bContext 1M
Setup Steps — NVIDIA NIM + OpenCode (full guide from tools/nvidia-nim/README.md)

Source: RuView/tools/nvidia-nim/README.md — verified live 2026-08-31. (open raw markdown)

# NVIDIA NIM + OpenCode Setup Guide

Hook your free NVIDIA Build key into OpenCode without the whack-a-mole.
Everything below was verified live on 2026-08-31 against `integrate.api.nvidia.com`.

---

## 1. Get the key

1. Go to https://build.nvidia.com, sign in.
2. Any model page → **Get API Key** → **Generate Key**.
3. Copy the key. It starts with `nvapi-`.

## 2. Find models your org can actually call

`GET /v1/models` lists ~83 models. **The catalog lies.** Being listed does not
mean your account may invoke it. Models fall into 4 buckets:

| Response | Meaning | Fix |
|---|---|---|
| `200` | Works | None |
| `404 Function '<uuid>' Not found for account` | Missing **Public API Endpoints** entitlement | Forum request (see §5) |
| `410 Gone` | Model hit end-of-life | Pick another model |
| Hangs / read timeout | Overloaded deployment | Retry later, or use smaller variant |

Probe them with the included script:

```bash
NVIDIA_API_KEY=nvapi-xxx node tools/nvidia-nim/diagnose.mjs
# or python
NVIDia_API_KEY=nvapi-xxx python scripts/nvidia_nim_diag.py
```

Working IDs as of 2026-08-31 (personal/free org):

- `moonshotai/kimi-k3` — this **is** Kimi 3 (not `kimi-k2.6`)
- `deepseek-ai/deepseek-v4-flash-0731`
- `nvidia/nemotron-3-nano-30b-a3b`
- `nvidia/nemotron-3-super-120b-a12b`
- `nvidia/nemotron-3.5-lightning-30b-a3b`
- `nvidia/nemotron-3-nano-omni-30b-a3b-reasoning`

Avoid: `deepseek-v4-pro-0813` (hangs), `meta/llama-3.1-*` (410 EOL),
`kimi-k2.6`, `nemotron-4-340b-instruct`, `nemotron-3-ultra-550b-a55b` (404 entitlement).

## 3. Configure OpenCode

`~/.config/opencode/opencode.json` — three traps matter:

```jsonc
{
  "enabled_providers": ["nvidia"],
  "provider": {
    "nvidia": {
      "name": "NVIDIA NIM",
      // TRAP 1: must be openai-compatible, NOT @ai-sdk/openai.
      // @ai-sdk/openai sends POST /v1/responses (OpenAI Responses API).
      // NIM has no /responses route -> "404 page not found" on EVERY model.
      "npm": "@ai-sdk/openai-compatible",
      "options": {
        // TRAP 2: include exactly one /v1, no trailing slash.
        "baseURL": "https://integrate.api.nvidia.com/v1"
      },
      "models": {
        "moonshotai/kimi-k3": {
          "name": "Kimi K3",
          "limit": { "context": 1048576, "output": 131072 }
          // TRAP 3: do NOT set "reasoning": true here.
          // reasoning:true makes OpenCode send `interleaved` /
          // reasoning params -> NIM answers 400:
          // "Validation: Unsupported parameter(s): `interleaved`"
        },
        "deepseek-ai/deepseek-v4-flash-0731": {
          "name": "DeepSeek V4 Flash",
          "limit": { "context": 1048576, "output": 131072 }
        },
        "nvidia/nemotron-3.5-lightning-30b-a3b": {
          "name": "Nemotron 3.5 Lightning",
          "limit": { "context": 1048576, "output": 131072 }
        }
      }
    }
  }
}
```

Store the key in `~/.config/opencode/auth.json` (or use `/connect` → NVIDIA in the TUI):

```json
{ "nvidia": { "api_key": "nvapi-xxxxxxxx" } }
```

Restart OpenCode fully (CMD+Q, reopen) — provider npm packages load at startup.

That's the whole recipe:

- `@ai-sdk/openai-compatible` → requests hit `/v1/chat/completions` (not `/v1/responses`)
- `baseURL` with a single `/v1` → no `/v1/v1` double-path 404
- no `reasoning: true` on NIM models → no 400 `interleaved`

## 4. Verify

```bash
NVIDIA_API_KEY=nvapi-xxx node tools/nvidia-nim/diagnose.mjs
```

You want: control model `200 OK`. Then in OpenCode's model picker choose
`nvidia/...` and send a message. Errors read:

- `404 page not found` (plain text, no JSON) → wrong npm package or wrong URL (traps 1/2)
- `400 Unsupported parameter(s): interleaved` → remove `reasoning: true` (trap 3)
- `404 Function ... Not found for account` → entitlement, see §5

## 6. Files in this repo

```
tools/nvidia-nim/nim-client.mjs   # baseURL normalizer + minimal client (JS)
tools/nvidia-nim/diagnose.mjs     # Node diagnostic
scripts/nvidia_nim_diag.py        # Python diagnostic
example.env                       # NVIDIA_API_KEY / NIM_BASE_URL template
```

The client guards the wrong-warp permanently:
`normalizeNimBaseUrl()` strips any trailing `/v1` so SDK auto-append can never
produce `/v1/v1/chat/completions`.
Quick checklist:
  • Get key at build.nvidia.com → Get API Key (starts nvapi-)
  • 4 buckets: 200 works / 404 entitlement forum / 410 EOL pick another / timeout retry
  • 3 traps: "npm": "@ai-sdk/openai-compatible" not @ai-sdk/openai · baseURL one /v1 · No "reasoning": true
  • Key into ~/.config/opencode/auth.json or /connect → NVIDIA
  • Verify: NVIDIA_API_KEY=nvapi-xxx node tools/nvidia-nim/diagnose.mjs
Worker AIs (Hermes/Openclaw-like full time digital employees with their own long term memory)
Compliance Seats
Extra Seats - Mac Mini Paseo

Why so many models? What is this, a beauty pageant?

Learn more about our Model Portfolio Strategy →