Skip to content
APEXVYRA

For developers

One endpoint. Every model you need.

POST a request, get back a structured run — model selection, tool calls, and retries handled for you. SDKs for Python, Node, and Go, or call the REST API directly.

bash
curl https://api.apexvyra.example/v1/runs \
  -H "Authorization: Bearer $APEXVYRA_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "agent": "refund-triage",
    "input": "Refund order #4821"
  }'
Response Live
{
  "id": "run_8841",
  "status": "succeeded",
  "model": "balanced",
  "output": "Refund of $42.00 issued for order #4821.",
  "latency_ms": 428,
  "cost_usd": 0.0021
}

One request, one run

Every run is a single request — and a full trace

You don't manage a queue, a retry policy, or a separate logging call. One request returns the result and everything that happened while producing it: which model ran, which tools it called, and what each step cost.

json
{
  "id": "run_8841",
  "trace": [
    { "step": "route", "model": "balanced" },
    { "step": "tool_call", "name": "lookup_order" },
    { "step": "guardrail", "check": "pii_redaction" }
  ]
}

Quickstart

From API key to first run in four steps

  1. 1

    Get your API key

    Create a free workspace and generate a key from Settings → API Keys.

  2. 2

    Install the SDK

    pip install apexvyra, npm install apexvyra, or go get apexvyra — or skip this and call the REST API directly.

  3. 3

    Make your first request

    Send an input string to an agent. You'll get a run object back with the result.

  4. 4

    Add routing, tools, and guardrails

    Layer in a routing policy, tool definitions, and guardrails as your agent grows — no changes to the calling code.

Built for the API

Handled for you, so your integration code stays small

Streaming responses

Token-by-token output over Server-Sent Events, or wait for the full run.

Structured outputs

Constrain a run's output to a JSON schema you define.

Function calling

Declare tools once; every routed model gets a compatible calling format.

Automatic retries

Transient provider failures retry with backoff before they reach you.

Rate limit handling

Requests queue and smooth out automatically as you approach a limit.

Idempotency keys

Retry a request safely — pass the same key and get the same run back.

Changelog

Recently shipped

  1. Streaming tool calls

    Tool calls now stream incrementally instead of arriving as one block at the end of a run.

  2. Go SDK

    A first-party Go client, matching the existing Python and Node SDKs.

  3. Idempotency keys

    Pass an idempotency key to /v1/runs to make retries safe.

Read the docs, then ship something.

Full API reference, SDK guides, and example agents.