# Use Local Models via API (/docs/development/use-local-models-via-api)



When developing against BullSequana AI, the preferred integration point is the `CoreAI API`.

Do not integrate directly with `LiteLLM` for application development. LiteLLM is an internal platform component and may change over time. The CoreAI API is the stable interface that BullSequana AI exposes to developers.

Recommended Background [#recommended-background]

This page is primarily for application developers and AI engineers integrating external tools or custom applications with the platform.

Helpful experience includes:

* API-based application development
* bearer-token or API-key authentication
* local development with Docker when reproducing platform-adjacent flows
* basic understanding of OpenAI-compatible clients and SDKs

What To Use [#what-to-use]

Use the API endpoint exposed by your BullSequana AI deployment:

```bash
export BSQAI_BASE_URL="https://llm-backend.<platform-domain>/v1"
```

The `llm-backend` subdomain stays the same across environments. Only the platform domain changes from one cluster to another.

This API already handles the platform integration behind the scenes:

* model routing
* authentication
* platform policy enforcement
* future component evolution behind the API boundary

Authentication Options [#authentication-options]

The backend supports both of these:

* `JWT bearer tokens`
* `BullSequana API keys` in the `sk-bsq-...` format

For interactive user access, JWT is a good fit.

For local developer tools, scripts, IDE plugins, and long-lived integrations, API keys are usually the cleaner option when your deployment exposes API key management.

How to pass credentials [#how-to-pass-credentials]

API keys can be sent in either of these ways:

* `Authorization: Bearer sk-bsq-...` — the same header used for JWTs
* `X-Api-Key: sk-bsq-...` — a dedicated API-key header

Both are accepted. JWT tokens always use `Authorization: Bearer <jwt-token>`.

When using OpenAI-compatible SDKs or tools, the `Authorization: Bearer` form is the most convenient because most clients already send the configured key in that header. The `X-Api-Key` header is available as an alternative when the `Authorization` header is reserved for another purpose or when you want an explicit separation between JWT and API-key authentication.

<Callout type="info" title="Requires platform 1.2.1 or later for the Bearer form">
  On platform versions before 1.2.1, API keys sent as `Authorization: Bearer` are rejected at the gateway with `Jwt is not in the form of Header.Payload.Signature` — the request never reaches the CoreAI API. On those versions, use the `X-Api-Key` header instead.
</Callout>

Option 1: Authenticate with JWT [#option-1-authenticate-with-jwt]

If your platform uses Keycloak-backed interactive access, obtain a JWT and pass it as the bearer token.

Example:

```bash
export BSQAI_TOKEN="<jwt-token>"

curl "$BSQAI_BASE_URL/models" \
  -H "Authorization: Bearer $BSQAI_TOKEN"
```

In local backend development, the sibling `coreai-llm-backend` repo already includes a helper script:

```bash
./scripts/get_jwt_token.sh --print-token
```

Option 2: Create and use an API key [#option-2-create-and-use-an-api-key]

API keys are created through the CoreAI API itself and are returned only once when created.

First call the API with a valid JWT:

```bash
curl -X POST "https://llm-backend.<platform-domain>/v1/api-keys" \
  -H "Authorization: Bearer <jwt-token>" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "Local Development"
  }'
```

Typical response shape:

```json
{
  "id": "a37e6b2e-728e-4ffc-aea4-e5415616a96a",
  "name": "Local Development",
  "key": "sk-bsq-v1-...",
  "masked": "sk-bsq-v1-********-e5b4a3",
  "created_at": "2024-01-15T10:30:00Z"
}
```

Then use that key for requests. Both header styles are accepted:

```bash
export BSQAI_API_KEY="sk-bsq-v1-..."

# Option A: Bearer header (works with most OpenAI-compatible SDKs)
curl "$BSQAI_BASE_URL/models" \
  -H "Authorization: Bearer $BSQAI_API_KEY"

# Option B: Dedicated API-key header
curl "$BSQAI_BASE_URL/models" \
  -H "X-Api-Key: $BSQAI_API_KEY"
```

Discover Available Models [#discover-available-models]

Before wiring an application, check which models your deployment exposes:

Available models are also visible in the `CoreAI Portal`.

If you want the API-level source of truth for the current environment, use:

```bash
curl "$BSQAI_BASE_URL/models" \
  -H "Authorization: Bearer $BSQAI_API_KEY"
```

This is the correct source of truth for model names in your environment.

Use the OpenAI-Compatible API [#use-the-openai-compatible-api]

The CoreAI API implements the OpenAI **Responses API**:

* `/v1/models`
* `/v1/responses`

That means SDKs and tools that speak the Responses API can be pointed to BullSequana AI with only a base URL and token change. The official OpenAI SDKs use the Responses API by default.

The legacy Chat Completions endpoint (`/v1/chat/completions`) is **not** available. Clients that only implement Chat Completions receive `404 Not Found`; they must be configured for (or updated to support) the Responses API. Note that the Responses API takes an `input` field, not a `messages` array — sending `messages` returns a `user_message_missing` error.

Responses API example (curl) [#responses-api-example-curl]

```bash
curl -X POST "$BSQAI_BASE_URL/responses" \
  -H "Authorization: Bearer $BSQAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "<model-name-from-/v1/models>",
    "input": [
      { "role": "user", "content": "Hello from BullSequana AI" }
    ]
  }'
```

Responses API example (Python) [#responses-api-example-python]

```python
from openai import OpenAI

client = OpenAI(
    api_key="sk-bsq-v1-...",
    base_url="https://llm-backend.<platform-domain>/v1",
)

response = client.responses.create(
    model="<model-name-from-/v1/models>",
    input="Explain the deployment model of BullSequana AI."
)

print(response)
```

IDE And Tooling Integrations [#ide-and-tooling-integrations]

For tools such as `OpenCode` or custom internal applications, point the tool to the CoreAI API, not to LiteLLM directly.

The correct pattern is:

* provider type: OpenAI (Responses API)
* base URL: your BullSequana AI CoreAI API
* bearer token: JWT or `sk-bsq-...` API key

The tool must use the OpenAI Responses API for the platform's model names. `Continue`, for example, only uses the Responses API for OpenAI's own model naming (o-series, GPT-5+) and routes all other model names to Chat Completions, so it cannot currently connect to the CoreAI API (see [Continue](/docs/development/developer-tools/continue)).

Recommended Integration Rule [#recommended-integration-rule]

Use this rule for all developer-facing integrations:

`Application or tool -> CoreAI API -> platform components`

Not:

`Application or tool -> LiteLLM`

That keeps the developer contract stable even if the platform team changes the internal inference or proxy layer later.

Related Pages [#related-pages]

* [CoreAI API](/docs/coreai/components/coreai-api)
* [Developer Tools](/docs/development/developer-tools)
* [Deploy apps on the platform](/docs/development/deploy-apps-on-the-platform)
