---
sidebar_label: Fireworks AI
description: Configure Fireworks AI's serverless chat and embedding models through their OpenAI-compatible API for LLM evaluation and testing with promptfoo
---

# Fireworks AI

[Fireworks AI](https://fireworks.ai) serves a broad catalogue of open models — Llama, Qwen, DeepSeek, Kimi, GLM, GPT-OSS, and more — through OpenAI-compatible chat and embedding endpoints.

The Fireworks AI provider supports all options available in the [OpenAI provider](/docs/providers/openai/).

## Setup

Create an API key from the Fireworks dashboard (**Settings → API Keys**) and expose it as an environment variable:

```sh
export FIREWORKS_API_KEY=your_api_key_here
```

The provider keeps Fireworks credentials isolated from OpenAI's: it reads `FIREWORKS_API_KEY` (never `OPENAI_API_KEY`) and never inherits `OPENAI_API_HOST` / `OPENAI_API_BASE_URL` / `OPENAI_ORGANIZATION`, so a stray OpenAI variable in your environment can't leak onto or reroute Fireworks requests.

## Provider format

- `fireworks:<model>` — chat completions, e.g. `fireworks:accounts/fireworks/models/gpt-oss-120b`
- `fireworks:embedding:<model>` — embeddings, e.g. `fireworks:embedding:fireworks/qwen3-embedding-8b`

Copy the exact identifier from the model's documentation. Chat models commonly use `accounts/fireworks/models/<model>`; embedding models can use a different namespace, and dedicated deployments use `accounts/<account>/deployments/<deployment>`. Check the [serverless catalogue](https://fireworks.ai/models?deployment=serverless) and your deployment configuration for availability.

Current Chinese model releases use these Fireworks IDs (prefix each with `fireworks:`):

| Model                                                                              | Model ID                                         |
| ---------------------------------------------------------------------------------- | ------------------------------------------------ |
| [DeepSeek V4.1 Flash](https://fireworks.ai/models/deepseek-ai/deepseek-v4p1-flash) | `accounts/fireworks/models/deepseek-v4p1-flash`  |
| [DeepSeek V4 Pro](https://fireworks.ai/models/deepseek-ai/deepseek-v4-pro-0813)    | `accounts/fireworks/models/deepseek-v4-pro-0813` |
| [GLM-5.3](https://fireworks.ai/models/fireworks/glm-5p3)                           | `accounts/fireworks/models/glm-5p3`              |
| [GLM-5.3 Flash](https://fireworks.ai/models/fireworks/glm-5p3-flash)               | `accounts/fireworks/models/glm-5p3-flash`        |
| [Kimi K3](https://fireworks.ai/models/fireworks/kimi-k3)                           | `accounts/fireworks/models/kimi-k3`              |
| [Qwen3.8 Max](https://fireworks.ai/models/fireworks/qwen3p8-max)                   | `accounts/fireworks/models/qwen3p8-max`          |
| [MiniMax M3](https://fireworks.ai/models/fireworks/minimax-m3)                     | `accounts/fireworks/models/minimax-m3`           |

## Example Usage

```yaml
providers:
  - id: fireworks:accounts/fireworks/models/gpt-oss-120b
    config:
      temperature: 0.2
      max_tokens: 1024
      apiKey: ... # optional; overrides FIREWORKS_API_KEY
```

:::note
Many of Fireworks's flagship models are reasoning models that emit hidden reasoning tokens before the visible answer. Set `max_tokens` high enough to leave room for both — otherwise the response can be truncated to empty output.
:::

Run the bundled example end-to-end:

```sh
npx promptfoo@latest init --example provider-fireworks
```

## Embeddings

Fireworks serves embedding models on the same key via the `fireworks:embedding:` prefix. For example, to grade a [`similar` assertion](/docs/configuration/expected-outputs/similar) with a Fireworks embedding model:

```yaml
defaultTest:
  options:
    provider:
      embedding:
        id: fireworks:embedding:fireworks/qwen3-embedding-8b
```

The [Fireworks embedding guide](https://docs.fireworks.ai/guides/querying-embeddings-models) uses `fireworks/qwen3-embedding-8b` for serverless requests. Keep existing model identifiers, dimensions, and preprocessing consistent with stored vectors; changing embedding models requires rebuilding the corresponding index and recalibrating similarity thresholds.

Embedding request options go under `config.passthrough`. For resizable models such as Qwen3, set `config.passthrough.dimensions` only when you intentionally choose an output dimension. Qwen3 vectors are not unit-normalized; normalize them before treating a raw dot product as cosine similarity, or use a cosine-similarity function that handles normalization.

Voyage embedding models require a dedicated deployment. Use `fireworks:embedding:accounts/<account>/deployments/<deployment>` with your deployment's identifiers, and set `config.passthrough.input_type` to `document` for corpus embeddings or `query` for search queries. Keep the model, dimensions, and role-specific preprocessing consistent across indexing and retrieval.

## Configuration

Because the provider extends the OpenAI provider, all [OpenAI configuration parameters](/docs/providers/openai/#configuring-parameters) apply. The most common options:

| Option                                             | Description                                                                                                                                                |
| -------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `apiKey`                                           | Fireworks API key (overrides the `FIREWORKS_API_KEY` environment variable).                                                                                |
| `apiBaseUrl`                                       | Base URL override. Can also be set with the `FIREWORKS_API_BASE_URL` environment variable.                                                                 |
| `apiHost`                                          | Host override for a proxy or gateway; resolves to `https://<apiHost>/v1`.                                                                                  |
| `temperature`, `max_tokens`, `top_p`, `top_k`, ... | Standard OpenAI-compatible sampling parameters.                                                                                                            |
| `cost`, `inputCost`, `outputCost`                  | Override promptfoo's cost estimate (USD per token). Use `inputCost` and `outputCost` for asymmetric pricing; `cost` is the shared fallback.                |
| `cacheReadInputCost`                               | Per-token rate for Fireworks server-side prompt-cache hits. Defaults to the full `inputCost` (no discount is assumed, since the discount varies by model). |

| Environment variable     | Description                                                        |
| ------------------------ | ------------------------------------------------------------------ |
| `FIREWORKS_API_KEY`      | Your Fireworks API key.                                            |
| `FIREWORKS_API_BASE_URL` | Override the base URL (defaults to the public Fireworks endpoint). |

### Cost tracking

Fireworks prices each model differently, so promptfoo can't infer a per-token rate. Supply `inputCost` and `outputCost` to surface spend estimates in your eval results:

```yaml
providers:
  - id: fireworks:accounts/fireworks/models/gpt-oss-120b
    config:
      inputCost: 0.00000015 # $0.15 / 1M input tokens
      outputCost: 0.0000006 # $0.60 / 1M output tokens
```

If you rely on Fireworks's server-side prompt caching, set `cacheReadInputCost` to the discounted cached-input rate; otherwise cached prompt tokens are billed at the full `inputCost`.

## API Details

- **Base URL**: `https://api.fireworks.ai/inference/v1`
- **API format**: OpenAI-compatible (`/chat/completions`, `/embeddings`)
- **Models**: [serverless model catalogue](https://fireworks.ai/models?deployment=serverless)
- Full [API documentation](https://docs.fireworks.ai)
