Skip to content
Engineering

Multimodal API compatibility from the adapter layer

Ian Eaves

Ian Eaves

2026-07-1314 min read

Multimodal API compatibility from the adapter layer
Engineering

Implementation notes from building compatibility across OpenAI, OpenRouter, Hugging Face, Cartesia, and ElevenLabs. Request shape, response shape, and usage units should be treated as three separate problems.

Implementation notes on request shape, response shape, and usage units

We recently implemented compatibility across several AI gateways and model-provider APIs, including OpenAI-style chat APIs, OpenRouter, Hugging Face Inference Providers, Cartesia, and ElevenLabs. The work was mostly straightforward in places where the underlying task looked like chat. It required more judgment in places where the task involved audio, images, realtime sessions, task-specific response objects, or provider-specific usage fields.

The distinction between these different modalities and model capabilities ends up really mattering if you are building against more than one provider. A request can be easy to send even when the response needs task-specific parsing. A response can be easy to parse while the usage object still needs provider-specific interpretation. A usage object can include a number called audio_tokens or total_tokens while leaving open what that number means for pricing, display, or comparison.

The useful framing for us was to separate compatibility into three layers:

  1. Request shape: what does the caller send?
  2. Response shape: what does the provider return?
  3. Usage and pricing semantics: what was counted, in what unit, and how should the caller compare it?
Three side-by-side boxes labeled Request shape, Response shape, and Usage semantics, each listing example fields the layer is responsible for
Figure 1Request shape, response shape, and usage semantics each diverge independently per provider. Compatibility on one does not imply compatibility on the others.

For basic chat completions, those three layers have some loose convergence. A caller usually sends messages. The messages usually contain roles and content. The response usually contains choices, messages, finish reasons, or a close equivalent. The usage object usually includes some version of input tokens, output tokens, and total tokens.

That is enough to build against in many text-chat cases but it is not enough to make multimodal compatibility automatic. Once the same adapter layer has to support speech-to-text, text-to-speech, image tasks, realtime audio, routed LLM calls, and multimodal chat, the three layers start to diverge.

This post is an after-action report from that implementation work. The goal is not to say one provider's API shape is better than another's. Most of the shapes we looked at were reasonable for the task they served. The main lesson was that request compatibility, response compatibility, and usage compatibility should be treated as separate problems.

Request Shape

The request layer is where OpenAI-style compatibility has been most useful. For chat, a shared request shape gets you a long way. A caller can usually reason in terms of messages, role, content, model selection, generation parameters, and sometimes tools.

That pattern is less complete once the task is no longer primarily chat. Hugging Face's provider docs describe this split directly. For LLM and VLM tasks, providers that support the OpenAI API can often avoid much of the custom integration work. For other tasks, including text-to-image and automatic speech recognition, Hugging Face documents task-specific schemas and notes that provider adapters may need to translate parameter names and output names.

That matched what we saw in practice. Chat requests often fit a familiar outer shape. Multimodal requests often carry task-specific fields that should remain visible.

A text-to-speech request is a simple example. The caller is not only asking a model to produce an output. The caller is asking for speech with a particular voice, format, language, timing behavior, and sometimes continuity across multiple inputs.

Cartesia's TTS API includes fields such as model_id, transcript, voice, output_format, context_id, timestamp options, pronunciation dictionary settings, and generation config.

json
// Cartesia TTS request, simplified
{
  "model_id": "sonic-model",
  "transcript": "Text to speak",
  "voice": {
    "id": "voice-id"
  },
  "output_format": {},
  "context_id": "context-id",
  "add_timestamps": true,
  "add_phoneme_timestamps": false,
  "generation_config": {
    "speed": 1
  }
}

ElevenLabs exposes a different but still voice-native shape. Its standard TTS route puts voice_id in the path and accepts fields such as text, model_id, output_format, and voice_settings in the body.

json
// ElevenLabs TTS request, simplified
POST /v1/text-to-speech/:voice_id

{
  "text": "Text to speak",
  "model_id": "eleven_multilingual_v2",
  "voice_settings": {}
}

In some cases these are not just naming differences, rather they encode task behavior. A voice API generally needs to describe the voice, the output audio format, and the timing or streaming behavior. An image API may need resolution, aspect ratio, source image, mask, seed, number of outputs, or output format. A realtime API may need session state, buffering behavior, turn detection, interruption behavior, and event-level controls.

For the adapter layer, the request-shape question was therefore practical. Which fields are common enough to normalize? Which fields need to remain explicit task parameters? Which fields can be passed through without pretending they are portable?

A chat-compatible request shape helped when the task was chat-like. For task-native multimodal APIs, it was usually cleaner to preserve the task's request model and normalize around it.

Response Shape

Even if two providers let you send similar requests, the output may not have the same shape or respect the same API concerns. While partly a function of the API it's also impacted by the underlying model-capability.

ASR timestamps are a good example of this. At the interface layer, timestamps can look like a response-format option. Transcription requests somewhere like HuggingFace might include a field like return_timestamps indicating that the response should include timestamped chunks alongside the raw transcription. In practice the ability to generate those timestamps might not exist, indeed often does not exist, intrinsically to the model.

Take the case of Qwen3-ASR. The vLLM-backed Qwen3-ASR example accepts audio through a chat-style request and returns the transcription through a completion-like response. The caller sends an audio message and reads the output from choices[0].message.content.

However, actually generating timestamps is delegated to a subordinate model designed to accept both the transcription and the raw speech. This would be something like Qwen3-ForcedAligner-0.6B used for timestamp prediction. So timestamps become not only a formatting choice on the ASR response but may actually require a second alignment path coordinated by the backend to generate the response expected by the user.

Pipeline diagram. Audio feeds Qwen3-ASR, which produces a transcript with no timing. The original audio and the transcript both feed a second model, Qwen3-ForcedAligner, which produces the timestamped output.
Figure 2Qwen3-ASR produces a transcript with no timing. Generating timestamps requires a second model that consumes both the transcript and the original audio.

These details expose complicated interactions between the various serving modes as well. While Qwen3-ASR supports streaming inference through the vLLM backend, it's not easily possible to stream timestamp aligned output because the subordinate timestamp generation requires a complicated transcript and associated text to complete. So an adapter cannot easily model timestamps as supported: true for the model family.

The same pattern shows up on the speech-generation side. vLLM-Omni exposes an OpenAI-compatible /v1/audio/speech endpoint for text-to-speech and for Qwen3-TTS, that interface can make several model operations look similar from the outside despite performing substantively different operations internally. With CustomVoice, the caller selects a predefined speaker and can optionally provide style instructions:

json
{
  "input": "Hello world",
  "task_type": "VoiceDesign",
  "instructions": "A warm, friendly female voice with a gentle tone"
}

With VoiceDesign, the caller describes the desired voice in natural language:

json
{
  "type": "speech",
  "audio": "<binary or url>",
  "format": "wav",
  "task_type": "VoiceDesign",
  "voice_source": {
    "type": "natural_language_description"
  }
}

With Base, the caller is doing voice cloning from reference audio:

json
{
  "input": "Hello, this is a cloned voice",
  "task_type": "Base",
  "ref_audio": "https://example.com/reference.wav",
  "ref_text": "Original transcript of the reference audio"
}

All three calls look completion-like and generate text-to-audio responses but the response semantics are different because the voice source is different. For CustomVoice, the voice source is a named or predefined speaker. For VoiceDesign, the voice source is a natural-language description. For Base, the voice source is reference audio. If the adapter only returns bytes, playback works but the response no longer explains which kind of speech-generation operation happened or even all of the artifacts (like the id for a persistent voice clone) which might have been produced.

Usage and Pricing Semantics

Usage was the hardest layer to normalize because it sits between API design and pricing. For text chat, the classic usage object is fairly easy to understand being composed principally of prompt tokens, completion tokens, and total tokens (and putting aside the tokenization differences between different models). Those fields map reasonably well to input text, output text, and total model-token volume.

That model becomes less clean for multimodal APIs. A multimodal request may include text, audio, images, video, files, tool results, cached context, or some mix of those. A multimodal response may produce text, audio, images, video, structured data, or intermediate reasoning. Those inputs and outputs can have very different computational and caching profiles.

So rather than just answering "how many tokens" the usage object has to answer questions like:

  • Was this input or output?
  • Which modality was counted?
  • Was the quantity measured in tokens, seconds, characters, images, frames, bytes, credits, or another unit?
  • If the provider reports media tokens, what does one token mean?
  • If the provider reports seconds, frames, or generated media count, what compute detail is being hidden?
  • Is the usage object meant for user display, provider billing, routing decisions, or all three?

OpenRouter is a useful example because it keeps an OpenAI-style chat surface while adding more usage and cost detail. Its usage accounting can include prompt tokens, completion tokens, reasoning tokens, cached tokens, cost, upstream inference cost, and token counts calculated with the model's native tokenizer.

json
// OpenRouter usage, simplified from its documented shape
{
  "usage": {
    "prompt_tokens": 194,
    "completion_tokens": 2,
    "total_tokens": 196,
    "prompt_tokens_details": {
      "cached_tokens": 0,
      "cache_write_tokens": 100,
      "audio_tokens": 0
    },
    "completion_tokens_details": {
      "reasoning_tokens": 0
    },
    "cost": 0.95,
    "cost_details": {
      "upstream_inference_cost": 19
    }
  }
}

OpenAI's chat usage shape also has more detail than the old three-field summary. It can include nested detail objects for cached tokens, audio tokens, reasoning tokens, and accepted or rejected prediction tokens.

json
// OpenAI-style chat usage, simplified
{
  "usage": {
    "prompt_tokens": 1200,
    "completion_tokens": 340,
    "total_tokens": 1540,
    "prompt_tokens_details": {
      "cached_tokens": 800,
      "audio_tokens": 0
    },
    "completion_tokens_details": {
      "reasoning_tokens": 0,
      "audio_tokens": 0
    }
  }
}

That helps, but the structure is still anchored around prompt and completion tokens. Once the output is not text, the words prompt and completion start to carry less information. A text output token, an audio output token, and a generated image are all output-side usage. They are not the same kind of thing.

Take a look at OpenAI's Realtime API usage response: Its usage object can break input and output tokens into modality-specific detail buckets, including text, audio, image, and cached tokens.

json
// OpenAI Realtime usage, simplified
{
  "usage": {
    "total_tokens": 253,
    "input_tokens": 132,
    "output_tokens": 121,
    "input_token_details": {
      "text_tokens": 119,
      "audio_tokens": 13,
      "image_tokens": 0,
      "cached_tokens": 64
    },
    "output_token_details": {
      "text_tokens": 30,
      "audio_tokens": 91
    }
  }
}

That is closer to what a user really needs. It separates input from output and gives modality-specific detail but it also shows the pricing problem.

OpenAI's Realtime api computes input tokens at a ratio of 100ms audio / token but output at 50ms audio / token. Others, like Alibaba, have variable conversion rates based on the referenced model. This ambiguity makes interrogating usage objects particularly tricky, particularly when comparing across model providers.

This was a broad pattern which emerges across multimodal pricing. The lack of structured units or standardized output objects make it very difficult to apply consistent output schemas especially when dealing with more custom model behaviors like the creation or usage of custom voices. This issue becomes especially pronounced when the visible API unit and the pricing unit are not always the same.

Seconds are easy for users to understand even though they are often an incomplete proxy for actual provider cost. A five-second audio output can vary by model, sample rate, codec, voice settings, and generation path. A five-second video can vary by resolution, frame rate, architecture, denoising steps, and batching behavior. One generated image can be cheap or expensive depending on resolution, model architecture, quality tier, and generation settings.

With newer serving architectures like vLLM-Omni make this more than simple bookkeeping. Any-to-any multimodal serving work increasingly describes these systems as multi-stage pipelines that can combine autoregressive LLM stages, diffusion stages, encoders, vocoders, and connector logic between stages. Each of these involve different underlying caching behaviors and compute profiles and different proxies can be more or less relevant depending on the balance amongst them.

For the caller, the result may be one audio response but for the provider, that response may have crossed several components before it was eventually returned. That is why usage standardization is harder than adding a few fields to total_tokens.

Usage Shape

The structure we kept wanting was itemized. Instead of making completion_tokens carry every output concept, usage could be reported as typed input and output items with explicit units.

jsonc
{
  "usage": {
    "inputs": [
      {
        "type": "text",
        "unit": "token",
        "quantity": 1200
      },
      {
        "type": "audio",
        "unit": "second",
        "quantity": 12.4,
        "provider_units": [
          {
            "unit": "audio_token",
            "quantity": 124,
            "unit_duration_ms": 100
          }
        ]
      }
    ],
    "outputs": [
      {
        "type": "text",
        "unit": "token",
        "quantity": 340
      },
      {
        "type": "audio",
        "unit": "second",
        "quantity": 8.7,
        "provider_units": [
          {
            "unit": "audio_token",
            "quantity": 174,
            "unit_duration_ms": 50
          }
        ]
      },
      {
        "type": "image",
        "unit": "image",
        "quantity": 1,
        "attributes": {
          "width": 1024,
          "height": 1024
        }
      }
    ],
    "totals": [
      {
        "unit": "token",
        "quantity": 1878
      }
    ]
  }
}

The important fields are small:

jsonc
{
  "type": "audio",
  "unit": "second",
  "quantity": 8.7
}

type preserves the modality. unit says how the quantity should be read. quantity gives the count. provider_units gives the provider room to expose native accounting units without making those units look universal.

With that structure, a text-only chat response can still produce familiar derived fields. completion_tokens can be derived from outputs[type=text, unit=token]. A speech response can expose seconds for user-facing display and audio tokens for provider-native billing. An image response can expose image count and resolution without pretending the output was a text completion.

This does not require every provider to expose the same internals. It only asks each usage item to say what was counted and how it was counted. To be clear this is still a work in progress on our end as well but our goal is to make this as transparent as possible for the end user.

Adapter Design

The implementation pattern that ended up working was not to choose one provider shape as the internal standard. Instead, we defined canonical objects for the three parts of the interaction:

  • request
  • response
  • usage

Each provider integration then became a pair of translators around those canonical objects.

On the way in, the integration accepts the provider's external shape and converts it into the canonical request.

text
Hugging Face request
-> Hugging Face request parser
-> canonical request
-> model execution

On the way out, the process runs in the other direction, but with one important addition.

When the model response is received, we compute canonical usage from the execution result. Then we convert both the canonical response and the canonical usage into the provider-specific response shape.

text
model execution result
-> canonical response
-> canonical usage
-> Hugging Face response formatter
-> Hugging Face-shaped response

For a Hugging Face-style integration, the full path looks roughly like this:

Three-band architecture diagram. Top band: Hugging Face request and response. Middle band: canonical request and canonical response plus usage. Bottom band: execute and model result. Arrows show parsing down from the edge into canonical, forwarding into the model, and formatting back up to the provider response.
Figure 3Each adapter translates between its provider's vernacular and the canonical layer. The model only ever sees canonical objects; provider-specific shapes stay at the edge.

That structure made the compatibility problem easier to reason about.

The Hugging Face adapter did not need the model backend to speak Hugging Face. The OpenRouter adapter did not need the model backend to speak OpenRouter. The OpenAI-compatible adapter did not need every internal task to behave like an OpenAI chat completion. Each external integration only needed to translate between its own vernacular and the canonical layer.

This also kept request, response, and usage separate.

A provider's request shape might be easy to support while its response shape required task-specific metadata. A provider's response shape might be easy to produce while its usage object required provider-specific accounting fields. A provider might use familiar chat-style request and response objects while still needing usage fields that did not map cleanly onto prompt_tokens, completion_tokens, and total_tokens. The canonical layer gave us a place to preserve the richer internal meaning before projecting it back into a provider-specific shape.

For requests, the canonical object had to describe the task without losing modality-specific inputs.

json
{
  "task": "speech_generation",
  "inputs": [
    {
      "type": "text",
      "text": "Hello world"
    }
  ],
  "parameters": {
    "voice_source": {
      "type": "named_voice",
      "voice": "vivian"
    },
    "format": "wav"
  }
}

For responses, the canonical object had to describe what the model actually produced, including any task variant or capability boundary.

json
{
  "type": "speech",
  "outputs": [
    {
      "type": "audio",
      "format": "wav",
      "data": "<binary or url>"
    }
  ],
  "metadata": {
    "task_type": "CustomVoice",
    "voice_source": {
      "type": "named_voice",
      "voice": "vivian"
    }
  }
}

If the caller expects an OpenAI-style response, the adapter can derive something like completion_tokens where that field makes sense. If the caller expects an OpenRouter-style response, the adapter can include cost and provider-specific usage details. If the caller expects a Hugging Face task response, the adapter can return the task-shaped result while still computing usage internally in the same canonical format.

The point is not that every provider should expose the canonical object directly, rather that the adapter layer needs some representation that is more stable than any one provider's API shape.

Provider-specific shapes are edge contracts. The canonical objects are our internal contracts.

When a new provider had a different request body, we wrote a new request translator. When a provider expected a different response shape, we wrote a new response formatter. When a provider used different usage units, we mapped from canonical usage into that provider's usage vocabulary. This way the model backend never needed to know which external API shape the caller used and because the responses were always converted at the edge our internal logic could be consistently structured around coherent and well defined typed objects.

Takeaways

The adapter work ended up being less about choosing one provider's API as the internal standard and more about deciding where translation should happen.

For chat-shaped APIs, the translation layer can be fairly thin. The external request, internal request, model call, internal response, and external response all roughly share the same structure. There are still differences, but the basic contract is familiar enough that an adapter mostly has to handle naming, validation, and a few provider-specific fields.

Multimodal APIs required a thicker adapter layer.

The request shape had to preserve the caller's external API contract. A Hugging Face-style request should be accepted as Hugging Face-shaped input. An OpenAI-compatible request should be accepted as OpenAI-shaped input. A task-native speech or image request should be accepted in the form that makes sense for that task.

Internally, though, those requests needed to become canonical requests. That gave the model backend one representation to execute against instead of forcing every model path to understand every provider's vocabulary.

The same pattern applied on the way out. The model result became a canonical response. Usage was computed in a canonical form. Then the adapter converted the canonical response and canonical usage back into the external shape the caller expected.

That architecture made the implementation tractable, but it also made the standardization gap clearer.

Where for chat, the provider-facing contract is usually standardized around openai completions, for multimodal APIs, canonicalization becomes part of the core implementation. The external API may look familiar while the model path depends on task variants, optional sidecar models, response capabilities, or serving-mode constraints. The response may contain the same final artifact while carrying different semantic meaning. The usage object may contain a familiar token field while relying on a provider-specific unit definition.

Pricing made that last point the most visible. A single product surface may need to price text tokens, cached tokens, reasoning tokens, audio seconds, audio tokens, generated images, video duration, frames, resolution, or provider credits and those units truly are not interchangeable. Some are user-readable. Some are provider-native. Some are closer to compute. Some are only useful after applying a conversion rule.

So the practical takeaway is not that every multimodal API should converge on one request body. The more useful target is a clearer separation between edge contracts and internal contracts.

External integrations can speak the provider's vernacular. The internal system can normalize request, response, and usage separately. Then each adapter can project the canonical objects back into the shape the caller expects.

That is roughly where multimodal API compatibility seems to be today. It can be implemented, but it requires explicit translation. The request and response contracts are still less uniform than chat. The usage and pricing contract is even less uniform because the units themselves vary by modality, provider, and model path.

A usage schema built around type, unit, and quantity would not fully remove that complexity but it would make the accounting layer easier to reason about for end users. Our advice to anyone building across gateways or providers is to normalize internally, preserve provider semantics at the edge, and avoid treating request compatibility, response compatibility, and usage compatibility as the same problem.