> ## Documentation Index
> Fetch the complete documentation index at: https://langchain-5e9cc07a-preview-change-1791323909-75c753a.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# ChatLiteLLM and ChatLiteLLMRouter integration

> Integrate with the ChatLiteLLM and ChatLiteLLMRouter chat model using LangChain Python.

[LiteLLM](https://github.com/BerriAI/litellm) is a library that simplifies calling Anthropic, Azure, Huggingface, Replicate, etc.

This page covers how to get started using LangChain with the LiteLLM I/O library.

This integration provides two chat model classes:

* [ChatLiteLLM](https://reference.langchain.com/python/langchain-litellm/chat_models/litellm/ChatLiteLLM): The main LangChain chat wrapper for LiteLLM.
* [ChatLiteLLMRouter](https://reference.langchain.com/python/langchain-litellm/chat_models/litellm_router/ChatLiteLLMRouter): A `ChatLiteLLM` wrapper that leverages LiteLLM's Router for load balancing and fallbacks.

The package also ships [LiteLLMEmbeddings](https://reference.langchain.com/python/langchain-litellm/embeddings/litellm/LiteLLMEmbeddings), [LiteLLMEmbeddingsRouter](https://reference.langchain.com/python/langchain-litellm/embeddings/litellm_router/LiteLLMEmbeddingsRouter), and [LiteLLMOCRLoader](https://reference.langchain.com/python/langchain-litellm/document_loaders/litellm_ocr/LiteLLMOCRLoader). See the [providers page](/oss/python/integrations/providers/litellm) for details.

## Overview

### Integration details

| Class | Package | Serializable | JS support | Downloads | Version |
| :- | :- | :-: | :-: | :-: | :-: |
| [ChatLiteLLM](https://reference.langchain.com/python/langchain-litellm/chat_models/litellm/ChatLiteLLM) | [`langchain-litellm`](https://pypi.org/project/langchain-litellm/) | ❌ | ❌ | ![PyPI - Downloads](https://img.shields.io/pypi/dm/langchain-litellm?style=flat-square\&label=%20) | ![PyPI - Version](https://img.shields.io/pypi/v/langchain-litellm?style=flat-square\&label=%20) |
| [ChatLiteLLMRouter](https://reference.langchain.com/python/langchain-litellm/chat_models/litellm_router/ChatLiteLLMRouter) | [`langchain-litellm`](https://pypi.org/project/langchain-litellm/) | ❌ | ❌ | ![PyPI - Downloads](https://img.shields.io/pypi/dm/langchain-litellm?style=flat-square\&label=%20) | ![PyPI - Version](https://img.shields.io/pypi/v/langchain-litellm?style=flat-square\&label=%20) |

### Model features

| [Tool calling](/oss/python/langchain/tools) | [Structured output](/oss/python/langchain/structured-output) | Image input | Audio input | Video input | [Token-level streaming](/oss/python/integrations/chat/litellm#async-and-streaming-functionality) | [Native async](/oss/python/integrations/chat/litellm#async-and-streaming-functionality) | [Token usage](/oss/python/langchain/models#token-usage) | [Logprobs](/oss/python/langchain/models#log-probabilities) |
| :-: | :-: | :-: | :-: | :-: | :-: | :-: | :-: | :-: |
| ✅ | ✅ | ✅ | ✅ | ❌ | ✅ | ✅ | ✅ | ✅ |

### Setup

To access `ChatLiteLLM` and `ChatLiteLLMRouter` models, you'll need to install the `langchain-litellm` package and create an OpenAI, Anthropic, Azure, Replicate, OpenRouter, Hugging Face, Together AI, or Cohere account. Then, you have to get an API key and export it as an environment variable.

## Credentials

You have to choose the LLM provider you want and sign up with them to get their API key.

### Example - Anthropic

Head to the [Claude console](https://console.anthropic.com) to sign up and generate a Claude API key. Once you've done this set the `ANTHROPIC_API_KEY` environment variable:

### Example - OpenAI

Head to [platform.openai.com/api-keys](https://platform.openai.com/api-keys) to sign up for OpenAI and generate an API key. Once you've done this, set the OPENAI\_API\_KEY environment variable.

```python theme={null}
## Set ENV variables
import os

os.environ["OPENAI_API_KEY"] = "your-openai-key"
os.environ["ANTHROPIC_API_KEY"] = "your-anthropic-key"
```

### Installation

The LangChain LiteLLM integration is available in the `langchain-litellm` package:

```python theme={null}
pip install -qU langchain-litellm
```

## Instantiation

### ChatLiteLLM

You can instantiate a `ChatLiteLLM` model by providing a `model` name [supported by LiteLLM](https://docs.litellm.ai/docs/providers).

```python theme={null}
from langchain_litellm import ChatLiteLLM

llm = ChatLiteLLM(model="gpt-5.4-nano", temperature=0.1)
```

### ChatLiteLLMRouter

You can also leverage LiteLLM's routing capabilities by defining your model list as specified in the [LiteLLM routing documentation](https://docs.litellm.ai/docs/routing).

```python theme={null}
from langchain_litellm import ChatLiteLLMRouter
from litellm import Router

model_list = [
    {
        "model_name": "gpt-5.5",
        "litellm_params": {
            "model": "azure/gpt-5.5",
            "api_key": "<your-api-key>",
            "api_version": "2024-10-21",
            "api_base": "https://<your-endpoint>.openai.azure.com/",
        },
    },
    {
        "model_name": "gpt-5.5",
        "litellm_params": {
            "model": "azure/gpt-5.5",
            "api_key": "<your-api-key>",
            "api_version": "2024-10-21",
            "api_base": "https://<your-endpoint>.openai.azure.com/",
        },
    },
]
litellm_router = Router(model_list=model_list)
llm = ChatLiteLLMRouter(router=litellm_router, model_name="gpt-5.5", temperature=0.1)
```

## Invocation

Whether you've instantiated a `ChatLiteLLM` or a `ChatLiteLLMRouter`, you can now use the ChatModel through LangChain's API.

```python theme={null}
response = await llm.ainvoke(
    "Classify the text into neutral, negative or positive. Text: I think the food was okay. Sentiment:"
)
print(response)
```

```text theme={null}
content='Neutral' additional_kwargs={} response_metadata={'token_usage': Usage(completion_tokens=2, prompt_tokens=30, total_tokens=32, completion_tokens_details=CompletionTokensDetailsWrapper(accepted_prediction_tokens=0, audio_tokens=0, reasoning_tokens=0, rejected_prediction_tokens=0, text_tokens=None), prompt_tokens_details=PromptTokensDetailsWrapper(audio_tokens=0, cached_tokens=0, text_tokens=None, image_tokens=None)), 'model': 'gpt-3.5-turbo', 'finish_reason': 'stop', 'model_name': 'gpt-3.5-turbo'} id='run-ab6a3b21-eae8-4c27-acb2-add65a38221a-0' usage_metadata={'input_tokens': 30, 'output_tokens': 2, 'total_tokens': 32}
```

## Async and streaming functionality

`ChatLiteLLM` and `ChatLiteLLMRouter` also support async and streaming functionality:

```python theme={null}
stream = await llm.astream_events("Hello, please explain how antibiotics work", version="v3")
async for token in stream.text:
    print(token, end="")
```

```text theme={null}
Antibiotics are medications that fight bacterial infections in the body. They work by targeting specific bacteria and either killing them or preventing their growth and reproduction.

There are several different mechanisms by which antibiotics work. Some antibiotics work by disrupting the cell walls of bacteria, causing them to burst and die. Others interfere with the protein synthesis of bacteria, preventing them from growing and reproducing. Some antibiotics target the DNA or RNA of bacteria, disrupting their ability to replicate.

It is important to note that antibiotics only work against bacterial infections and not viral infections. It is also crucial to take antibiotics as prescribed by a healthcare professional and to complete the full course of treatment, even if symptoms improve before the medication is finished. This helps to prevent antibiotic resistance, where bacteria become resistant to the effects of antibiotics.
```

## Advanced features

### Gemini Enterprise Agent Platform grounding (Google Search)

Use Google Search grounding with Gemini Enterprise Agent Platform models (e.g., `gemini-3.6-flash`). Grounding metadata is returned in `response_metadata`, for both batch and streaming calls.

<Note>
  Reading streamed grounding metadata from `response_metadata` requires `langchain-litellm>=0.11.0`. Before that, `ChatLiteLLM` returns it in `additional_kwargs`, and `ChatLiteLLMRouter` does not return it when streaming.
</Note>

**Setup**

```python theme={null}
import os
from langchain_litellm import ChatLiteLLM

os.environ["VERTEXAI_PROJECT"] = "your-project-id"
os.environ["VERTEXAI_LOCATION"] = "us-central1"

llm = ChatLiteLLM(model="vertex_ai/gemini-2.5-flash", temperature=0)
```

**Batch usage**

```python theme={null}
# Invoke with Google Search tool enabled
response = llm.invoke(
    "What is the current stock price of Google?",
    tools=[{"googleSearch": {}}]
)

# Access citations & metadata
provider_fields = response.response_metadata.get("provider_specific_fields")
if provider_fields:
    # Vertex returns a list; the first item contains the grounding info
    print(provider_fields[0])
```

**Streaming usage**

```python theme={null}
stream = llm.stream_events(
    "What is the current stock price of Google?",
    version="v3",
    tools=[{"googleSearch": {}}],
)
for token in stream.text:
    print(token, end="", flush=True)
# Metadata is available on the full output message
output = stream.output
if "provider_specific_fields" in output.response_metadata:
    print("\n[Metadata Found]:", output.response_metadata["provider_specific_fields"])
```

### Responses API

<Note>
  `use_responses_api` requires `langchain-litellm>=0.10.0`.
</Note>

Set `use_responses_api=True` to send `ChatLiteLLM` calls to the provider's [Responses API](https://docs.litellm.ai/docs/response_api) instead of Chat Completions. LiteLLM translates each request and reply, so messages, tools, streaming, and structured output work as usual. OpenAI's built-in tools, such as web search, need this route:

```python theme={null}
from langchain_litellm import ChatLiteLLM

llm = ChatLiteLLM(model="openai/gpt-6-astra", use_responses_api=True)
llm_with_tools = llm.bind_tools([{"type": "web_search"}])

response = llm_with_tools.invoke("What was a positive news story from today?")
print(response.text)
```

* A model LiteLLM cannot send to a Responses API raises `ValueError` before any request goes out.
* LiteLLM drops Chat Completions-only parameters on this route, such as `stop`, `n`, and `seed`.
* When `use_responses_api` is `None` (the default) or `False`, LiteLLM picks the API itself, and it already sends some models, such as `gpt-5-pro`, to the Responses API.

#### Reasoning items

<Note>
  Keeping reasoning items requires `langchain-litellm>=0.11.0`.
</Note>

A reasoning model on the Responses API returns its reasoning as [reasoning items](https://developers.openai.com/api/docs/guides/reasoning). Sending a turn's item back on the next request lets the model continue its reasoning through a tool loop rather than start again after each tool call. An item goes back only with its encrypted content, which OpenAI returns by default on stateless requests, sent with `"store": False`. OpenAI also accepts an explicit `include` for that content:

```python theme={null}
from langchain.messages import HumanMessage
from langchain.tools import tool
from langchain_litellm import ChatLiteLLM


@tool
def get_weather(city: str) -> str:
    """Get the weather for a city."""
    return f"It is sunny in {city}."


llm = ChatLiteLLM(
    model="openai/gpt-6-astra",
    use_responses_api=True,
    model_kwargs={
        "extra_body": {"store": False, "include": ["reasoning.encrypted_content"]}
    },
)
llm_with_tools = llm.bind_tools([get_weather])

messages = [
    HumanMessage(
        "Get the weather for the city that hosted the Summer Olympics four years "
        "before London did."
    )
]
ai_message = llm_with_tools.invoke(messages)
reasoning_items = ai_message.additional_kwargs.get("reasoning_items", [])
print([item["id"] for item in reasoning_items])

messages.append(ai_message)
messages.extend(get_weather.invoke(call) for call in ai_message.tool_calls)
response = llm_with_tools.invoke(messages)
print(response.text)
```

A reply keeps its reasoning items in `additional_kwargs["reasoning_items"]`, and only the items that carry encrypted content. A turn the model answers without reasoning has none. Each kept item carries an `origin` key that names the request that issued it, so keep that key when you store a conversation.

A later request sends a turn's item back when the turn holds one item and the request is configured with the same model, base URL, and credentials as the request that issued it. A request that does not match sends no items and loses only that turn's reasoning. When the request could reach another endpoint, such as through a fallback, LiteLLM's response cache, or `litellm.use_litellm_proxy`, it neither sends nor keeps items.

OpenAI rejects an item it cannot decrypt with a 400 `invalid_encrypted_content` error. The match covers only the settings `ChatLiteLLM` can read, so an item can still reach an account that cannot decrypt it through a gateway that routes to other accounts, such as a LiteLLM proxy, or through credentials LiteLLM finds on its own, such as OpenAI workload identity.

Some turns send no items back:

* Calls LiteLLM sends to the Responses API on its own, without `use_responses_api` or a `responses/` model name, such as a call to `gpt-5-pro`. Their replies keep no items.
* A streamed turn that holds more than one item. When LiteLLM does not stream a reply, it keeps only the last item before the reply's text and the last item before its tool calls, so the same turn read with `invoke` can send its item back.

On `ChatLiteLLMRouter`, items go back only when no fallback is set and every deployment of the group shares one model, base URL, set of credentials, and set of tool and reasoning settings. Name the deployments `<provider>/responses/<model>`, as in [Router deployments](#router-deployments). A deployment LiteLLM sends to the Responses API on its own, such as `openai/gpt-5-pro`, keeps no items, with or without `use_responses_api`.

#### Router deployments

The Router picks a deployment on each call, so there is no single model name for `use_responses_api` to reroute. Name each deployment's model `<provider>/responses/<model>` to send it to the Responses API:

```python theme={null}
from langchain_litellm import ChatLiteLLMRouter
from litellm import Router

model_list = [
    {
        "model_name": "gpt-6-astra",
        "litellm_params": {"model": "openai/responses/gpt-6-astra"},
    },
]
litellm_router = Router(model_list=model_list)
llm = ChatLiteLLMRouter(router=litellm_router, model_name="gpt-6-astra")
llm_with_tools = llm.bind_tools([{"type": "web_search"}])

response = llm_with_tools.invoke("What was a positive news story from today?")
print(response.text)
```

From `langchain-litellm` 0.11.0, `use_responses_api=True` on `ChatLiteLLMRouter` checks the deployments instead of rerouting them. The call goes out unchanged when LiteLLM sends every deployment of the called group to a Responses API, and raises `ValueError` before any request otherwise, naming each deployment that does not go there. It also raises when the call may reach deployments outside the group, such as through a fallback or a `model_group_alias`. Earlier versions raise `ValueError` for the flag on `ChatLiteLLMRouter`.

***

## API reference

For detailed documentation of all `ChatLiteLLM` and `ChatLiteLLMRouter` features and configurations, see the [langchain-litellm](https://reference.langchain.com/python/langchain-litellm/) API reference.

***

<div className="source-links">
  <Callout icon="terminal-2">
    [Connect these docs](/use-these-docs) to your agent of choice via MCP for real-time answers.
  </Callout>

  <Callout icon="edit">
    [Edit this page on GitHub](https://github.com/langchain-ai/docs/edit/main/src/oss/python/integrations/chat/litellm.mdx) or [file an issue](https://github.com/langchain-ai/docs/issues/new/choose).
  </Callout>
</div>
