Automated OpenWiki documentation update. OpenWiki result: success When the result is `failure`, this PR intentionally preserves only the pages completed before the failure. Merge it to make that progress the baseline for the next scheduled run. Co-authored-by: npentrel <5212232+npentrel@users.noreply.github.com>
33 KiB
type, title, description, tags, verified, sources, generated
| type | title | description | tags | verified | sources | generated | ||||||||||||||||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Architecture | Chat Model Interface and Lifecycle | Document BaseChatModel protocol, input/output handling, streaming, and integration points with callbacks and model profiling. Covers the init_chat_model() factory, provider registry, and model instantiation. |
|
|
|
|
Overview
The chat model system is the core interface for integrating large language models into LangChain applications. BaseChatModel is the abstract protocol that all chat model implementations inherit from. It defines the contract for synchronous and asynchronous invoke/streaming behavior, callback integration, rate limiting, structured output binding, and capability discovery via model profiles.
Chat models convert conversational message history into AI responses, supporting both simple generation (invoke) and streaming output (stream). The framework unifies sync/async patterns, handles caching transparently, routes to streaming or non-streaming backends based on configuration and attached callbacks, and provides extension points for custom behavior via method overrides.
The init_chat_model() factory provides a unified interface to instantiate any supported chat model by provider name (e.g., 'openai:gpt-4o'), with automatic provider detection, dependency management, and fallback strategies.
Core Interface: BaseChatModel
Location: repo://libs/core/langchain_core/language_models/chat_models.py#L284-L2400
BaseChatModel inherits from BaseLanguageModel[AIMessage] and is a Runnable that accepts LanguageModelInput and produces AIMessage outputs. It is designed for subclassing; implementations must override _generate (required) and optionally _llm_type, _stream, and _agenerate.
Input and Output Types
LanguageModelInput (repo://libs/core/langchain_core/language_models/base.py#L140) is a union type:
LanguageModelInput = PromptValue | str | Sequence[MessageLikeRepresentation]
- string: Converted to a
StringPromptValue(simple user message) - list of messages: Converted to a
ChatPromptValue(full conversation history) - PromptValue: Already a structured prompt (passed through)
The _convert_input method normalizes all input forms to a PromptValue for downstream processing.
Output: All invoke/stream methods return AIMessage or AIMessageChunk (for streaming). Chat results are wrapped in ChatGeneration objects (holding message + generation metadata) aggregated into ChatResult.
Synchronous Methods
invoke (repo://libs/core/langchain_core/language_models/chat_models.py#L474-L499) is the primary synchronous entrypoint:
def invoke(
self,
input: LanguageModelInput,
config: RunnableConfig | None = None,
*,
stop: list[str] | None = None,
**kwargs: Any,
) -> AIMessage
- Converts input to
PromptValue, then to messages - Calls
generate_prompt(which internally calls_generate_with_cache) - Extracts and returns the first generation's message
- Propagates
run_id, callbacks, tags, and metadata from config
stream (repo://libs/core/langchain_core/language_models/chat_models.py#L726-L856) yields AIMessageChunk objects as they arrive:
def stream(
self,
input: LanguageModelInput,
config: RunnableConfig | None = None,
*,
stop: list[str] | None = None,
**kwargs: Any,
) -> Iterator[AIMessageChunk]
- Checks if streaming is enabled and implemented via
_should_stream() - Falls back to
invokeif streaming is disabled or not implemented - For streaming-enabled models, calls
_stream()directly and yields chunks - Wraps output in callback lifecycle:
on_chat_model_start,on_llm_new_token(per chunk),on_llm_endoron_llm_error - Applies rate limiting if configured
- Normalizes messages and handles streaming-specific output formatting (e.g.,
output_version="v1") - Yields a final empty chunk with
chunk_position="last"when streaming completes
Asynchronous Methods
ainvoke (repo://libs/core/langchain_core/language_models/chat_models.py#L501-L523) is the async variant:
async def ainvoke(
self,
input: LanguageModelInput,
config: RunnableConfig | None = None,
*,
stop: list[str] | None = None,
**kwargs: Any,
) -> AIMessage
- Awaits
agenerate_prompt - Otherwise mirrors
invokebehavior
astream (repo://libs/core/langchain_core/language_models/chat_models.py#L857-L990) is the async streaming variant:
async def astream(
self,
input: LanguageModelInput,
config: RunnableConfig | None = None,
*,
stop: list[str] | None = None,
**kwargs: Any,
) -> AsyncIterator[AIMessageChunk]
- Checks
_should_stream(async_api=True)to route to_astreamor fallback - Otherwise mirrors
streambehavior with async callback dispatch
Streaming Architecture
Stream Decision Logic
_should_stream() (repo://libs/core/langchain_core/language_models/chat_models.py#L549-L585) determines whether to use the streaming code path:
def _should_stream(
self,
*,
async_api: bool,
run_manager: CallbackManagerForLLMRun | AsyncCallbackManagerForLLMRun | None = None,
**kwargs: Any,
) -> bool
Returns True if:
- Streaming is not disabled (
_streaming_disabled()returnsFalse) - Streaming method is implemented for the requested variant (sync/async)
- Any of these are true:
- Explicit
stream=Truekwarg - Instance-level
streaming=Trueattribute - A v1-style
_StreamingCallbackHandleris attached
- Explicit
Returns False (fallback to non-streaming) if:
disable_streaming=True(hard disable)disable_streaming="tool_calling"and tools are providedstream=Falseexplicitly- Streaming is not implemented and async falls back to sync
Stream Implementation Methods
_stream() (repo://libs/core/langchain_core/language_models/chat_models.py#L2255-L2273) is the sync streaming hook (optional override):
def _stream(
self,
messages: list[BaseMessage],
stop: list[str] | None = None,
run_manager: CallbackManagerForLLMRun | None = None,
**kwargs: Any,
) -> Iterator[ChatGenerationChunk]
- Subclasses override to implement native streaming
- Default raises
NotImplementedError(fallback to_generate) - Receives run_manager for per-token callbacks
_astream() (repo://libs/core/langchain_core/language_models/chat_models.py#L2275-L2311) is the async streaming hook (optional override):
async def _astream(
self,
messages: list[BaseMessage],
stop: list[str] | None = None,
run_manager: AsyncCallbackManagerForLLMRun | None = None,
**kwargs: Any,
) -> AsyncIterator[ChatGenerationChunk]
- Default implementation runs
_stream()in an executor and yields results - Subclasses can override for native async streaming
ChatModelStream and AsyncChatModelStream
Location: repo://libs/core/langchain_core/language_models/chat_model_stream.py
For the v3 event protocol (stream_events(version="v3")), models return a ChatModelStream (sync) or AsyncChatModelStream (async) that expose typed projections for incremental content:
.text: Accumulates text content blocks.reasoning: Accumulates reasoning/chain-of-thought content.tool_calls: Accumulates parsed tool call blocks.usage: Accumulates token usage info.output: Final assembledAIMessage
Each projection can be iterated for deltas or awaited for the final value. Internally, these accumulators track incoming protocol events and merge them into structured output.
Generation and Caching
Core Generation Methods
_generate() (repo://libs/core/langchain_core/language_models/chat_models.py#L2208-L2226) is the required abstract method all subclasses must implement:
@abstractmethod
def _generate(
self,
messages: list[BaseMessage],
stop: list[str] | None = None,
run_manager: CallbackManagerForLLMRun | None = None,
**kwargs: Any,
) -> ChatResult
- Calls the underlying model's API
- Returns a
ChatResultwith a list ofChatGenerationobjects - Must handle errors internally or propagate them
- Receives normalized messages and a run manager for callbacks
_agenerate() (repo://libs/core/langchain_core/language_models/chat_models.py#L2228-L2253) is the optional async override:
async def _agenerate(
self,
messages: list[BaseMessage],
stop: list[str] | None = None,
run_manager: AsyncCallbackManagerForLLMRun | None = None,
**kwargs: Any,
) -> ChatResult
- Default implementation runs
_generatein an executor - Subclasses override for native async API support
Cached Generation
_generate_with_cache() and _agenerate_with_cache() wrap the core methods with:
- Prompt caching: Checks if
self.cacheor globalget_llm_cache()has cached results for the input - Cache hits: Returns cached generations and replays them as v2 events if a v2 handler is attached
- Cache misses: Routes through streaming or non-streaming path
- Protocol routing: Dispatches to v2 events (
_should_use_protocol_streaming) or v1 callback path (_should_stream)
Batch Methods
generate() and agenerate() accept a list of message lists and use internal caching/streaming to batch-process prompts:
def generate(
self,
messages: list[list[BaseMessage]],
stop: list[str] | None = None,
callbacks: Callbacks = None,
**kwargs: Any,
) -> LLMResult
Returns an LLMResult with generations grouped by input prompt and combined llm_output.
Callback Lifecycle
Chat models integrate with the callback system to emit structured events throughout execution:
LLM Run Lifecycle
-
on_chat_model_start(or fallbackon_llm_start):- Fires when
invoke,stream, orgeneratebegins - Receives serialized model config, formatted input messages, invocation params, and batch size
- Returns run manager(s) bound to the operation
- Fires when
-
on_llm_new_token(streaming only):- Fires once per streamed token/chunk
- Receives token string and
ChatGenerationChunkmetadata - Allows real-time output capture
-
on_llm_end:- Fires when generation completes successfully
- Receives final
LLMResultwith all generations and metadata
-
on_llm_error:- Fires if generation raises an exception
- Receives the exception and partial
LLMResult(if available) _generate_response_from_error()extracts response metadata from HTTP errors
-
on_stream_event(v2/v3 protocol):- Fires for each content-block protocol event during streaming
- Allows fine-grained event observation for advanced tracing
Callback Configuration
Callbacks are configured via RunnableConfig:
config = {
"callbacks": [my_handler], # Callbacks for this run
"tags": ["agent", "tools"], # Labels for filtering
"metadata": {"user_id": "123"}, # Context data
"run_name": "my_run", # Human-readable run name
"run_id": uuid.uuid4(), # Explicit run ID (optional)
}
result = model.invoke(input, config=config)
Inheritable metadata and LangSmith params are extracted via _get_invocation_params() and _get_ls_params().
Structured Output and Tool Binding
with_structured_output()
Location: repo://libs/core/langchain_core/language_models/chat_models.py#L2385-L2565
with_structured_output() wraps a chat model to constrain output to a specified schema:
def with_structured_output(
self,
schema: dict[str, Any] | type,
*,
include_raw: bool = False,
**kwargs: Any,
) -> Runnable[LanguageModelInput, dict[str, Any] | BaseModel]
How it works:
- Delegates to
bind_tools([schema], tool_choice="any", ...) - Chains the result through an output parser:
- If schema is a Pydantic class:
PydanticToolsParser→ Pydantic instance - If schema is a dict:
JsonOutputKeyToolsParser→ dict
- If schema is a Pydantic class:
- If
include_raw=True: Wraps output in{"raw": AIMessage, "parsed": ..., "parsing_error": ...} - If parsing fails and
include_raw=False: Raises exception
Prerequisites: Requires the model to implement bind_tools() (not all models support this).
bind_tools()
Location: repo://libs/core/langchain_core/language_models/chat_models.py#L2366-L2383
def bind_tools(
self,
tools: Sequence[dict[str, Any] | type | Callable[..., Any] | BaseTool],
*,
tool_choice: str | None = None,
**kwargs: Any,
) -> Runnable[LanguageModelInput, AIMessage]
- Abstract method; must be implemented by subclasses that support tool calling
- Binds a list of tools to the model
- Returns a bound runnable that includes tool definitions in the API request
tool_choice="any"forces the model to call at least one tool
Model Profiles and Capabilities
Location: repo://libs/core/langchain_core/language_models/model_profile.py
The profile field on BaseChatModel holds metadata about model capabilities:
class ModelProfile(TypedDict, total=False):
# Metadata
name: str # Human-readable model name
status: str # 'active', 'deprecated', etc.
release_date: str # ISO 8601
last_updated: str # ISO 8601
open_weights: bool # Weights publicly available?
# Input constraints
max_input_tokens: int # Context window size
text_inputs: bool
image_inputs: bool
image_url_inputs: bool
pdf_inputs: bool
audio_inputs: bool
video_inputs: bool
image_tool_message: bool # Images in ToolMessage?
pdf_tool_message: bool # PDFs in ToolMessage?
# Output constraints
max_output_tokens: int
text_outputs: bool
image_outputs: bool
audio_outputs: bool
video_outputs: bool
# Capabilities
tool_calling: bool # Supports function calling?
tool_choice: bool # Supports tool_choice parameter?
tool_call_streaming: bool # Returns structured tool_call_chunks when streaming?
structured_output: bool # Native structured output support?
reasoning_output: bool # Reasoning/chain-of-thought?
reasoning_effort_levels: list[str] # ['low', 'medium', 'high']
reasoning_effort_default: str
temperature: bool # Supports temperature parameter?
attachment: bool # Supports file attachments?
Auto-loading: Profiles are resolved via _resolve_model_profile() (subclass override) and cached in the profile field. Unrecognized keys trigger a warning via _warn_unknown_profile_keys().
Partner Pattern Integration
Partner packages (e.g., langchain-openai) override _resolve_model_profile() to load model-specific metadata from their own profile data. The base validator _set_model_profile (Pydantic mode="after") automatically populates the field if not explicitly set.
Model Initialization: init_chat_model()
Location: repo://libs/langchain_v1/langchain/chat_models/base.py#L197-L533
init_chat_model() is a unified factory function that instantiates any supported chat model by provider name, with automatic dependency management, provider inference, and fallback strategies.
Basic Usage
from langchain.chat_models import init_chat_model
# Fixed model with explicit provider
model = init_chat_model("openai:gpt-4o", temperature=0.7)
# Bare model name with inferred provider
model = init_chat_model("claude-3-sonnet", api_key="...")
# Configurable model (deferred instantiation)
configurable = init_chat_model(configurable_fields=("model", "model_provider"))
result = configurable.invoke("Hello", config={"configurable": {"model": "gpt-4o"}})
Function Signature
def init_chat_model(
model: str | None = None,
*,
model_provider: str | None = None,
configurable_fields: Literal["any"] | list[str] | tuple[str, ...] | None = None,
config_prefix: str | None = None,
**kwargs: Any,
) -> BaseChatModel | _ConfigurableModel
Parameters:
model: Model name string with optional provider prefix (e.g.,'openai:gpt-4o'or bare'gpt-4o'). IfNoneandconfigurable_fieldsis unset, defaults to configurable mode.model_provider: Provider name if not inferred frommodelprefix. Examples:'openai','anthropic','bedrock','google_vertexai'.configurable_fields: Which parameters are runtime-configurable:None: Fixed model (default ifmodelis specified)"any": All fields configurable (security risk with untrusted configs)list[str] | tuple[str, ...]: Specified fields only (e.g.,["model", "temperature"])- Default if
modelisNone:("model", "model_provider")
config_prefix: Prefix for config keys (e.g.,"llm"→config["configurable"]["llm_model"])**kwargs: Model-specific parameters (temperature,max_tokens,api_key,base_url, etc.)
Returns:
BaseChatModel: A ready-to-use chat model ifmodelis specified andconfigurable_fieldsisNone_ConfigurableModel: A deferred-initialization runnable if fields are configurable; calls_init_chat_model_helper()at invoke time
Provider Inference
When model is specified without a provider prefix, _attempt_infer_model_provider() guesses the provider from model name patterns:
| Pattern | Provider | Example |
|---|---|---|
gpt-..., o1..., o3... |
openai |
gpt-4o, o1-pro |
claude... |
anthropic |
claude-3-sonnet |
command... |
cohere |
command-light |
accounts/fireworks... |
fireworks |
accounts/fireworks/models/mixtral-8x7b |
mistral..., mixtral... |
mistralai |
mistral-large |
deepseek... |
deepseek |
deepseek-chat |
grok... |
xai |
grok-2 |
sonar... |
perplexity |
sonar-pro |
solar... |
upstage |
solar-1-mini |
gemini... |
google_vertexai (default) |
gemini-2.0-flash |
amazon..., anthropic..., meta... |
bedrock |
amazon.nova-pro |
Prefer explicit provider prefix form (e.g., 'openai:gpt-4o') when possible to avoid inference ambiguity.
Built-in Provider Registry
Location: repo://libs/langchain_v1/langchain/chat_models/base.py#L56-L99
The _BUILTIN_PROVIDERS dict maps provider names to their import paths and instantiation functions:
_BUILTIN_PROVIDERS: dict[str, tuple[str, str, Callable[..., BaseChatModel]]] = {
"anthropic": ("langchain_anthropic", "ChatAnthropic", _call),
"anthropic_bedrock": ("langchain_aws", "ChatAnthropicBedrock", _call),
"azure_ai": ("langchain_azure_ai.chat_models", "AzureAIOpenAIApiChatModel", _call),
"azure_openai": ("langchain_openai", "AzureChatOpenAI", _call),
"baseten": ("langchain_baseten", "ChatBaseten", _call),
"bedrock": ("langchain_aws", "ChatBedrock", _call),
"bedrock_converse": ("langchain_aws", "ChatBedrockConverse", _call),
"bedrock_mantle_anthropic": ("langchain_aws", "ChatAnthropicMantle", _call),
"bedrock_mantle_openai": ("langchain_aws", "ChatOpenAIMantle", _call),
"cohere": ("langchain_cohere", "ChatCohere", _call),
"deepseek": ("langchain_deepseek", "ChatDeepSeek", _call),
"fireworks": ("langchain_fireworks", "ChatFireworks", _call),
"google_anthropic_vertex": ("langchain_google_vertexai.model_garden", "ChatAnthropicVertex", _call),
"google_genai": ("langchain_google_genai", "ChatGoogleGenerativeAI", _call),
"google_vertexai": ("langchain_google_vertexai", "ChatVertexAI", _call),
"groq": ("langchain_groq", "ChatGroq", _call),
"huggingface": ("langchain_huggingface", "ChatHuggingFace", lambda cls, model, **kwargs: cls.from_model_id(...)),
"ibm": ("langchain_ibm", "ChatWatsonx", lambda cls, model, **kwargs: cls(model_id=model, ...)),
"langsmith": ("langchain_openai", "ChatOpenAI", _init_langsmith),
"litellm": ("langchain_litellm", "ChatLiteLLM", _call),
"meta": ("langchain_meta", "ChatMetaModel", _call),
"mistralai": ("langchain_mistralai", "ChatMistralAI", _call),
"nvidia": ("langchain_nvidia_ai_endpoints", "ChatNVIDIA", _call),
"ollama": ("langchain_ollama", "ChatOllama", _call),
"openai": ("langchain_openai", "ChatOpenAI", _call),
"openrouter": ("langchain_openrouter", "ChatOpenRouter", _call),
"perplexity": ("langchain_perplexity", "ChatPerplexity", _call),
"together": ("langchain_together", "ChatTogether", _call),
"upstage": ("langchain_upstage", "ChatUpstage", _call),
"xai": ("langchain_xai", "ChatXAI", _call),
}
Each entry is a tuple: (module_path, class_name, creator_func):
module_path: Fully qualified Python module (e.g.,'langchain_openai'or'langchain_azure_ai.chat_models')class_name: Name of the model class to instantiatecreator_func: Callable that acceptsclsand model kwargs and returns aBaseChatModelinstance (usually_callor a custom wrapper)
Dynamic Loading and Error Handling
_get_chat_model_creator() (repo://libs/langchain_v1/langchain/chat_models/base.py#L152-L193) retrieves and caches the creator function:
- Validates that
provideris in_BUILTIN_PROVIDERS - Imports the module via
_import_module() - Retrieves the class from the module
- Returns a partial function bound to the class
- Caches result to avoid repeated imports
_import_module() wraps importlib.import_module() with helpful error messages:
- If import fails: Raises
ImportErrorwith suggestion topip install <package> - Extracts package name from module path (e.g.,
'langchain_openai'→'langchain-openai') - Message format:
"Initializing {class} requires the {pkg} package. Please install it with 'pip install {pkg}'"
Fallback for backwards compatibility:
ollamaprovider attemptslangchain-ollamafirst, then falls back tolangchain-communityfor legacy users- Other providers raise immediately if their package is not installed
Configurable Models
When configurable_fields is specified, init_chat_model() returns a _ConfigurableModel instance that defers model instantiation:
# Fixed model instantiated immediately
fixed = init_chat_model("gpt-4o", temperature=0.5)
# Configurable model deferred until invoke
configurable = init_chat_model(
"gpt-4o",
configurable_fields=("temperature", "max_tokens"),
config_prefix="llm"
)
# Actual instantiation happens here with runtime config
result = configurable.invoke(
"Hello",
config={
"configurable": {
"llm_temperature": 0.7, # Override default
"llm_max_tokens": 100,
}
}
)
_ConfigurableModel implements:
invoke(),ainvoke(),stream(),astream(): Call_model(config)to instantiate the underlying model, then delegatebatch(),abatch(): Use underlying model's batch if single config, else use Runnable.batch (parallelized)- Declarative methods:
bind_tools(),with_structured_output()queue operations and apply them after model instantiation - Config merging:
with_config()applies per-invocation config and returns a new_ConfigurableModelwith merged defaults
Model-Specific Initialization Wrappers
Some providers require custom instantiation:
huggingface: UsesChatHuggingFace.from_model_id(model_id=model, **kwargs)instead of__init__ibm: UsesChatWatsonx(model_id=model, **kwargs)(parameter name differs from standardmodel)langsmith: Uses_init_langsmith()wrapper that applies LangSmith gateway config and setsuse_responses_api=True
Error Handling and Fallback Strategies
| Scenario | Behavior | Recovery |
|---|---|---|
| Provider not in registry | Raises ValueError with list of supported providers |
Pass model_provider explicitly if using an unlisted provider |
| Module import fails | Raises ImportError with pip install suggestion |
Install the required integration package |
| Class not found in module | Raises AttributeError (implicit) |
Verify provider is correctly registered |
| Invalid kwargs | Raises TypeError or ValidationError from model class |
Check model-specific parameter names and types |
| Missing API credentials | Model init succeeds; call fails at invoke time | Set env vars or pass credentials in kwargs |
Configuration and State
Core Fields
class BaseChatModel(BaseLanguageModel[AIMessage], ABC):
rate_limiter: BaseRateLimiter | None = Field(default=None, exclude=True)
disable_streaming: bool | Literal["tool_calling"] = False
# False: use streaming if available
# True: always use non-streaming (invoke)
# "tool_calling": use non-streaming only when tools are passed
output_version: str | None = None
# 'v0': provider-specific format (lazy-parse via content_blocks)
# 'v1': standardized format (merged into content)
profile: ModelProfile | None = Field(default=None, exclude=True)
# Capability metadata (auto-loaded if not provided)
cache: BaseCache | None = None # Inherited from BaseLanguageModel
callbacks: list[BaseCallbackHandler] | None = None
verbose: bool = False
tags: list[str] | None = None
metadata: dict[str, Any] | None = None
Required Properties
_llm_type(property, abstract): Unique model type identifier (e.g.,"openai","anthropic")_identifying_params(property, optional): Dict of model configuration for tracing (e.g.,{"model": "gpt-4", "temperature": 0.7})
Token Counting
Location: repo://libs/core/langchain_core/language_models/base.py#L434-L463
BaseLanguageModel provides token counting via get_token_ids() and get_num_tokens():
def get_token_ids(self, text: str) -> list[int]:
"""Return token IDs for the given text."""
if self.custom_get_token_ids is not None:
return self.custom_get_token_ids(text)
return _get_token_ids_default_method(text) # GPT-2 fallback
def get_num_tokens(self, text: str) -> int:
"""Return token count for the given text."""
return len(self.get_token_ids(text))
Default behavior:
- Uses GPT-2 tokenizer as fallback (requires
transformerspackage) - Warns once that counts may be inaccurate for non-GPT-2 models
- Custom implementations can override
custom_get_token_idsfield for model-specific accuracy
Partner package support:
- Implementations (e.g.,
ChatOpenAI) overrideget_token_ids()to use native tokenizers - Ensures accurate counts for pricing and context-window management
Implementation Requirements
Subclasses must implement:
| Method/Property | Description | Required | Notes |
|---|---|---|---|
_generate() |
Core generation logic | ✓ | Calls provider API, returns ChatResult |
_llm_type |
Model type identifier | ✓ | String like "openai", "anthropic" |
_identifying_params |
Config dict for tracing | ✗ | Used by _get_llm_string() and serialization |
_stream() |
Sync streaming | ✗ | Optional; if not implemented, stream falls back to invoke |
_agenerate() |
Native async generation | ✗ | Optional; defaults to running _generate in executor |
_astream() |
Native async streaming | ✗ | Optional; defaults to running _stream in executor |
bind_tools() |
Tool binding for structured output | ✗ | Required only if with_structured_output() is needed |
Example: Custom Chat Model
from langchain_core.language_models.chat_models import BaseChatModel
from langchain_core.messages import BaseMessage, AIMessage
from langchain_core.outputs import ChatResult, ChatGeneration
from langchain_core.callbacks import CallbackManagerForLLMRun
class MyCustomChatModel(BaseChatModel):
"""Custom chat model for demonstration."""
model_name: str = "my-model"
temperature: float = 0.7
def _generate(
self,
messages: list[BaseMessage],
stop: list[str] | None = None,
run_manager: CallbackManagerForLLMRun | None = None,
**kwargs: Any,
) -> ChatResult:
"""Generate a response from the messages."""
# Call your model API here
response_text = f"Echo: {messages[-1].content}"
message = AIMessage(content=response_text)
generation = ChatGeneration(message=message)
return ChatResult(generations=[generation])
def _stream(
self,
messages: list[BaseMessage],
stop: list[str] | None = None,
run_manager: CallbackManagerForLLMRun | None = None,
**kwargs: Any,
) -> Iterator[ChatGenerationChunk]:
"""Stream tokens from the model."""
text = f"Echo: {messages[-1].content}"
for char in text:
chunk = ChatGenerationChunk(
message=AIMessageChunk(content=char)
)
yield chunk
@property
def _llm_type(self) -> str:
"""Return the model type identifier."""
return "my-custom-model"
@property
def _identifying_params(self) -> dict[str, Any]:
"""Return identifying parameters for tracing."""
return {
"model_name": self.model_name,
"temperature": self.temperature,
}
Advanced Patterns
Streaming with Callbacks
from langchain_core.callbacks import StreamingStdOutCallbackHandler
handler = StreamingStdOutCallbackHandler()
config = {"callbacks": [handler]}
# Streams token-by-token to stdout
for chunk in model.stream("Tell me a joke", config=config):
pass # Handler prints as chunks arrive
Structured Output with Validation
from pydantic import BaseModel
class Answer(BaseModel):
text: str
confidence: float
structured_model = model.with_structured_output(Answer)
result = structured_model.invoke("What is 2+2?") # -> Answer(text="4", confidence=0.99)
Caching and Rate Limiting
from langchain_core.caches import InMemoryCache
from langchain_core.rate_limiters import InMemoryRateLimiter
model = ChatOpenAI(
model="gpt-4",
cache=InMemoryCache(), # Cache results
rate_limiter=InMemoryRateLimiter(requests_per_second=10) # Limit requests
)
# Subsequent identical calls hit the cache
result1 = model.invoke("Hello")
result2 = model.invoke("Hello") # Cached, no API call
Conditional Streaming
model_with_fallback = ChatOpenAI().with_fallbacks([ChatAnthropic()])
# Use streaming only when a handler requests it
config = {"callbacks": [MyStreamingHandler()]}
model_with_fallback.invoke("Prompt", config=config)
Using init_chat_model() in Agents
from langchain.chat_models import init_chat_model
from langchain.agents import create_tool_calling_agent
from langchain.agents import AgentExecutor
# Create a configurable model that can be swapped at runtime
llm = init_chat_model(
configurable_fields=("model", "model_provider"),
temperature=0.7
)
# Build agent with the model
agent = create_tool_calling_agent(llm, tools, prompt)
executor = AgentExecutor(agent=agent, tools=tools)
# Default with OpenAI
result = executor.invoke({"input": "What is the weather?"})
# Switch to Anthropic at runtime
result = executor.invoke(
{"input": "What is the weather?"},
config={"configurable": {"model": "claude-3-sonnet"}}
)
Key Invariants and Guarantees
-
Input normalization: All input forms (string, message list, PromptValue) are normalized to messages before
_generate/_streamare called. -
Message IDs: Each streamed message chunk and final message gets a unique ID (derived from run_id) for tracing.
-
Callback ordering: Callbacks fire in order:
on_chat_model_start→on_llm_new_token(per chunk) →on_llm_endoron_llm_error. -
Streaming fallback: If streaming is not implemented or disabled,
streamseamlessly falls back toinvokeand yields the result as a single chunk. -
Cache transparency: Cache hits are completely transparent—same lifecycle callbacks fire as for cache misses.
-
Async/sync equivalence: Async methods mirror sync behavior; default async implementations run sync methods in an executor.
-
Response metadata: Each generation accumulates metadata (tokens, finish_reason, etc.) in
message.response_metadata. -
Error handling: Exceptions during generation trigger
on_llm_errorand propagate to the caller; error metadata is extracted from HTTP responses if available. -
Provider isolation: Each provider integration is independently installed and loaded; missing packages are reported with clear
pip installinstructions. -
Configurable deferred binding: Declarative methods (
bind_tools,with_structured_output) on configurable models queue operations and apply them at instantiation time, ensuring consistent behavior across runtime model swaps.