mirror of
https://github.com/langchain-ai/langchain.git
synced 2026-10-05 09:25:14 +03:00
Added `FireworksPromptCachingMiddleware` to improve prompt-cache reuse across calls in the same agent thread. Explicit affinity settings take precedence, and no affinity is generated without a thread ID. --- Fireworks agents need consistent routing to reuse a replica's prompt cache across turns. `FireworksPromptCachingMiddleware` supplies session affinity from a SHA-256 hash of `config.configurable.thread_id`, while respecting explicit `user`, `prompt_cache_key`, and `x-session-affinity` settings on the selected model or request. Affinity is scoped to the call and applied by `ChatFireworks` when invoking the API. Generated affinity stays out of shared request settings, and model-local headers remain scoped to their owning model, including during fallback. This works with either ordering of the caching and fallback middleware for a Fireworks primary model. Related fallback cleanup for explicitly supplied cache settings is in #40886, stacked on this PR. This PR works independently of that change. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
FAQ
Looking for an integration not listed here? Check out the integrations documentation and the note in the libs/ README about third-party maintained packages.
Integration docs
For full documentation, see the primary and API reference docs for integrations.