Files
langchain/docs/extras/expression_language/interface.ipynb
T
Nuno Campos 1cbe7f5450 Small changes to runnable docs (#11293)
<!-- Thank you for contributing to LangChain!

Replace this entire comment with:
  - **Description:** a description of the change, 
  - **Issue:** the issue # it fixes (if applicable),
  - **Dependencies:** any dependencies required for this change,
- **Tag maintainer:** for a quicker response, tag the relevant
maintainer (see below),
- **Twitter handle:** we announce bigger features on Twitter. If your PR
gets announced, and you'd like a mention, we'll gladly shout you out!

Please make sure your PR is passing linting and testing before
submitting. Run `make format`, `make lint` and `make test` to check this
locally.

See contribution guidelines for more information on how to write/run
tests, lint, etc:

https://github.com/langchain-ai/langchain/blob/master/.github/CONTRIBUTING.md

If you're adding a new integration, please include:
1. a test for the integration, preferably unit tests that do not rely on
network access,
2. an example notebook showing its use. It lives in `docs/extras`
directory.

If no one reviews your PR within a few days, please @-mention one of
@baskaryan, @eyurtsev, @hwchase17.
 -->
2023-10-02 16:27:11 +01:00

11 KiB

Cell:
[Cell type raw - unsupported, skipped]

In an effort to make it as easy as possible to create custom chains, we've implemented a "Runnable" protocol that most components implement. This is a standard interface with a few different methods, which makes it easy to define custom chains as well as making it possible to invoke them in a standard way. The standard interface exposed includes:

  • stream: stream back chunks of the response
  • invoke: call the chain on an input
  • batch: call the chain on a list of inputs

These also have corresponding async methods:

  • astream: stream back chunks of the response async
  • ainvoke: call the chain on an input async
  • abatch: call the chain on a list of inputs async

The type of the input varies by component:

Component Input Type
Prompt Dictionary
Retriever Single string
LLM, ChatModel Single string, list of chat messages or a PromptValue
Tool Single string, or dictionary, depending on the tool
OutputParser The output of an LLM or ChatModel

The output type also varies by component:

Component Output Type
LLM String
ChatModel ChatMessage
Prompt PromptValue
Retriever List of documents
Tool Depends on the tool
OutputParser Depends on the parser

Let's take a look at these methods! To do so, we'll create a super simple PromptTemplate + ChatModel chain.

In [1]:
from langchain.prompts import ChatPromptTemplate
from langchain.chat_models import ChatOpenAI
In [2]:
model = ChatOpenAI()
In [3]:
prompt = ChatPromptTemplate.from_template("tell me a joke about {topic}")
In [4]:
chain = prompt | model

Stream

In [8]:
for s in chain.stream({"topic": "bears"}):
    print(s.content, end="", flush=True)
Sure, here's a bear-themed joke for you:

Why don't bears wear shoes?

Because they have bear feet!

Invoke

In [9]:
chain.invoke({"topic": "bears"})
Out [9]:
AIMessage(content="Why don't bears wear shoes?\n\nBecause they already have bear feet!", additional_kwargs={}, example=False)

Batch

In [19]:
chain.batch([{"topic": "bears"}, {"topic": "cats"}])
Out [19]:
[AIMessage(content="Why don't bears ever wear shoes?\n\nBecause they have bear feet!", additional_kwargs={}, example=False),
 AIMessage(content="Why don't cats play poker in the wild?\n\nToo many cheetahs!", additional_kwargs={}, example=False)]

You can set the number of concurrent requests by using the max_concurrency parameter

In [5]:
chain.batch([{"topic": "bears"}, {"topic": "cats"}], config={"max_concurrency": 5})
Out [5]:
[AIMessage(content="Why don't bears wear shoes?\n\nBecause they have bear feet!", additional_kwargs={}, example=False),
 AIMessage(content="Why don't cats play poker in the wild?\n\nToo many cheetahs!", additional_kwargs={}, example=False)]

Async Stream

In [13]:
async for s in chain.astream({"topic": "bears"}):
    print(s.content, end="", flush=True)
Why don't bears wear shoes?

Because they have bear feet!

Async Invoke

In [16]:
await chain.ainvoke({"topic": "bears"})
Out [16]:
AIMessage(content="Sure, here you go:\n\nWhy don't bears wear shoes?\n\nBecause they have bear feet!", additional_kwargs={}, example=False)

Async Batch

In [18]:
await chain.abatch([{"topic": "bears"}])
Out [18]:
[AIMessage(content="Why don't bears wear shoes?\n\nBecause they have bear feet!", additional_kwargs={}, example=False)]

Parallelism

Let's take a look at how LangChain Expression Language support parallel requests as much as possible. For example, when using a RunnableMap (often written as a dictionary) it executes each element in parallel.

In [7]:
from langchain.schema.runnable import RunnableMap
chain1 = ChatPromptTemplate.from_template("tell me a joke about {topic}") | model
chain2 = ChatPromptTemplate.from_template("write a short (2 line) poem about {topic}") | model
combined = RunnableMap({
    "joke": chain1,
    "poem": chain2,
})
In [11]:
%%time
chain1.invoke({"topic": "bears"})
Out [11]:
CPU times: user 31.7 ms, sys: 8.59 ms, total: 40.3 ms
Wall time: 1.05 s
AIMessage(content="Why don't bears like fast food?\n\nBecause they can't catch it!", additional_kwargs={}, example=False)
In [12]:
%%time
chain2.invoke({"topic": "bears"})
Out [12]:
CPU times: user 42.9 ms, sys: 10.2 ms, total: 53 ms
Wall time: 1.93 s
AIMessage(content="In forest's embrace, bears roam free,\nSilent strength, nature's majesty.", additional_kwargs={}, example=False)
In [13]:
%%time
combined.invoke({"topic": "bears"})
Out [13]:
CPU times: user 96.3 ms, sys: 20.4 ms, total: 117 ms
Wall time: 1.1 s
{'joke': AIMessage(content="Why don't bears wear socks?\n\nBecause they have bear feet!", additional_kwargs={}, example=False),
 'poem': AIMessage(content="In forest's embrace,\nMajestic bears leave their trace.", additional_kwargs={}, example=False)}
In [ ]: