mirror of
https://github.com/langchain-ai/langchain.git
synced 2026-10-09 03:15:19 +03:00
+26









Erick Friis
Harrison Chase
jacoblee93
Leonid Ganeline
Leonid Kuligin
Averi Kitsch
Nuno Campos
Nuno Campos
Bagatur
Eugene Yurtsev
Martín Gotelli Ferenaz
Fayfox
Eugene Yurtsev
Dawson Bauer
Ravindu Somawansa
Dhruv Chawla
ccurme
Bagatur
WeichenXu
Benito Geordie
kartikTAI
Kartik Sarangmath
Sevin F. Varoglu
MacanPN
Prashanth Rao
Hyeongchan Kim
sdan
Guangdong Liu
Rahul Triptahi
Rahul Tripathi
pjb157
Eun Hye Kim
kaijietti
Pengcheng Liu
Tomer Cagan
Christophe Bornet
21d14549a9
current python.langchain.com is building from branch `v0.1`. Iterate on v0.2 docs here. --------- Signed-off-by: Weichen Xu <weichen.xu@databricks.com> Signed-off-by: Rahul Tripathi <rauhl.psit.ec@gmail.com> Co-authored-by: Harrison Chase <hw.chase.17@gmail.com> Co-authored-by: jacoblee93 <jacoblee93@gmail.com> Co-authored-by: Leonid Ganeline <leo.gan.57@gmail.com> Co-authored-by: Leonid Kuligin <lkuligin@yandex.ru> Co-authored-by: Averi Kitsch <akitsch@google.com> Co-authored-by: Nuno Campos <nuno@langchain.dev> Co-authored-by: Nuno Campos <nuno@boringbits.io> Co-authored-by: Bagatur <22008038+baskaryan@users.noreply.github.com> Co-authored-by: Eugene Yurtsev <eyurtsev@gmail.com> Co-authored-by: Martín Gotelli Ferenaz <martingotelliferenaz@gmail.com> Co-authored-by: Fayfox <admin@fayfox.com> Co-authored-by: Eugene Yurtsev <eugene@langchain.dev> Co-authored-by: Dawson Bauer <105886620+djbauer2@users.noreply.github.com> Co-authored-by: Ravindu Somawansa <ravindu.somawansa@gmail.com> Co-authored-by: Dhruv Chawla <43818888+Dominastorm@users.noreply.github.com> Co-authored-by: ccurme <chester.curme@gmail.com> Co-authored-by: Bagatur <baskaryan@gmail.com> Co-authored-by: WeichenXu <weichen.xu@databricks.com> Co-authored-by: Benito Geordie <89472452+benitoThree@users.noreply.github.com> Co-authored-by: kartikTAI <129414343+kartikTAI@users.noreply.github.com> Co-authored-by: Kartik Sarangmath <kartik@thirdai.com> Co-authored-by: Sevin F. Varoglu <sfvaroglu@octoml.ai> Co-authored-by: MacanPN <martin.triska@gmail.com> Co-authored-by: Prashanth Rao <35005448+prrao87@users.noreply.github.com> Co-authored-by: Hyeongchan Kim <kozistr@gmail.com> Co-authored-by: sdan <git@sdan.io> Co-authored-by: Guangdong Liu <liugddx@gmail.com> Co-authored-by: Rahul Triptahi <rahul.psit.ec@gmail.com> Co-authored-by: Rahul Tripathi <rauhl.psit.ec@gmail.com> Co-authored-by: pjb157 <84070455+pjb157@users.noreply.github.com> Co-authored-by: Eun Hye Kim <ehkim1440@gmail.com> Co-authored-by: kaijietti <43436010+kaijietti@users.noreply.github.com> Co-authored-by: Pengcheng Liu <pcliu.fd@gmail.com> Co-authored-by: Tomer Cagan <tomer@tomercagan.com> Co-authored-by: Christophe Bornet <cbornet@hotmail.com>
129 lines
3.9 KiB
Plaintext
129 lines
3.9 KiB
Plaintext
# Text embedding models
|
|
|
|
:::info
|
|
Head to [Integrations](/docs/integrations/text_embedding/) for documentation on built-in integrations with text embedding model providers.
|
|
:::
|
|
|
|
The Embeddings class is a class designed for interfacing with text embedding models. There are lots of embedding model providers (OpenAI, Cohere, Hugging Face, etc) - this class is designed to provide a standard interface for all of them.
|
|
|
|
Embeddings create a vector representation of a piece of text. This is useful because it means we can think about text in the vector space, and do things like semantic search where we look for pieces of text that are most similar in the vector space.
|
|
|
|
The base Embeddings class in LangChain provides two methods: one for embedding documents and one for embedding a query. The former, `.embed_documents`, takes as input multiple texts, while the latter, `.embed_query`, takes a single text. The reason for having these as two separate methods is that some embedding providers have different embedding methods for documents (to be searched over) vs queries (the search query itself).
|
|
`.embed_query` will return a list of floats, whereas `.embed_documents` returns a list of lists of floats.
|
|
|
|
## Get started
|
|
|
|
### Setup
|
|
|
|
import Tabs from '@theme/Tabs';
|
|
import TabItem from '@theme/TabItem';
|
|
|
|
<Tabs>
|
|
<TabItem value="openai" label="OpenAI" default>
|
|
To start we'll need to install the OpenAI partner package:
|
|
|
|
```bash
|
|
pip install langchain-openai
|
|
```
|
|
|
|
Accessing the API requires an API key, which you can get by creating an account and heading [here](https://platform.openai.com/account/api-keys). Once we have a key we'll want to set it as an environment variable by running:
|
|
|
|
```bash
|
|
export OPENAI_API_KEY="..."
|
|
```
|
|
|
|
If you'd prefer not to set an environment variable you can pass the key in directly via the `api_key` named parameter when initiating the OpenAI LLM class:
|
|
|
|
```python
|
|
from langchain_openai import OpenAIEmbeddings
|
|
|
|
embeddings_model = OpenAIEmbeddings(api_key="...")
|
|
```
|
|
|
|
Otherwise you can initialize without any params:
|
|
```python
|
|
from langchain_openai import OpenAIEmbeddings
|
|
|
|
embeddings_model = OpenAIEmbeddings()
|
|
```
|
|
|
|
</TabItem>
|
|
<TabItem value="cohere" label="Cohere">
|
|
|
|
To start we'll need to install the Cohere SDK package:
|
|
|
|
```bash
|
|
pip install langchain-cohere
|
|
```
|
|
|
|
Accessing the API requires an API key, which you can get by creating an account and heading [here](https://dashboard.cohere.com/api-keys). Once we have a key we'll want to set it as an environment variable by running:
|
|
|
|
```shell
|
|
export COHERE_API_KEY="..."
|
|
```
|
|
|
|
If you'd prefer not to set an environment variable you can pass the key in directly via the `cohere_api_key` named parameter when initiating the Cohere LLM class:
|
|
|
|
```python
|
|
from langchain_cohere import CohereEmbeddings
|
|
|
|
embeddings_model = CohereEmbeddings(cohere_api_key="...")
|
|
```
|
|
|
|
Otherwise you can initialize without any params:
|
|
```python
|
|
from langchain_cohere import CohereEmbeddings
|
|
|
|
embeddings_model = CohereEmbeddings()
|
|
```
|
|
|
|
</TabItem>
|
|
</Tabs>
|
|
|
|
### `embed_documents`
|
|
#### Embed list of texts
|
|
|
|
Use `.embed_documents` to embed a list of strings, recovering a list of embeddings:
|
|
|
|
```python
|
|
embeddings = embeddings_model.embed_documents(
|
|
[
|
|
"Hi there!",
|
|
"Oh, hello!",
|
|
"What's your name?",
|
|
"My friends call me World",
|
|
"Hello World!"
|
|
]
|
|
)
|
|
len(embeddings), len(embeddings[0])
|
|
```
|
|
|
|
<CodeOutputBlock language="python">
|
|
|
|
```
|
|
(5, 1536)
|
|
```
|
|
|
|
</CodeOutputBlock>
|
|
|
|
### `embed_query`
|
|
#### Embed single query
|
|
Use `.embed_query` to embed a single piece of text (e.g., for the purpose of comparing to other embedded pieces of texts).
|
|
|
|
```python
|
|
embedded_query = embeddings_model.embed_query("What was the name mentioned in the conversation?")
|
|
embedded_query[:5]
|
|
```
|
|
|
|
<CodeOutputBlock language="python">
|
|
|
|
```
|
|
[0.0053587136790156364,
|
|
-0.0004999046213924885,
|
|
0.038883671164512634,
|
|
-0.003001077566295862,
|
|
-0.00900818221271038]
|
|
```
|
|
|
|
</CodeOutputBlock>
|