mirror of
https://github.com/langchain-ai/langchain.git
synced 2026-10-08 10:55:18 +03:00
#docs: text splitters improvements
Changes are only in the Jupyter notebooks.
- added links to the source packages and a short description of these
packages
- removed " Text Splitters" suffixes from the TOC elements (they made
the list of the text splitters messy)
- moved text splitters, based on the length function into a separate
list. They can be mixed with any classes from the "Text Splitters", so
it is a different classification.
## Who can review?
@hwchase17 - project lead
@eyurtsev
@vowelparrot
NOTE: please, check out the results of the `Python code` text splitter
example (text_splitters/examples/python.ipynb). It looks suboptimal.
3.9 KiB
3.9 KiB
In [1]:
from langchain.text_splitter import MarkdownTextSplitterIn [2]:
markdown_text = """
# 🦜️🔗 LangChain
⚡ Building applications with LLMs through composability ⚡
## Quick Install
```bash
# Hopefully this code block isn't split
pip install langchain
```
As an open source project in a rapidly developing field, we are extremely open to contributions.
"""
markdown_splitter = MarkdownTextSplitter(chunk_size=100, chunk_overlap=0)In [3]:
docs = markdown_splitter.create_documents([markdown_text])In [4]:
docsOut [4]:
[Document(page_content='# 🦜️🔗 LangChain\n\n⚡ Building applications with LLMs through composability ⚡', metadata={}),
Document(page_content="Quick Install\n\n```bash\n# Hopefully this code block isn't split\npip install langchain", metadata={}),
Document(page_content='As an open source project in a rapidly developing field, we are extremely open to contributions.', metadata={})]In [5]:
markdown_splitter.split_text(markdown_text)Out [5]:
['# 🦜️🔗 LangChain\n\n⚡ Building applications with LLMs through composability ⚡', "Quick Install\n\n```bash\n# Hopefully this code block isn't split\npip install langchain", 'As an open source project in a rapidly developing field, we are extremely open to contributions.']
In [ ]: