Files
langchain/docs/modules/indexes/text_splitters/examples/markdown.ipynb
T
Leonid Ganeline c998569c8f docs: text splitters improvements (#4490)
#docs: text splitters improvements

Changes are only in the Jupyter notebooks.
- added links to the source packages and a short description of these
packages
- removed " Text Splitters" suffixes from the TOC elements (they made
the list of the text splitters messy)
- moved text splitters, based on the length function into a separate
list. They can be mixed with any classes from the "Text Splitters", so
it is a different classification.

## Who can review?
        @hwchase17 - project lead
        @eyurtsev
        @vowelparrot

NOTE: please, check out the results of the `Python code` text splitter
example (text_splitters/examples/python.ipynb). It looks suboptimal.
2023-05-17 21:33:34 -07:00

3.9 KiB

Markdown

Markdown is a lightweight markup language for creating formatted text using a plain-text editor.

MarkdownTextSplitter splits text along Markdown headings, code blocks, or horizontal rules. It's implemented as a simple subclass of RecursiveCharacterSplitter with Markdown-specific separators. See the source code to see the Markdown syntax expected by default.

  1. How the text is split: by list of markdown specific separators
  2. How the chunk size is measured: by number of characters
In [1]:
from langchain.text_splitter import MarkdownTextSplitter
In [2]:
markdown_text = """
# 🦜️🔗 LangChain

⚡ Building applications with LLMs through composability ⚡

## Quick Install

```bash
# Hopefully this code block isn't split
pip install langchain
```

As an open source project in a rapidly developing field, we are extremely open to contributions.
"""
markdown_splitter = MarkdownTextSplitter(chunk_size=100, chunk_overlap=0)
In [3]:
docs = markdown_splitter.create_documents([markdown_text])
In [4]:
docs
Out [4]:
[Document(page_content='# 🦜️🔗 LangChain\n\n⚡ Building applications with LLMs through composability ⚡', metadata={}),
 Document(page_content="Quick Install\n\n```bash\n# Hopefully this code block isn't split\npip install langchain", metadata={}),
 Document(page_content='As an open source project in a rapidly developing field, we are extremely open to contributions.', metadata={})]
In [5]:
markdown_splitter.split_text(markdown_text)
Out [5]:
['# 🦜️🔗 LangChain\n\n⚡ Building applications with LLMs through composability ⚡',
 "Quick Install\n\n```bash\n# Hopefully this code block isn't split\npip install langchain",
 'As an open source project in a rapidly developing field, we are extremely open to contributions.']
In [ ]: