mirror of
https://github.com/langchain-ai/langchain.git
synced 2026-10-08 02:45:22 +03:00
#docs: text splitters improvements
Changes are only in the Jupyter notebooks.
- added links to the source packages and a short description of these
packages
- removed " Text Splitters" suffixes from the TOC elements (they made
the list of the text splitters messy)
- moved text splitters, based on the length function into a separate
list. They can be mixed with any classes from the "Text Splitters", so
it is a different classification.
## Who can review?
@hwchase17 - project lead
@eyurtsev
@vowelparrot
NOTE: please, check out the results of the `Python code` text splitter
example (text_splitters/examples/python.ipynb). It looks suboptimal.
3.1 KiB
3.1 KiB
In [1]:
from langchain.text_splitter import PythonCodeTextSplitterIn [2]:
python_text = """
class Foo:
def bar():
def foo():
def testing_func():
def bar():
"""
python_splitter = PythonCodeTextSplitter(chunk_size=30, chunk_overlap=0)In [3]:
docs = python_splitter.create_documents([python_text])In [4]:
docsOut [4]:
[Document(page_content='Foo:\n\n def bar():', lookup_str='', metadata={}, lookup_index=0),
Document(page_content='foo():\n\ndef testing_func():', lookup_str='', metadata={}, lookup_index=0),
Document(page_content='bar():', lookup_str='', metadata={}, lookup_index=0)]In [3]:
python_splitter.split_text(python_text)Out [3]:
['Foo:\n\n def bar():', 'foo():\n\ndef testing_func():', 'bar():']
In [ ]: