Files
langchain/docs/modules/indexes/text_splitters/examples/python.ipynb
T
Leonid Ganeline c998569c8f docs: text splitters improvements (#4490)
#docs: text splitters improvements

Changes are only in the Jupyter notebooks.
- added links to the source packages and a short description of these
packages
- removed " Text Splitters" suffixes from the TOC elements (they made
the list of the text splitters messy)
- moved text splitters, based on the length function into a separate
list. They can be mixed with any classes from the "Text Splitters", so
it is a different classification.

## Who can review?
        @hwchase17 - project lead
        @eyurtsev
        @vowelparrot

NOTE: please, check out the results of the `Python code` text splitter
example (text_splitters/examples/python.ipynb). It looks suboptimal.
2023-05-17 21:33:34 -07:00

3.1 KiB

Python Code

PythonCodeTextSplitter splits text along python class and method definitions. It's implemented as a simple subclass of RecursiveCharacterSplitter with Python-specific separators. See the source code to see the Python syntax expected by default.

  1. How the text is split: by list of python specific separators
  2. How the chunk size is measured: by number of characters
In [1]:
from langchain.text_splitter import PythonCodeTextSplitter
In [2]:
python_text = """
class Foo:

    def bar():
    
    
def foo():

def testing_func():

def bar():
"""
python_splitter = PythonCodeTextSplitter(chunk_size=30, chunk_overlap=0)
In [3]:
docs = python_splitter.create_documents([python_text])
In [4]:
docs
Out [4]:
[Document(page_content='Foo:\n\n    def bar():', lookup_str='', metadata={}, lookup_index=0),
 Document(page_content='foo():\n\ndef testing_func():', lookup_str='', metadata={}, lookup_index=0),
 Document(page_content='bar():', lookup_str='', metadata={}, lookup_index=0)]
In [3]:
python_splitter.split_text(python_text)
Out [3]:
['Foo:\n\n    def bar():', 'foo():\n\ndef testing_func():', 'bar():']
In [ ]: