Files
langchain/docs/modules/indexes/text_splitters/examples/tiktoken_splitter.ipynb
T
2023-03-26 19:49:46 -07:00

2.0 KiB

TiktokenText Splitter

  1. How the text is split: by tiktoken tokens
  2. How the chunk size is measured: by tiktoken tokens
In [3]:
# This is a long document we can split up.
with open('../../../state_of_the_union.txt') as f:
    state_of_the_union = f.read()
In [4]:
from langchain.text_splitter import TokenTextSplitter
In [5]:
text_splitter = TokenTextSplitter(chunk_size=10, chunk_overlap=0)
In [6]:
texts = text_splitter.split_text(state_of_the_union)
print(texts[0])
Madam Speaker, Madam Vice President, our
In [ ]: