Files
langchain/docs/modules/document_loaders/examples/microsoft_word.ipynb
T

2.9 KiB

Microsoft Word

This notebook shows how to load text from Microsoft word documents.

In [1]:
from langchain.document_loaders import UnstructuredDocxLoader
In [2]:
loader = UnstructuredDocxLoader('example_data/fake.docx')
In [3]:
data = loader.load()
In [4]:
data
Out [4]:
[Document(page_content='Lorem ipsum dolor sit amet.', lookup_str='', metadata={'source': 'example_data/fake.docx'}, lookup_index=0)]

Retain Elements

Under the hood, Unstructured creates different "elements" for different chunks of text. By default we combine those together, but you can easily keep that separation by specifying mode="elements".

In [2]:
loader = UnstructuredDocxLoader('example_data/fake.docx', mode="elements")
In [3]:
data = loader.load()
In [4]:
data
Out [4]:
[Document(page_content='Lorem ipsum dolor sit amet.', lookup_str='', metadata={'source': 'example_data/fake.docx'}, lookup_index=0)]
In [ ]: