Files
langchain/docs/modules/document_loaders/examples/word_document.ipynb
T
Klein Tahiraj d3d4503ce2 Remove redundant .docx loader (closes #1716) + update how_to_guides.rst (#1891)
In https://github.com/hwchase17/langchain/issues/1716 , it was
identified that there were two .py files performing similar tasks. As a
resolution, one of the files has been removed, as its purpose had
already been fulfilled by the other file. Additionally, the init has
been updated accordingly.

Furthermore, the how_to_guides.rst file has been updated to include
links to documentation that was previously missing. This was deemed
necessary as the existing list on
https://langchain.readthedocs.io/en/latest/modules/document_loaders/how_to_guides.html
was incomplete, causing confusion for users who rely on the full list of
documentation on the left sidebar of the website.
2023-03-22 15:19:42 -07:00

2.8 KiB

Word Documents

This covers how to load Word documents into a document format that we can use downstream.

In [1]:
from langchain.document_loaders import UnstructuredWordDocumentLoader
In [2]:
loader = UnstructuredWordDocumentLoader("example_data/fake.docx")
In [3]:
data = loader.load()
In [4]:
data
Out [4]:
[Document(page_content='Lorem ipsum dolor sit amet.', lookup_str='', metadata={'source': 'fake.docx'}, lookup_index=0)]

Retain Elements

Under the hood, Unstructured creates different "elements" for different chunks of text. By default we combine those together, but you can easily keep that separation by specifying mode="elements".

In [5]:
loader = UnstructuredWordDocumentLoader("example_data/fake.docx", mode="elements")
In [6]:
data = loader.load()
In [7]:
data[0]
Out [7]:
Document(page_content='Lorem ipsum dolor sit amet.', lookup_str='', metadata={'source': 'fake.docx', 'filename': 'fake.docx', 'category': 'Title'}, lookup_index=0)