Files
langchain/docs/modules/indexes/document_loaders/examples/html.ipynb
T
Zander ChaseandLeonid Ganeline aa38355999 Vwp/docs improved document loaders (#4006)
Huge thanks to @leo-gan for improving the document loaders notebooks

---------

Co-authored-by: Leonid Ganeline <leo.gan.57@gmail.com>
2023-05-02 15:24:53 -07:00

2.9 KiB

HTML

This covers how to load HTML documents into a document format that we can use downstream.

In [1]:
from langchain.document_loaders import UnstructuredHTMLLoader
In [2]:
loader = UnstructuredHTMLLoader("example_data/fake-content.html")
In [3]:
data = loader.load()
In [4]:
data
Out [4]:
[Document(page_content='My First Heading\n\nMy first paragraph.', lookup_str='', metadata={'source': 'example_data/fake-content.html'}, lookup_index=0)]

Loading HTML with BeautifulSoup4

We can also use BeautifulSoup4 to load HTML documents using the BSHTMLLoader. This will extract the text from the HTML into page_content, and the page title as title into metadata.

In [1]:
from langchain.document_loaders import BSHTMLLoader
In [2]:
loader = BSHTMLLoader("example_data/fake-content.html")
data = loader.load()
data
Out [2]:
[Document(page_content='\n\nTest Title\n\n\nMy First Heading\nMy first paragraph.\n\n\n', metadata={'source': 'example_data/fake-content.html', 'title': 'Test Title'})]