Files
2023-03-26 19:49:46 -07:00

2.9 KiB

HTML

This covers how to load HTML documents into a document format that we can use downstream.

In [1]:
from langchain.document_loaders import UnstructuredHTMLLoader
In [2]:
loader = UnstructuredHTMLLoader("example_data/fake-content.html")
In [3]:
data = loader.load()
In [4]:
data
Out [4]:
[Document(page_content='My First Heading\n\nMy first paragraph.', lookup_str='', metadata={'source': 'example_data/fake-content.html'}, lookup_index=0)]

Loading HTML with BeautifulSoup4

We can also use BeautifulSoup4 to load HTML documents using the BSHTMLLoader. This will extract the text from the html into page_content, and the page title as title into metadata.

In [16]:
from langchain.document_loaders import BSHTMLLoader
In [17]:
loader = BSHTMLLoader("example_data/fake-content.html")
data = loader.load()
data
Out [17]:
[Document(page_content='\n\nTest Title\n\n\nMy First Heading\nMy first paragraph.\n\n\n', lookup_str='', metadata={'source': 'example_data/fake-content.html', 'title': 'Test Title'}, lookup_index=0)]
In [ ]: