Files
langchain/docs/modules/indexes/document_loaders/examples/url.ipynb
T
2023-04-13 22:15:03 -07:00

4.6 KiB

URL

This covers how to load HTML documents from a list of URLs into a document format that we can use downstream.

In [1]:
 from langchain.document_loaders import UnstructuredURLLoader
In [2]:
urls = [
    "https://www.understandingwar.org/backgrounder/russian-offensive-campaign-assessment-february-8-2023",
    "https://www.understandingwar.org/backgrounder/russian-offensive-campaign-assessment-february-9-2023"
]
In [3]:
loader = UnstructuredURLLoader(urls=urls)
In [4]:
data = loader.load()

Selenium URL Loader

This covers how to load HTML documents from a list of URLs using the SeleniumURLLoader.

Using selenium allows us to load pages that require JavaScript to render.

Setup

To use the SeleniumURLLoader, you will need to install selenium and unstructured.

In [ ]:
from langchain.document_loaders import SeleniumURLLoader
In [ ]:
urls = [
    "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
    "https://goo.gl/maps/NDSHwePEyaHMFGwh8"
]
In [ ]:
loader = SeleniumURLLoader(urls=urls)
In [ ]:
data = loader.load()

Playwright URL Loader

This covers how to load HTML documents from a list of URLs using the PlaywrightURLLoader.

As in the Selenium case, Playwright allows us to load pages that need JavaScript to render.

Setup

To use the PlaywrightURLLoader, you will need to install playwright and unstructured. Additionally, you will need to install the Playwright Chromium browser:

In [ ]:
# Install playwright
!pip install "playwright"
!pip install "unstructured"
!playwright install
In [ ]:
from langchain.document_loaders import PlaywrightURLLoader
In [ ]:
urls = [
    "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
    "https://goo.gl/maps/NDSHwePEyaHMFGwh8"
]
In [ ]:
loader = PlaywrightURLLoader(urls=urls, remove_selectors=["header", "footer"])
In [ ]:
data = loader.load()