Files
langchain/docs/extras/integrations/document_loaders/azure_document_intelligence.ipynb
T
Lars von Wedel 6d82503eb1 Add parser and loader for Azure document intelligence service. (#10136)
Hi,

this PR contains loader / parser for Azure Document intelligence which
is a ML-based service to ingest arbitrary PDFs / images, even if
scanned. The loader generates Documents by pages of the original
document. This is my first contribution to LangChain.

Unfortunately I could not find the correct place for test cases. Happy
to add one if you can point me to the location, but as this is a
cloud-based service, a test would require network access and credentials
- so might be of limited help.

Dependencies: The needed dependency was already part of pyproject.toml,
no change.
Twitter: feel free to mention @LarsAC on the announcement
2023-09-03 14:25:39 -07:00

3.4 KiB

Azure Document Intelligence

Azure Document Intelligence (formerly known as Azure Forms Recognizer) is machine-learning based service that extracts text (including handwriting), tables or key-value-pairs from scanned documents or images.

This current implementation of a loader using Document Intelligence is able to incorporate content page-wise and turn it into LangChain documents.

Document Intelligence supports PDF, JPEG, PNG, BMP, or TIFF.

Further documentation is available at https://learn.microsoft.com/en-us/azure/ai-services/document-intelligence/?view=doc-intel-3.1.0.

In [ ]:
%pip install langchain azure-ai-formrecognizer -q

Example 1

The first example uses a local file which will be sent to Azure Document Intelligence.

First, an instance of a DocumentAnalysisClient is created with endpoint and key for the Azure service.

In [ ]:
from azure.ai.formrecognizer import DocumentAnalysisClient
from azure.core.credentials import AzureKeyCredential

document_analysis_client = DocumentAnalysisClient(
                endpoint="<service_endpoint>", credential=AzureKeyCredential("<service_key>")
            )

With the initialized document analysis client, we can proceed to create an instance of the DocumentIntelligenceLoader:

In [9]:
from langchain.document_loaders.pdf import DocumentIntelligenceLoader
loader = DocumentIntelligenceLoader(
    "<Local_filename>",
    client=document_analysis_client,
    model="<model_name>") # e.g. prebuilt-document

documents = loader.load()

The output contains each page of the source document as a LangChain document:

In [18]:
documents
Out [18]:
[Document(page_content='...', metadata={'source': '...', 'page': 1})]