mirror of
https://github.com/langchain-ai/langchain.git
synced 2026-10-09 11:25:21 +03:00
Hi, this PR contains loader / parser for Azure Document intelligence which is a ML-based service to ingest arbitrary PDFs / images, even if scanned. The loader generates Documents by pages of the original document. This is my first contribution to LangChain. Unfortunately I could not find the correct place for test cases. Happy to add one if you can point me to the location, but as this is a cloud-based service, a test would require network access and credentials - so might be of limited help. Dependencies: The needed dependency was already part of pyproject.toml, no change. Twitter: feel free to mention @LarsAC on the announcement
3.4 KiB
3.4 KiB
In [ ]:
%pip install langchain azure-ai-formrecognizer -qIn [ ]:
from azure.ai.formrecognizer import DocumentAnalysisClient
from azure.core.credentials import AzureKeyCredential
document_analysis_client = DocumentAnalysisClient(
endpoint="<service_endpoint>", credential=AzureKeyCredential("<service_key>")
)In [9]:
from langchain.document_loaders.pdf import DocumentIntelligenceLoader
loader = DocumentIntelligenceLoader(
"<Local_filename>",
client=document_analysis_client,
model="<model_name>") # e.g. prebuilt-document
documents = loader.load()In [18]:
documentsOut [18]:
[Document(page_content='...', metadata={'source': '...', 'page': 1})]