Files
langchain/docs/extras/integrations/document_loaders/mongodb.ipynb
T
715ffda28b mongodb doc loader init (#10645)
- **Description:** A Document Loader for MongoDB
  - **Issue:** n/a
  - **Dependencies:** Motor, the async driver for MongoDB
  - **Tag maintainer:** n/a
  - **Twitter handle:** pigpenblue

Note that an initial mongodb document loader was created 4 months ago,
but the [PR ](https://github.com/langchain-ai/langchain/pull/4285)was
never pulled in. @leo-gan had commented on that PR, but given it is
extremely far behind the master branch and a ton has changed in
Langchain since then (including repo name and structure), I rewrote the
branch and issued a new PR with the expectation that the old one can be
closed.

Please reference that old PR for comments/context, but it can be closed
in favor of this one. Thanks!

---------

Co-authored-by: Bagatur <baskaryan@gmail.com>
Co-authored-by: Eugene Yurtsev <eyurtsev@gmail.com>
2023-09-29 11:44:07 -04:00

3.9 KiB

MongoDB

MongoDB is a NoSQL , document-oriented database that supports JSON-like documents with a dynamic schema.

Overview

The MongoDB Document Loader returns a list of Langchain Documents from a MongoDB database.

The Loader requires the following parameters:

  • MongoDB connection string
  • MongoDB database name
  • MongoDB collection name
  • (Optional) Content Filter dictionary

The output takes the following format:

  • pageContent= Mongo Document
  • metadata={'database': '[database_name]', 'collection': '[collection_name]'}

Load the Document Loader

In [24]:
# add this import for running in jupyter notebook
import nest_asyncio
nest_asyncio.apply()
In [ ]:
from langchain.document_loaders.mongodb import MongodbLoader
In [25]:
loader = MongodbLoader(connection_string="mongodb://localhost:27017/",
                       db_name="sample_restaurants", 
                       collection_name="restaurants",
                       filter_criteria={"borough": "Bronx", "cuisine": "Bakery" },
                       ) 
In [26]:
docs = loader.load()

len(docs)
Out [26]:
25359
In [27]:
docs[0]
Out [27]:
Document(page_content="{'_id': ObjectId('5eb3d668b31de5d588f4292a'), 'address': {'building': '2780', 'coord': [-73.98241999999999, 40.579505], 'street': 'Stillwell Avenue', 'zipcode': '11224'}, 'borough': 'Brooklyn', 'cuisine': 'American', 'grades': [{'date': datetime.datetime(2014, 6, 10, 0, 0), 'grade': 'A', 'score': 5}, {'date': datetime.datetime(2013, 6, 5, 0, 0), 'grade': 'A', 'score': 7}, {'date': datetime.datetime(2012, 4, 13, 0, 0), 'grade': 'A', 'score': 12}, {'date': datetime.datetime(2011, 10, 12, 0, 0), 'grade': 'A', 'score': 12}], 'name': 'Riviera Caterer', 'restaurant_id': '40356018'}", metadata={'database': 'sample_restaurants', 'collection': 'restaurants'})