> ## Documentation Index
> Fetch the complete documentation index at: https://langchain-5e9cc07a-preview-change-1791323909-75c753a.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# Dell powerscale integration

> Integrate with the Dell powerscale document loader using LangChain Python.

[Dell PowerScale](https://www.dell.com/en-us/shop/powerscale-family/sf/powerscale) is an enterprise scale-out storage system that hosts the industry-leading OneFS filesystem, which can be hosted on-premises or deployed in the cloud.

This document loader utilizes unique capabilities from PowerScale that can determine which files have been modified since an application's last run and only returns modified files for processing. This will eliminate the need to re-process (chunk and embed) files that have not been changed, improving the overall data ingestion workflow.

This loader requires PowerScale's MetadataIQ feature enabled. Additional information can be found on our GitHub Repo: [https://github.com/dell/powerscale-rag-connector](https://github.com/dell/powerscale-rag-connector)

## Overview

### Integration details

| Class | Package | Local | Serializable | [JS support](https://js.langchain.com/docs/integrations/document_loaders/web_loaders/__module_name___loader) |
| :- | :- | :-: | :-: | :-: |
| [`PowerScaleDocumentLoader`](https://github.com/dell/powerscale-rag-connector/blob/main/src/powerscale_rag_connector/PowerScaleDocumentLoader.py) | [`powerscale-rag-connector`](https://github.com/dell/powerscale-rag-connector) | ✅ | ❌ | ❌ |
| [`PowerScaleUnstructuredLoader`](https://github.com/dell/powerscale-rag-connector/blob/main/src/powerscale_rag_connector/PowerScaleUnstructuredLoader.py) | [`powerscale-rag-connector`](https://github.com/dell/powerscale-rag-connector) | ✅ | ❌ | ❌ |

### Loader features

| Source | Document Lazy Loading | Native Async Support |
| :-: | :-: | :-: |
| `PowerScaleDocumentLoader` | ✅ | ❌ |
| `PowerScaleUnstructuredLoader` | ✅ | ❌ |

## Setup

This document loader requires the use of a Dell PowerScale system with MetadataIQ enabled. Additional information can be found on our GitHub page: [https://github.com/dell/powerscale-rag-connector](https://github.com/dell/powerscale-rag-connector).

### Installation

The document loader lives in an external pip package and can be installed using standard tooling.

```bash theme={null}
pip install -qU powerscale-rag-connector[langchain]
```

## Initialization

Now we can instantiate the document loader:

### Generic document loader

Our generic document loader can be used to incrementally load all files from PowerScale in the following manner:

```python theme={null}
from powerscale_rag_connector import PowerScaleDocumentLoader

loader = PowerScaleDocumentLoader(
    es_host_url="http://elasticsearch:9200",
    es_index_name="metadataiq",
    es_api_key="your-api-key",
    folder_path="/ifs/data",
)
```

### UnstructuredLoader

Optionally, the `PowerScaleUnstructuredLoader` can be used to locate the changed files *and* automatically process the files producing elements of the source file. This is done using LangChain's `UnstructuredLoader` class.

```python theme={null}
from powerscale_rag_connector import PowerScaleUnstructuredLoader

# Or load files with the Unstructured Loader
loader = PowerScaleUnstructuredLoader(
    es_host_url="http://elasticsearch:9200",
    es_index_name="metadataiq",
    es_api_key="your-api-key",
    folder_path="/ifs/data",
    # 'basic' merges elements into larger contiguous chunks
    # Use 'by_title' for section-title boundaries,
    # or omit chunking_strategy to get one Document per element
    chunking_strategy="basic",
)
```

The fields:

* `es_host_url` is the endpoint to MetadataIQ Elasticsearch database
* `es_index_name` is the name of the index where PowerScale writes its file system metadata
* `es_api_key` is the **encoded** version of your Elasticsearch API key
* `folder_path` is the path on PowerScale to be queried for changes

## Load

Internally, all code is asynchronous with PowerScale and MetadataIQ and the load and lazy load methods will return a Python generator. We recommend using the lazy load function.

```python theme={null}
for doc in loader.load():
    print(doc)
```

```python theme={null}
[Document(page_content='' metadata={'source': '/ifs/pdfs/1994-Graph.Theoretic.Obstacles.to.Perfect.Hashing.TR0257.pdf', 'snapshot': 20834, 'lin': 1099511628345, 'change_types': ['ENTRY_ADDED']}),
Document(page_content='' metadata={'source': '/ifs/pdfs/New.sendfile-FreeBSD.20.Feb.2015.pdf', 'snapshot': 20920, 'lin': 1099511628346, 'change_types': ['ENTRY_MODIFIED']}),
Document(page_content='' metadata={'source': '/ifs/pdfs/FAST-Fast.Architecture.Sensitive.Tree.Search.on.Modern.CPUs.and.GPUs-Slides.pdf', 'snapshot': 20924, 'lin': 1099511628347, 'change_types': ['ENTRY_ADDED']})]
```

### Returned object

Both document loaders will keep track of what files were previously returned to your application. When called again, the document loader will only return new or modified files since your previous run. The connector uses the file's birth time (`btime`) to distinguish between `ENTRY_ADDED` and `ENTRY_MODIFIED`:

* `ENTRY_ADDED` — the file was created after the last checkpoint

* `ENTRY_MODIFIED` — the file existed before the last checkpoint and has been modified

* The `metadata` fields in the returned [`Document`](https://reference.langchain.com/python/langchain-core/documents/base/Document) will return the path on PowerScale that contains the modified file. You will use this path to read the data via NFS (or S3) and process the data in your application (e.g.: create chunks and embedding).

* The `source` field is the path on PowerScale and not necessarily on your local system (depending on your mount strategy); OneFS expresses the entire storage system as a single tree rooted at `/ifs`.

* The `snapshot` field is the PowerScale snapshot ID associated with the file.

* The `lin` field is the PowerScale logical inode number — a unique, stable identifier for each file on the filesystem.

* The `change_types` property will inform you on what change occurred since the last one.

Your RAG application can use the information from `change_types` to add or update entries in your chunk and vector store.

When using `PowerScaleUnstructuredLoader`, the `page_content` field will be filled with data from the Unstructured Loader.

## Lazy load

Internally, all code is asynchronous with PowerScale and MetadataIQ and the load and lazy load methods will return a Python generator. We recommend using the lazy load function.

```python theme={null}
for doc in loader.lazy_load():
    print(doc)  # do something specific with the document
```

The same [`Document`](https://reference.langchain.com/python/langchain-core/documents/base/Document) is returned as the load function with all the same properties mentioned above.

## Additional examples

Additional examples and code can be found on our public GitHub webpage: [https://github.com/dell/powerscale-rag-connector/tree/main/examples](https://github.com/dell/powerscale-rag-connector/tree/main/examples) that provide full working examples.

* [PowerScale LangChain Document Loader](https://github.com/dell/powerscale-rag-connector/blob/main/examples/powerscale_langchain_doc_loader.py) - Working example of our standard document loader
* [PowerScale LangChain Unstructured Loader](https://github.com/dell/powerscale-rag-connector/blob/main/examples/powerscale_langchain_unstructured_loader.py) - Working example of our standard document loader using unstructured loader for chunking and embedding
* [PowerScale NVIDIA Ingest LangChain Document Loader](https://github.com/dell/powerscale-rag-connector/blob/main/examples/powerscale_nvingest_langchain_doc_loader.py) - Working example of our document loader with NVIDIA NvIngest for chunking and embedding

***

## API reference

For detailed documentation of all PowerScale Document Loader features and configurations, head to the GitHub page: [https://github.com/dell/powerscale-rag-connector/](https://github.com/dell/powerscale-rag-connector/).

***

<div className="source-links">
  <Callout icon="terminal-2">
    [Connect these docs](/use-these-docs) to your agent of choice via MCP for real-time answers.
  </Callout>

  <Callout icon="edit">
    [Edit this page on GitHub](https://github.com/langchain-ai/docs/edit/main/src/oss/python/integrations/document_loaders/powerscale.mdx) or [file an issue](https://github.com/langchain-ai/docs/issues/new/choose).
  </Callout>
</div>
