FAQs
Answers to questions about general features, troubleshooting, usage, and more.
- General features
- What sets RAGFlow apart from other RAG products?
- Where can I find the RAGFlow version, and how do I interpret it?
- Why does RAGFlow use Elasticsearch as its default document engine?
- What are the differences between cloud.ragflow.io and a locally deployed open-source RAGFlow service?
- Why does RAGFlow require substantial system resources?
- Which architectures or devices does RAGFlow support?
- Do you offer an API for integration with third-party applications?
- Do you support stream output?
- What are the key differences between Search and Chat?
- Troubleshooting
- How do I build a RAGFlow image from scratch?
- Why does PDF parsing fail when DeepDoc model files are missing?
- Does the open-source 1.0 Go DeepDoc backend support GPU inference?
- Why can't RAGFlow access my Ollama model?
- Why can't Nginx reach the Go API?
- What does
network anomaly: There is an abnormality in your network and you cannot connect to the servermean? - Why are document tasks waiting in the queue?
- Why does document parsing stall below 1%?
- Why does PDF parsing stall near completion without any errors in the logs?
- What does
Index failuremean? - How do I check RAGFlow logs?
- How do I check the status of each RAGFlow component?
- Why can't RAGFlow connect to Elasticsearch?
- Why does the Elasticsearch container fail to start with
Elasticsearch did not exit normally? - How do I configure
vm.max_map_counton macOS? - Why does the Go API return a 404 response?
- Do you provide examples of using DeepDoc to parse PDFs or other files?
- Why can't RAGFlow find a required file?
- Usage
- How do I run RAGFlow with a locally deployed LLM?
- How do I add an LLM that is not directly supported?
- How do I change the file size limit?
- How do I get an API key for integration with third-party applications?
- How do I upgrade RAGFlow?
- How do I switch the document engine to Infinity?
- Where are uploaded files stored in RAGFlow?
- How do I tune document parsing and embedding throughput?
- Why can't I retrieve relevant content in Search or Chat even though the document was parsed successfully?
- Why does the same query produce different results in Retrieval testing, Search, and Chat?
- How can I tell whether a problem comes from document parsing, chunking, retrieval, or answer generation?
- Can the same dataset be used by Search, Chat, and Agent?
- Should I adjust chunking, retrieval settings, or models first?
- What needs to be reprocessed after changing the embedding model, reranker, or LLM?
- Why can the model still hallucinate when the answer exists in the dataset?
- How do I reduce my chat assistant's response latency?
- How do I reduce my Agent's response latency?
- How do I use MinerU to parse PDF documents?
- How do I configure MinerU-specific settings?
- How do I use MinerU with a vLLM server for document parsing?
- How do I use an external Docling Serve server for document parsing?
- How do I use PaddleOCR for document parsing?
- How do I use Ollama with RAGFlow for local LLM inference?
General features
What sets RAGFlow apart from other RAG products?
RAGFlow provides an end-to-end RAG platform that goes beyond basic document chunking and retrieval. Its key strengths include:
- Deep document understanding for complex content such as text, tables, and images.
- Flexible retrieval and traceable answers with multiple retrieval strategies, reranking, and citations.
- Agentic capabilities for multi-step retrieval, reasoning, tool use, and workflows.
- Knowledge compilation for organizing documents into structured knowledge artifacts.
- Broad model and integration support for building different RAG applications.
Where can I find the RAGFlow version, and how do I interpret it?
You can find the RAGFlow version number on the System page of the UI:
If you build RAGFlow from source, the Go admin and ingestor log their versions at startup. The server also prints its version with bin/ragflow_server --api --version.
For example, v0.27.1-14-g6daae7f2 means 14 commits after the v0.27.1 tag, at a commit whose abbreviated ID is 6daae7f2. The actual value comes from the packaged VERSION file or, when that file is absent, from git describe; it may be a tag alone or unknown if neither source supplies a version.
Why does RAGFlow use Elasticsearch as its default document engine?
Elasticsearch meets RAGFlow's core hybrid search requirements, including full-text search, vector search, phrase search, and advanced ranking capabilities.
The Go backend also supports Infinity as a document engine. Infinity is an AI-native database developed by InfiniFlow and optimized for RAG workloads. Support levels and available features may vary between document engines.
What are the differences between cloud.ragflow.io and a locally deployed open-source RAGFlow service?
cloud.ragflow.io is the hosted RAGFlow service. It provides managed infrastructure and subscription-based limits and features, while a locally deployed open-source service gives you control over deployment, data, models, storage, and system resources.
REST API access on cloud.ragflow.io depends on the subscription plan. Check the current plan details on the RAGFlow website before relying on API access. A locally deployed service exposes the RAGFlow HTTP APIs directly and uses API keys created in its UI.
Why does RAGFlow require substantial system resources?
RAGFlow runs multiple components for document parsing, embedding, full-text and vector indexing, retrieval, task processing, metadata storage, caching, and object storage. Some document parsers also load or download machine-learning models and may require significant CPU and memory resources.
Actual resource usage depends on the selected document engine, parser, model provider, document volume, and workload. See the quickstart prerequisites for the recommended starting configuration.
Which architectures or devices does RAGFlow support?
The documented Go image build targets Linux AMD64. Building on another architecture requires a matching dependency image and native libraries, and the selected document engine must support that architecture. The commands in the quickstart guide target AMD64.
Do you offer an API for integration with third-party applications?
See the RAGFlow HTTP API Reference for integration endpoints.
Do you support stream output?
Yes. RAGFlow supports both streaming and non-streaming responses. The interactive Chat and Agent pages display responses as they are generated, while the embedded chat configuration provides an Enable streaming responses option.
For HTTP API calls, use the endpoint's stream parameter to select the response mode. Defaults can differ between endpoints, so check the corresponding API reference:
What are the key differences between Search and Chat?
- Search is designed for direct knowledge retrieval. It retrieves and ranks relevant content from one or more datasets and presents the results for users to review.
- Chat is designed for conversational question answering. It retrieves relevant knowledge and uses an LLM to generate answers, supporting multi-turn conversations and more advanced retrieval capabilities such as Agentic Retrieval.
Use Search when you want to find and inspect relevant knowledge directly, and Chat when you want generated answers based on that knowledge.
Troubleshooting
How do I build a RAGFlow image from scratch?
Build the Go image from the repository root with Dockerfile, then set RAGFLOW_IMAGE in docker/.env. See the quickstart guide for the current commands and prerequisites.
Why does PDF parsing fail when DeepDoc model files are missing?
The Go API and ingestor initialize the in-process DeepDoc backend. Check their logs for model initialization errors and verify that the configured DEEPDOC_MODEL_DIR contains the required model files, including det.ort, layout.ort, tsr.ort, rec.ort, and ocr.res. If you use the shared model asset directory, check MODEL_ASSETS_DIR instead. Restore the missing files from the deployment's model assets, then restart the affected service.
Does the open-source 1.0 Go DeepDoc backend support GPU inference?
No. The RAGFlow open-source 1.0 Go DeepDoc backend runs layout analysis, OCR, and table recognition on CPU.
Why can't RAGFlow access my Ollama model?
Check that Ollama is running, the model has been downloaded, the model name is correct, and the Ollama Base URL is reachable from the RAGFlow container.
If the request times out while loading the model, check the Ollama logs and available memory. Try a smaller model or allocate more memory if necessary.
See Deploy a local LLM for more information.
For setup instructions, see How do I use Ollama with RAGFlow for local LLM inference?
Why can't Nginx reach the Go API?
Check that the Go API and Admin services started, then review the RAGFlow container and Nginx logs. In the target Go configuration, Nginx forwards API requests to port 9380 and Admin requests to port 9381. Check the active ragflow.conf against docker/nginx/ragflow.conf.golang and verify that those services are listening before changing the proxy configuration.
What does network anomaly: There is an abnormality in your network and you cannot connect to the server mean?
You cannot log in until the services are initialized. Review the RAGFlow container logs (docker compose --env-file docker/.env -f docker/docker-compose.yml logs -f ragflow-cpu from the repository root). The Go services log these startup messages:
Starting api server: ...
Server starting on port: 9380
Starting admin server: ...
Starting RAGFlow admin HTTP server on port: 9381
Starting ingestor server: ...
Ingestor ... initialized
RAGFlow ingestion service version: ...
These messages come from separate processes and may appear in a different order. Check the logs for startup errors, use the system health API to inspect dependencies, and verify that the ingestor has registered with admin before treating document parsing as ready.
Why are document tasks waiting in the queue?
The Go ingestor consumes document tasks from NATS JetStream. A growing queue can mean that no ingestor is consuming, workers are busy, or tasks are repeatedly failing.
- Check that the NATS and RAGFlow containers are running, and inspect their logs for connection or consumer errors.
- Check the ingestor logs for
Ingestor ... initialized,NATS stream RAGFLOW_TASKS ready, pull errors, and task failures. Confirm that the ingestor is registered with admin. - Check whether workers are processing tasks and whether the task status or document progress changes. Investigate the first task error before retrying it.
Do not clear the queue as a first troubleshooting step: doing so can discard pending work.
Why does document parsing stall below 1%?
Click the red cross beside the 'parsing status' bar, then restart the parsing process to see if the issue remains. If the issue persists and your RAGFlow is deployed locally, try the following:
-
Check the RAGFlow container logs for API, admin, and ingestor startup errors:
docker compose --env-file docker/.env -f docker/docker-compose.yml logs -f ragflow-cpu -
Confirm that the Go ingestor is running and sending heartbeats to admin. Check admin's service status for the ingestor; an API health response alone does not confirm that workers are consuming tasks.
-
Check the NATS connection and JetStream consumer in the logs. Look for pull errors, consumer initialization failures, and repeated task failures.
-
Check the parser and tokenizer assets separately. Missing DeepDoc model files prevent the API or ingestor from starting because the in-process DeepDoc backend is required. A missing
cl100k_base.tiktokentable does not normally prevent startup, but ingestion fails for models that declare that tokenizer until the table is restored. -
Check whether ingestor workers are processing tasks and whether task progress changes. If a worker stops making progress, inspect its task error and available memory before retrying.
Why does PDF parsing stall near completion without any errors in the logs?
Click the red cross beside the 'parsing status' bar, then restart the parsing process to see if the issue remains. If the issue persists, check the ingestor and container logs for an out-of-memory error. Make sure the Docker runtime or host has enough memory available for the RAGFlow service, and reduce concurrent parsing work if necessary.
The standard Go Compose file does not apply MEM_LIMIT to ragflow-cpu; that variable limits selected dependency services. Changing it does not increase the memory available to the Go ingestor. Adjust the Docker runtime or host allocation, or add an explicit service-level memory limit in a Compose override if your environment requires one.
What does Index failure mean?
An index failure means that RAGFlow could not write the processed document chunks to the configured document engine.
Check the task logs, embedding model, document engine connection, and system health. If you use Elasticsearch, Infinity, or another document engine supported by the current Go backend, verify that the corresponding service is running and accessible.
How do I check RAGFlow logs?
From the repository root, follow the Go service logs written to the mounted log directory:
tail -f docker/ragflow-logs/*.log
To view the container's combined output, use docker compose --env-file docker/.env -f docker/docker-compose.yml logs -f ragflow-cpu.
How do I check the status of each RAGFlow component?
-
From the repository root, check the Go deployment's containers:
docker compose --env-file docker/.env -f docker/docker-compose.yml psReview the RAGFlow service logs:
docker compose --env-file docker/.env -f docker/docker-compose.yml logs -f ragflow-cpuCheck the
natsand selected document engine services in the same Compose project when diagnosing ingestion or indexing. -
Use the system health API to check the database, Kvrocks cache, document engine, object storage, and NATS message queue.
A running container does not necessarily mean that the service inside it is healthy. Check the health API and logs for connection, port, DNS, and configuration errors.
Why can't RAGFlow connect to Elasticsearch?
-
Check the status of the Elasticsearch Docker container:
docker compose --env-file docker/.env -f docker/docker-compose.yml ps es01A healthy Elasticsearch container should report
healthy. The default Go Compose deployment exposes Elasticsearch on host port1200, while the container listens on port9200. -
Follow the system health API to check the health status of the Elasticsearch service.
IMPORTANTThe status of a Docker container does not necessarily reflect the status of the service inside it. A service can be unhealthy even when its container is up and running. Possible causes include network failures, incorrect ports, and DNS or configuration errors.
-
If your container keeps restarting, ensure
vm.max_map_count>= 262144. On Linux, update /etc/sysctl.conf to keep the change permanent. For macOS, see the following FAQ.
Why does the Elasticsearch container fail to start with Elasticsearch did not exit normally?
On Linux, an insufficient vm.max_map_count value is a common cause. Check the Elasticsearch logs first. If they report that vm.max_map_count is too low, set it to at least 262144 and add the setting to /etc/sysctl.conf so it persists after a reboot. If the logs report a different error, check memory, disk space, permissions, and the Elasticsearch data directory instead.
How do I configure vm.max_map_count on macOS?
vm.max_map_count is a Linux kernel parameter required only by Elasticsearch. It does not affect RAGFlow deployments that use Infinity as the document engine.
On Docker Desktop, set the value in its Linux virtual machine:
docker run --rm --privileged alpine sysctl -w vm.max_map_count=262144
This setting is temporary and is reset when Docker Desktop restarts.
On Colima, check the current value inside its virtual machine:
colima ssh -- sysctl vm.max_map_count
To set the value temporarily:
colima ssh -- sudo sysctl -w vm.max_map_count=262144
For a persistent setting, run colima start --edit and add a provision script to colima.yaml:
provision:
- mode: system
script: |
#!/bin/bash
sysctl -w vm.max_map_count=262144
Then restart Colima:
colima stop
colima start
Contributed by @helloxjade.
Why does the Go API return a 404 response?
The Go API returns HTTP 404 with Not Found: <path> when a request does not match a route. Your URL may point to the wrong service.
For normal UI and API access, use http://<IP_OF_YOUR_MACHINE> when SVR_WEB_HTTP_PORT is 80, or include the configured web port. Nginx forwards API traffic to the Go API on internal port 9380 and Admin traffic to internal port 9381.
The Go Compose deployment uses SVR_HTTP_PORT and ADMIN_SVR_HTTP_PORT for the directly published API and Admin host ports. For normal browser and API access, prefer the Nginx web port.
Do you provide examples of using DeepDoc to parse PDFs or other files?
Yes. The Go parser implementations are in internal/parser/parser, and the Go DeepDoc integration is in internal/deepdoc. The following tests provide concrete usage examples:
internal/deepdoc/parser/pdf/parser_pipeline_integration_test.godemonstrates the PDF parsing and post-processing pipeline.internal/deepdoc/parser/docx/parser_integration_test.godemonstrates DOCX parsing.internal/parser/parser/pdf_parser_pages_e2e_test.godemonstrates configuring and calling the higher-level PDF parser adapter.
RAGFlow's ingestor uses these components when it processes documents. These files are integration or end-to-end tests, so review their prerequisites and build constraints before running them directly.
Why can't RAGFlow find a required file?
Check the complete Go error message and surrounding logs to identify the missing path.
- If a model file is missing, check whether the required model was downloaded successfully.
- If an uploaded document is missing, check the configured object storage service.
- If a temporary or local file is missing, check the container volume mappings and file permissions.
Use the system health API to verify the configured object storage and other dependencies.
Usage
How do I run RAGFlow with a locally deployed LLM?
You can use Ollama or Xinference to deploy a local LLM. See Deploy a local LLM for configuration details.
How do I add an LLM that is not directly supported?
If your model is not currently supported but has APIs compatible with those of OpenAI, click OpenAI-API-Compatible on the Model providers page to configure your model:
How do I change the file size limit?
The Go file-manager upload endpoint allows up to 1 GiB per file by default. The dataset document upload path has a separate 128 MiB per-file limit. Check which upload endpoint rejected the file before changing deployment settings.
To change the file-manager upload limit in a Go Compose deployment, set MAX_CONTENT_LENGTH in docker/.env to the desired number of bytes. The default 1 GiB value is 1073741824.
If you raise the limit above 1 GiB, also increase client_max_body_size in docker/nginx/nginx.conf. The Go image contains its own copy of this file, so make the change effective by either rebuilding RAGFLOW_IMAGE or enabling the ./nginx/nginx.conf:/etc/nginx/nginx.conf bind mount for the deployed service in docker/docker-compose.yml, then recreate that service. Lowering MAX_CONTENT_LENGTH below 1 GiB does not require changing the default Nginx limit.
These settings do not raise the dataset document upload path's fixed 128 MiB limit.
How do I get an API key for integration with third-party applications?
In the RAGFlow UI, click your avatar, open the API page, and copy or create an API key. Use that key to authenticate requests to the RAGFlow HTTP API.
How do I upgrade RAGFlow?
For a Go Compose deployment:
- Back up the database, object storage, and deployment configuration.
- From the repository root, stop the deployment without deleting its volumes:
docker compose --env-file docker/.env -f docker/docker-compose.yml down - Update the repository to the RAGFlow version you intend to deploy.
- Set
RAGFLOW_IMAGEin docker/.env to a compatible Go image. If you build from source, buildDockerfile, tag the image, and use that tag asRAGFLOW_IMAGE. - Pull the configured images when using registry images, then start the deployment:
docker compose --env-file docker/.env -f docker/docker-compose.yml pull
docker compose --env-file docker/.env -f docker/docker-compose.yml up -d - Check the service state and logs, then call
/api/v1/system/healthzthrough the configured web endpoint to verify the dependencies.
The Go Compose entrypoint runs the standalone ragflow_server --migrate action before starting the server processes. Starting ragflow_server --api, --admin, or --ingestor directly does not run the complete migration sequence. Keep the backup until the upgrade has been verified, and do not enable RAGFLOW_DEV_MODE in production to bypass downgrade protection. Never add -v to docker compose down unless you intend to delete the deployment's volumes.
How do I switch the document engine to Infinity?
To switch your Go deployment's document engine from Elasticsearch to Infinity, run these commands from the repository root:
Existing document indexes are not transferred to Infinity automatically. Back up your data and plan to reprocess or reindex documents after switching engines.
-
Stop the current Go Compose deployment without deleting its volumes:
docker compose --env-file docker/.env -f docker/docker-compose.yml down -
Set the following value in docker/.env. Its
COMPOSE_PROFILESsetting includesDOC_ENGINE, so this selects the Infinity service when the deployment starts:DOC_ENGINE=infinity -
Start the Go deployment and check its services:
docker compose --env-file docker/.env -f docker/docker-compose.yml up -d
docker compose --env-file docker/.env -f docker/docker-compose.yml ps -
Reprocess or reindex the documents so their searchable content is written to Infinity. Verify retrieval before retiring the old Elasticsearch data.
Where are uploaded files stored in RAGFlow?
Uploaded files are stored in the configured object storage backend. The Go backend uses MinIO by default and also supports Amazon S3, Alibaba Cloud OSS, and Google Cloud Storage.
The internal bucket and object path depend on how and where the file was uploaded. Manage uploaded files through RAGFlow instead of relying on a fixed storage path.
How do I tune document parsing and embedding throughput?
Document indexing processes 32 chunks per batch. The embedding request size comes from the model's batch_size capability, or defaults to 16 when the model does not provide one. To tune embedding requests, set TOKENIZER_EMBEDDING_BATCH_SIZE to a positive integer. Larger batches can use more memory and may exceed the model provider's request limit, so increase the value gradually and verify parsing on representative documents.
Why can't I retrieve relevant content in Search or Chat even though the document was parsed successfully?
Successful parsing only means that the document has been parsed and chunked. It does not guarantee that relevant content can be retrieved.
First, use Retrieval testing to check whether the expected chunks can be retrieved. If not, check the chunking results, embedding model, similarity threshold, and reranker settings. If the expected chunks are retrieved but Chat still gives an incorrect answer, check the Chat retrieval settings, system prompt, and LLM.
Why does the same query produce different results in Retrieval testing, Search, and Chat?
Retrieval testing is mainly used to evaluate whether relevant chunks can be retrieved from a dataset. Search further filters and ranks retrieved content according to its configuration, while Chat uses the retrieved content as context for an LLM to generate an answer.
Therefore, even when using the same dataset, differences in retrieval settings, reranking, and answer generation can lead to different results.
How can I tell whether a problem comes from document parsing, chunking, retrieval, or answer generation?
Check the RAG pipeline step by step:
- Verify that the document has been parsed correctly.
- Check whether the resulting chunks contain the expected content.
- Use Retrieval testing to verify that the relevant chunks can be retrieved.
- If retrieval works correctly but Chat or Agent produces an unexpected answer, check the application settings, system prompt, and LLM.
This helps identify which stage of the pipeline is causing the problem.
Can the same dataset be used by Search, Chat, and Agent?
Yes. The same dataset can be used by different knowledge applications and Agents.
The dataset manages and indexes the underlying knowledge, while Search, Chat, and Agent use that knowledge in different ways. Changing application-level settings generally does not modify the original documents or chunks in the dataset.
Should I adjust chunking, retrieval settings, or models first?
Start by verifying that the document is parsed and chunked correctly. Then use Retrieval testing to evaluate retrieval quality.
If retrieval needs improvement, adjust settings such as the similarity threshold, vector similarity weight, or reranker. If retrieval results are already relevant but the generated answer is still unsatisfactory, consider adjusting the prompt or changing the LLM.
What needs to be reprocessed after changing the embedding model, reranker, or LLM?
- When a dataset already contains chunks, changing the embedding model triggers a compatibility check: RAGFlow samples existing chunks, embeds them with the candidate model, and compares the new vectors with the stored vectors. The switch is allowed when the average similarity is at least 0.9 and does not require reparsing. If the models are incompatible, remove the existing chunks and parse the documents again, or create a new dataset with the new model.
- Changing the reranker affects only the ranking of retrieved results and does not require reparsing the documents.
- Changing the LLM affects query understanding and answer generation and does not normally require reparsing the dataset.
Why can the model still hallucinate when the answer exists in the dataset?
Having the correct information in the dataset does not guarantee that it will be retrieved or that the LLM will strictly follow the retrieved context.
Check whether the content has been indexed correctly, whether the relevant chunks are retrieved, and whether the model uses the retrieved context appropriately when generating its answer.
How do I reduce my chat assistant's response latency?
To reduce response latency, consider the following:
- Use a faster LLM with lower inference latency.
- Reduce the number of retrieved chunks by adjusting retrieval parameters such as Top N.
- Keep prompts concise and avoid unnecessary context.
- Disable optional features that require additional model calls when they are not needed.
- Use a reranker only when it provides a meaningful improvement in retrieval quality.
- Make sure the deployed model service has sufficient computing resources and low network latency.
How do I reduce my Agent's response latency?
Agent response time depends on the number of components, model calls, and external services involved in the workflow.
To improve response speed:
- Use faster models for components that do not require strong reasoning capabilities.
- Reduce unnecessary LLM, retrieval, tool, and HTTP calls.
- Simplify the Agent workflow and avoid excessively long execution paths.
- Limit the number of iterations in loops or reasoning-intensive components.
- Reduce the amount of context passed between components where possible.
- Run independent operations in parallel when the workflow supports it.
- Make sure external APIs and model services used by the Agent have low latency.
How do I use MinerU to parse PDF documents?
RAGFlow sends PDF documents to a remote MinerU service and polls for the parsed result. To use it:
- Prepare a reachable MinerU API service that provides the
/file_parseand/tasksendpoints used by RAGFlow. - In docker/.env or on the Model providers page in the UI, configure RAGFlow as a remote client to MinerU:
MINERU_APISERVER: The MinerU API endpoint (e.g.,http://mineru-host:8886).MINERU_API_KEY: An API key if the MinerU service requires one.MINERU_BACKEND: The MinerU backend (matches current MinerU API / CLI-bvalues):"pipeline"(default)"vlm-engine""hybrid-engine""vlm-http-client""hybrid-http-client".
- In the web UI, navigate to your dataset's Configuration page and find the Ingestion pipeline section:
- If you decide to use a chunking method from the Built-in dropdown, ensure it supports PDF parsing, then select MinerU from the PDF parser dropdown.
- If you use a custom ingestion pipeline instead, select MinerU in the PDF parser section of the Parser component.
You can configure MinerU through the Model providers page instead of setting environment variables. Values configured for the parser take precedence over MINERU_APISERVER, MINERU_API_KEY, and MINERU_BACKEND.
How do I configure MinerU-specific settings?
Use the following environment variables when MinerU is not configured directly in the parser:
| Environment variable | Description | Default | Example |
|---|---|---|---|
MINERU_APISERVER | URL of the MinerU API service | unset | MINERU_APISERVER=http://your-mineru-server:8886 |
MINERU_API_KEY | API key, when required by MinerU | unset | MINERU_API_KEY=your-key |
MINERU_BACKEND | MinerU parsing backend | pipeline | MINERU_BACKEND=pipeline|vlm-engine|hybrid-engine|vlm-http-client|hybrid-http-client |
- Set
MINERU_APISERVERto point RAGFlow to your MinerU API server. - Set
MINERU_API_KEYif your MinerU service requires authentication. - Set
MINERU_BACKENDto specify the backend sent with the parse request.
For a custom Parser component, the numeric mineru_timeout_seconds setup field controls how long RAGFlow polls for the result. Its default is 30 seconds; zero or a negative value also falls back to 30 seconds.
For other environment variables supported by MinerU itself, see the MinerU environment variable documentation.
How do I use MinerU with a vLLM server for document parsing?
Set MINERU_BACKEND to vlm-http-client or hybrid-http-client to use a downstream OpenAI-compatible server such as vLLM. Configure the downstream server URL on the MinerU service itself.
- Ensure a MinerU API service is reachable (for example
http://mineru-host:8886). - Configure the MinerU service to use a reachable OpenAI-compatible server.
- Configure the following in docker/.env (or your shell if running from source):
MINERU_APISERVER=http://mineru-host:8886MINERU_BACKEND="vlm-http-client"(or"hybrid-http-client")
- Select MinerU as the PDF parser in the dataset's ingestion settings, then check the ingestor logs if the remote parse request fails.
For these remote backends, the RAGFlow ingestor needs network access to MinerU. The downstream model service is configured and reached by MinerU.
How do I use an external Docling Serve server for document parsing?
The Go parser uses a remote Docling Serve endpoint. Set the endpoint in the parser configuration (docling_server_url) or provide it through DOCLING_SERVER_URL in docker/.env. If the service requires bearer authentication, also set docling_api_key in the parser configuration or provide DOCLING_API_KEY:
DOCLING_SERVER_URL=http://your-docling-serve-host:5001
DOCLING_API_KEY=your-api-key
The Go parser sends PDFs to Docling Serve using /v1/convert/source and also tries /v1alpha/convert/source for older servers. It sends DOCLING_API_KEY as a Bearer token. If neither the parser configuration nor the environment variable supplies a URL, Docling parsing fails with a configuration error. Leave the API key unset when the service does not require authentication.
How do I use PaddleOCR for document parsing?
RAGFlow includes PaddleOCR as an optional remote PDF parser. The Go implementation supports both the asynchronous PaddleOCR Job API and a synchronous self-hosted PaddleOCR.local service.
There are two main ways to configure and use PaddleOCR in RAGFlow:
1. Using the official PaddleOCR API
This method uses PaddleOCR's official API service with an access token.
Step 1: Configure RAGFlow
-
Via Environment Variables:
# In your docker/.env file:
PADDLEOCR_API_URL=https://paddleocr.aistudio-app.com/api
PADDLEOCR_ALGORITHM=PaddleOCR-VL
PADDLEOCR_ACCESS_TOKEN=your-access-token-here -
Via UI:
- Navigate to Model providers page
- Add a new OCR model with factory type "PaddleOCR"
- Configure the following fields:
- PaddleOCR API URL: The API base URL, such as
https://paddleocr.aistudio-app.com/api, without the/v2/ocr/jobssuffix - PaddleOCR Algorithm: Select the algorithm corresponding to the API endpoint
- AI Studio Access Token: Your access token for the PaddleOCR API
- PaddleOCR API URL: The API base URL, such as
Step 2: Usage in Dataset Configuration
- In your dataset's Configuration page, find the Ingestion pipeline section
- If using built-in chunking methods that support PDF parsing, select PaddleOCR from the PDF parser dropdown
- If using custom ingestion pipeline, select PaddleOCR in the Parser component
Notes:
- When configuring the asynchronous PaddleOCR model provider through
PADDLEOCR_API_URLor the Model providers page, use an API base URL that includes/api, such ashttps://paddleocr.aistudio-app.com/api, but does not include/v2/ocr/jobs. The provider appends/v2/ocr/jobs. - When configuring the PDF parser directly with
paddleocr_base_urlin a custom Parser component orPADDLEOCR_BASE_URLin the environment, use the service origin without/api, such ashttps://paddleocr.aistudio-app.com. The direct parser appends/api/v2/ocr/jobs. - Access tokens can be obtained from the AI Studio platform.
- This method requires internet connectivity to reach the official PaddleOCR API.
2. Using a self-hosted PaddleOCR service
For a synchronous self-hosted PaddleOCR service, add the PaddleOCR.local provider and configure its base URL. The current Go provider appends /layout-parsing to that URL.
A self-hosted service that implements the asynchronous PaddleOCR Job API can instead use the regular PaddleOCR provider described above.
Step 1: Deploy PaddleOCR Service
Provide a PaddleOCR service reachable from the RAGFlow container. For a synchronous PaddleOCR.local service running on the Docker host, an example base URL is:
http://host.docker.internal:8080
Enter the base URL without the /layout-parsing suffix because the Go provider appends it when sending the request.
Step 2: Configure RAGFlow
-
Via UI:
- Navigate to Model providers page
- Add a new OCR model with factory type PaddleOCR.local
- Configure the following fields:
- PaddleOCR API URL: The base URL of your self-hosted service
- AI Studio Access Token: Set a bearer token if your service requires one; otherwise leave it empty
Step 3: Usage in Dataset Configuration
- In your dataset's Configuration page, find the Ingestion pipeline section
- If using built-in chunking methods that support PDF parsing, select PaddleOCR from the PDF parser dropdown
- If using custom ingestion pipeline, select PaddleOCR in the Parser component
Asynchronous PaddleOCR environment variable summary
| Environment Variable | Description | Default | Required |
|---|---|---|---|
PADDLEOCR_API_URL | Asynchronous model-provider base URL, including /api but without /v2/ocr/jobs | unset | Yes, when configuring the model provider through environment variables |
PADDLEOCR_BASE_URL | Direct PDF-parser service origin, without /api/v2/ocr/jobs | unset | Yes, when configuring the parser directly through environment variables |
PADDLEOCR_ALGORITHM | Algorithm to use for parsing | "PaddleOCR-VL" | No |
PADDLEOCR_ACCESS_TOKEN | Bearer token for the PaddleOCR service | unset | When the service requires authentication |
PADDLEOCR_API_URL configures and auto-provisions the asynchronous PaddleOCR model provider. PADDLEOCR_BASE_URL is the fallback used by the direct PDF parser. These variables are not required when configuring a provider through the UI and do not select PaddleOCR.local.
How do I use Ollama with RAGFlow for local LLM inference?
RAGFlow supports Ollama as a local model provider for private, offline inference.
Step 1: Start Ollama and pull a model
Start Ollama in one terminal:
export OLLAMA_HOST=0.0.0.0
ollama serve
Binding Ollama to 0.0.0.0 makes it reachable on every network interface. Do not expose this port directly to the public internet; restrict access with a firewall or private network.
While the service is running, pull the model in another terminal:
ollama pull llama3
Step 2: Add Ollama in RAGFlow
- Go to Settings > Model providers > Ollama.
- Set the Base URL to
http://host.docker.internal:11434for the provided Go Compose deployment, orhttp://localhost:11434when RAGFlow and Ollama run directly on the same host. If you use a different container setup, use an address resolvable and reachable from the RAGFlow container. - Enter the model name (e.g.,
llama3) and click Save.
Step 3: Use Ollama in your assistant
- Open an assistant's Configuration page and select the Ollama model under Chat model.