Skip to main content

FAQs

Answers to questions about general features, troubleshooting, usage, and more.


General features


What sets RAGFlow apart from other RAG products?

RAGFlow provides an end-to-end RAG platform that goes beyond basic document chunking and retrieval. Its key strengths include:

  • Deep document understanding for complex content such as text, tables, and images.
  • Flexible retrieval and traceable answers with multiple retrieval strategies, reranking, and citations.
  • Agentic capabilities for multi-step retrieval, reasoning, tool use, and workflows.
  • Knowledge compilation for organizing documents into structured knowledge artifacts.
  • Broad model and integration support for building different RAG applications.

Where can I find the RAGFlow version, and how do I interpret it?

You can find the RAGFlow version number on the System page of the UI:

Image

If you build RAGFlow from source, the Go admin and ingestor log their versions at startup. The server also prints its version with bin/ragflow_server --api --version.

For example, v0.27.1-14-g6daae7f2 means 14 commits after the v0.27.1 tag, at a commit whose abbreviated ID is 6daae7f2. The actual value comes from the packaged VERSION file or, when that file is absent, from git describe; it may be a tag alone or unknown if neither source supplies a version.


Why does RAGFlow use Elasticsearch as its default document engine?

Elasticsearch meets RAGFlow's core hybrid search requirements, including full-text search, vector search, phrase search, and advanced ranking capabilities.

The Go backend also supports Infinity as a document engine. Infinity is an AI-native database developed by InfiniFlow and optimized for RAG workloads. Support levels and available features may vary between document engines.


What are the differences between cloud.ragflow.io and a locally deployed open-source RAGFlow service?

cloud.ragflow.io is the hosted RAGFlow service. It provides managed infrastructure and subscription-based limits and features, while a locally deployed open-source service gives you control over deployment, data, models, storage, and system resources.

REST API access on cloud.ragflow.io depends on the subscription plan. Check the current plan details on the RAGFlow website before relying on API access. A locally deployed service exposes the RAGFlow HTTP APIs directly and uses API keys created in its UI.


Why does RAGFlow require substantial system resources?

RAGFlow runs multiple components for document parsing, embedding, full-text and vector indexing, retrieval, task processing, metadata storage, caching, and object storage. Some document parsers also load or download machine-learning models and may require significant CPU and memory resources.

Actual resource usage depends on the selected document engine, parser, model provider, document volume, and workload. See the quickstart prerequisites for the recommended starting configuration.


Which architectures or devices does RAGFlow support?

The documented Go image build targets Linux AMD64. Building on another architecture requires a matching dependency image and native libraries, and the selected document engine must support that architecture. The commands in the quickstart guide target AMD64.


Do you offer an API for integration with third-party applications?

See the RAGFlow HTTP API Reference for integration endpoints.


Do you support stream output?

Yes. RAGFlow supports both streaming and non-streaming responses. The interactive Chat and Agent pages display responses as they are generated, while the embedded chat configuration provides an Enable streaming responses option.

For HTTP API calls, use the endpoint's stream parameter to select the response mode. Defaults can differ between endpoints, so check the corresponding API reference:


What are the key differences between Search and Chat?

  • Search is designed for direct knowledge retrieval. It retrieves and ranks relevant content from one or more datasets and presents the results for users to review.
  • Chat is designed for conversational question answering. It retrieves relevant knowledge and uses an LLM to generate answers, supporting multi-turn conversations and more advanced retrieval capabilities such as Agentic Retrieval.

Use Search when you want to find and inspect relevant knowledge directly, and Chat when you want generated answers based on that knowledge.


Troubleshooting


How do I build a RAGFlow image from scratch?

Build the Go image from the repository root with Dockerfile, then set RAGFLOW_IMAGE in docker/.env. See the quickstart guide for the current commands and prerequisites.


Why does PDF parsing fail when DeepDoc model files are missing?

The Go API and ingestor initialize the in-process DeepDoc backend. Check their logs for model initialization errors and verify that the configured DEEPDOC_MODEL_DIR contains the required model files, including det.ort, layout.ort, tsr.ort, rec.ort, and ocr.res. If you use the shared model asset directory, check MODEL_ASSETS_DIR instead. Restore the missing files from the deployment's model assets, then restart the affected service.


Does the open-source 1.0 Go DeepDoc backend support GPU inference?

No. The RAGFlow open-source 1.0 Go DeepDoc backend runs layout analysis, OCR, and table recognition on CPU.


Why can't RAGFlow access my Ollama model?

Check that Ollama is running, the model has been downloaded, the model name is correct, and the Ollama Base URL is reachable from the RAGFlow container.

If the request times out while loading the model, check the Ollama logs and available memory. Try a smaller model or allocate more memory if necessary.

See Deploy a local LLM for more information.

For setup instructions, see How do I use Ollama with RAGFlow for local LLM inference?


Why can't Nginx reach the Go API?

Check that the Go API and Admin services started, then review the RAGFlow container and Nginx logs. In the target Go configuration, Nginx forwards API requests to port 9380 and Admin requests to port 9381. Check the active ragflow.conf against docker/nginx/ragflow.conf.golang and verify that those services are listening before changing the proxy configuration.


What does network anomaly: There is an abnormality in your network and you cannot connect to the server mean?

anomaly

You cannot log in until the services are initialized. Review the RAGFlow container logs (docker compose --env-file docker/.env -f docker/docker-compose.yml logs -f ragflow-cpu from the repository root). The Go services log these startup messages:

Starting api server: ...
Server starting on port: 9380
Starting admin server: ...
Starting RAGFlow admin HTTP server on port: 9381
Starting ingestor server: ...
Ingestor ... initialized
RAGFlow ingestion service version: ...

These messages come from separate processes and may appear in a different order. Check the logs for startup errors, use the system health API to inspect dependencies, and verify that the ingestor has registered with admin before treating document parsing as ready.


Why are document tasks waiting in the queue?

The Go ingestor consumes document tasks from NATS JetStream. A growing queue can mean that no ingestor is consuming, workers are busy, or tasks are repeatedly failing.

  1. Check that the NATS and RAGFlow containers are running, and inspect their logs for connection or consumer errors.
  2. Check the ingestor logs for Ingestor ... initialized, NATS stream RAGFLOW_TASKS ready, pull errors, and task failures. Confirm that the ingestor is registered with admin.
  3. Check whether workers are processing tasks and whether the task status or document progress changes. Investigate the first task error before retrying it.

Do not clear the queue as a first troubleshooting step: doing so can discard pending work.


Why does document parsing stall below 1%?

stall

Click the red cross beside the 'parsing status' bar, then restart the parsing process to see if the issue remains. If the issue persists and your RAGFlow is deployed locally, try the following:

  1. Check the RAGFlow container logs for API, admin, and ingestor startup errors:

    docker compose --env-file docker/.env -f docker/docker-compose.yml logs -f ragflow-cpu
  2. Confirm that the Go ingestor is running and sending heartbeats to admin. Check admin's service status for the ingestor; an API health response alone does not confirm that workers are consuming tasks.

  3. Check the NATS connection and JetStream consumer in the logs. Look for pull errors, consumer initialization failures, and repeated task failures.

  4. Check the parser and tokenizer assets separately. Missing DeepDoc model files prevent the API or ingestor from starting because the in-process DeepDoc backend is required. A missing cl100k_base.tiktoken table does not normally prevent startup, but ingestion fails for models that declare that tokenizer until the table is restored.

  5. Check whether ingestor workers are processing tasks and whether task progress changes. If a worker stops making progress, inspect its task error and available memory before retrying.


Why does PDF parsing stall near completion without any errors in the logs?

Click the red cross beside the 'parsing status' bar, then restart the parsing process to see if the issue remains. If the issue persists, check the ingestor and container logs for an out-of-memory error. Make sure the Docker runtime or host has enough memory available for the RAGFlow service, and reduce concurrent parsing work if necessary.

note

The standard Go Compose file does not apply MEM_LIMIT to ragflow-cpu; that variable limits selected dependency services. Changing it does not increase the memory available to the Go ingestor. Adjust the Docker runtime or host allocation, or add an explicit service-level memory limit in a Compose override if your environment requires one.

nearcompletion


What does Index failure mean?

An index failure means that RAGFlow could not write the processed document chunks to the configured document engine.

Check the task logs, embedding model, document engine connection, and system health. If you use Elasticsearch, Infinity, or another document engine supported by the current Go backend, verify that the corresponding service is running and accessible.


How do I check RAGFlow logs?

From the repository root, follow the Go service logs written to the mounted log directory:

tail -f docker/ragflow-logs/*.log

To view the container's combined output, use docker compose --env-file docker/.env -f docker/docker-compose.yml logs -f ragflow-cpu.


How do I check the status of each RAGFlow component?

  1. From the repository root, check the Go deployment's containers:

    docker compose --env-file docker/.env -f docker/docker-compose.yml ps

    Review the RAGFlow service logs:

    docker compose --env-file docker/.env -f docker/docker-compose.yml logs -f ragflow-cpu

    Check the nats and selected document engine services in the same Compose project when diagnosing ingestion or indexing.

  2. Use the system health API to check the database, Kvrocks cache, document engine, object storage, and NATS message queue.

A running container does not necessarily mean that the service inside it is healthy. Check the health API and logs for connection, port, DNS, and configuration errors.


Why can't RAGFlow connect to Elasticsearch?

  1. Check the status of the Elasticsearch Docker container:

    docker compose --env-file docker/.env -f docker/docker-compose.yml ps es01

    A healthy Elasticsearch container should report healthy. The default Go Compose deployment exposes Elasticsearch on host port 1200, while the container listens on port 9200.

  2. Follow the system health API to check the health status of the Elasticsearch service.

    IMPORTANT

    The status of a Docker container does not necessarily reflect the status of the service inside it. A service can be unhealthy even when its container is up and running. Possible causes include network failures, incorrect ports, and DNS or configuration errors.

  3. If your container keeps restarting, ensure vm.max_map_count >= 262144. On Linux, update /etc/sysctl.conf to keep the change permanent. For macOS, see the following FAQ.


Why does the Elasticsearch container fail to start with Elasticsearch did not exit normally?

On Linux, an insufficient vm.max_map_count value is a common cause. Check the Elasticsearch logs first. If they report that vm.max_map_count is too low, set it to at least 262144 and add the setting to /etc/sysctl.conf so it persists after a reboot. If the logs report a different error, check memory, disk space, permissions, and the Elasticsearch data directory instead.


How do I configure vm.max_map_count on macOS?

vm.max_map_count is a Linux kernel parameter required only by Elasticsearch. It does not affect RAGFlow deployments that use Infinity as the document engine.

On Docker Desktop, set the value in its Linux virtual machine:

docker run --rm --privileged alpine sysctl -w vm.max_map_count=262144

This setting is temporary and is reset when Docker Desktop restarts.

On Colima, check the current value inside its virtual machine:

colima ssh -- sysctl vm.max_map_count

To set the value temporarily:

colima ssh -- sudo sysctl -w vm.max_map_count=262144

For a persistent setting, run colima start --edit and add a provision script to colima.yaml:

provision:
- mode: system
script: |
#!/bin/bash
sysctl -w vm.max_map_count=262144

Then restart Colima:

colima stop
colima start

Contributed by @helloxjade.


Why does the Go API return a 404 response?

The Go API returns HTTP 404 with Not Found: <path> when a request does not match a route. Your URL may point to the wrong service.

For normal UI and API access, use http://<IP_OF_YOUR_MACHINE> when SVR_WEB_HTTP_PORT is 80, or include the configured web port. Nginx forwards API traffic to the Go API on internal port 9380 and Admin traffic to internal port 9381.

The Go Compose deployment uses SVR_HTTP_PORT and ADMIN_SVR_HTTP_PORT for the directly published API and Admin host ports. For normal browser and API access, prefer the Nginx web port.


Do you provide examples of using DeepDoc to parse PDFs or other files?

Yes. The Go parser implementations are in internal/parser/parser, and the Go DeepDoc integration is in internal/deepdoc. The following tests provide concrete usage examples:

RAGFlow's ingestor uses these components when it processes documents. These files are integration or end-to-end tests, so review their prerequisites and build constraints before running them directly.


Why can't RAGFlow find a required file?

Check the complete Go error message and surrounding logs to identify the missing path.

  • If a model file is missing, check whether the required model was downloaded successfully.
  • If an uploaded document is missing, check the configured object storage service.
  • If a temporary or local file is missing, check the container volume mappings and file permissions.

Use the system health API to verify the configured object storage and other dependencies.


Usage


How do I run RAGFlow with a locally deployed LLM?

You can use Ollama or Xinference to deploy a local LLM. See Deploy a local LLM for configuration details.


How do I add an LLM that is not directly supported?

If your model is not currently supported but has APIs compatible with those of OpenAI, click OpenAI-API-Compatible on the Model providers page to configure your model:

openai-api-compatible


How do I change the file size limit?

The Go file-manager upload endpoint allows up to 1 GiB per file by default. The dataset document upload path has a separate 128 MiB per-file limit. Check which upload endpoint rejected the file before changing deployment settings.

To change the file-manager upload limit in a Go Compose deployment, set MAX_CONTENT_LENGTH in docker/.env to the desired number of bytes. The default 1 GiB value is 1073741824.

If you raise the limit above 1 GiB, also increase client_max_body_size in docker/nginx/nginx.conf. The Go image contains its own copy of this file, so make the change effective by either rebuilding RAGFLOW_IMAGE or enabling the ./nginx/nginx.conf:/etc/nginx/nginx.conf bind mount for the deployed service in docker/docker-compose.yml, then recreate that service. Lowering MAX_CONTENT_LENGTH below 1 GiB does not require changing the default Nginx limit.

These settings do not raise the dataset document upload path's fixed 128 MiB limit.


How do I get an API key for integration with third-party applications?

In the RAGFlow UI, click your avatar, open the API page, and copy or create an API key. Use that key to authenticate requests to the RAGFlow HTTP API.


How do I upgrade RAGFlow?

For a Go Compose deployment:

  1. Back up the database, object storage, and deployment configuration.
  2. From the repository root, stop the deployment without deleting its volumes:
    docker compose --env-file docker/.env -f docker/docker-compose.yml down
  3. Update the repository to the RAGFlow version you intend to deploy.
  4. Set RAGFLOW_IMAGE in docker/.env to a compatible Go image. If you build from source, build Dockerfile, tag the image, and use that tag as RAGFLOW_IMAGE.
  5. Pull the configured images when using registry images, then start the deployment:
    docker compose --env-file docker/.env -f docker/docker-compose.yml pull
    docker compose --env-file docker/.env -f docker/docker-compose.yml up -d
  6. Check the service state and logs, then call /api/v1/system/healthz through the configured web endpoint to verify the dependencies.

The Go Compose entrypoint runs the standalone ragflow_server --migrate action before starting the server processes. Starting ragflow_server --api, --admin, or --ingestor directly does not run the complete migration sequence. Keep the backup until the upgrade has been verified, and do not enable RAGFLOW_DEV_MODE in production to bypass downgrade protection. Never add -v to docker compose down unless you intend to delete the deployment's volumes.


How do I switch the document engine to Infinity?

To switch your Go deployment's document engine from Elasticsearch to Infinity, run these commands from the repository root:

WARNING

Existing document indexes are not transferred to Infinity automatically. Back up your data and plan to reprocess or reindex documents after switching engines.

  1. Stop the current Go Compose deployment without deleting its volumes:

    docker compose --env-file docker/.env -f docker/docker-compose.yml down
  2. Set the following value in docker/.env. Its COMPOSE_PROFILES setting includes DOC_ENGINE, so this selects the Infinity service when the deployment starts:

    DOC_ENGINE=infinity
  3. Start the Go deployment and check its services:

    docker compose --env-file docker/.env -f docker/docker-compose.yml up -d
    docker compose --env-file docker/.env -f docker/docker-compose.yml ps
  4. Reprocess or reindex the documents so their searchable content is written to Infinity. Verify retrieval before retiring the old Elasticsearch data.


Where are uploaded files stored in RAGFlow?

Uploaded files are stored in the configured object storage backend. The Go backend uses MinIO by default and also supports Amazon S3, Alibaba Cloud OSS, and Google Cloud Storage.

The internal bucket and object path depend on how and where the file was uploaded. Manage uploaded files through RAGFlow instead of relying on a fixed storage path.


How do I tune document parsing and embedding throughput?

Document indexing processes 32 chunks per batch. The embedding request size comes from the model's batch_size capability, or defaults to 16 when the model does not provide one. To tune embedding requests, set TOKENIZER_EMBEDDING_BATCH_SIZE to a positive integer. Larger batches can use more memory and may exceed the model provider's request limit, so increase the value gradually and verify parsing on representative documents.


Why can't I retrieve relevant content in Search or Chat even though the document was parsed successfully?

Successful parsing only means that the document has been parsed and chunked. It does not guarantee that relevant content can be retrieved.

First, use Retrieval testing to check whether the expected chunks can be retrieved. If not, check the chunking results, embedding model, similarity threshold, and reranker settings. If the expected chunks are retrieved but Chat still gives an incorrect answer, check the Chat retrieval settings, system prompt, and LLM.


Why does the same query produce different results in Retrieval testing, Search, and Chat?

Retrieval testing is mainly used to evaluate whether relevant chunks can be retrieved from a dataset. Search further filters and ranks retrieved content according to its configuration, while Chat uses the retrieved content as context for an LLM to generate an answer.

Therefore, even when using the same dataset, differences in retrieval settings, reranking, and answer generation can lead to different results.


How can I tell whether a problem comes from document parsing, chunking, retrieval, or answer generation?

Check the RAG pipeline step by step:

  1. Verify that the document has been parsed correctly.
  2. Check whether the resulting chunks contain the expected content.
  3. Use Retrieval testing to verify that the relevant chunks can be retrieved.
  4. If retrieval works correctly but Chat or Agent produces an unexpected answer, check the application settings, system prompt, and LLM.

This helps identify which stage of the pipeline is causing the problem.


Can the same dataset be used by Search, Chat, and Agent?

Yes. The same dataset can be used by different knowledge applications and Agents.

The dataset manages and indexes the underlying knowledge, while Search, Chat, and Agent use that knowledge in different ways. Changing application-level settings generally does not modify the original documents or chunks in the dataset.


Should I adjust chunking, retrieval settings, or models first?

Start by verifying that the document is parsed and chunked correctly. Then use Retrieval testing to evaluate retrieval quality.

If retrieval needs improvement, adjust settings such as the similarity threshold, vector similarity weight, or reranker. If retrieval results are already relevant but the generated answer is still unsatisfactory, consider adjusting the prompt or changing the LLM.


What needs to be reprocessed after changing the embedding model, reranker, or LLM?

  • When a dataset already contains chunks, changing the embedding model triggers a compatibility check: RAGFlow samples existing chunks, embeds them with the candidate model, and compares the new vectors with the stored vectors. The switch is allowed when the average similarity is at least 0.9 and does not require reparsing. If the models are incompatible, remove the existing chunks and parse the documents again, or create a new dataset with the new model.
  • Changing the reranker affects only the ranking of retrieved results and does not require reparsing the documents.
  • Changing the LLM affects query understanding and answer generation and does not normally require reparsing the dataset.

Why can the model still hallucinate when the answer exists in the dataset?

Having the correct information in the dataset does not guarantee that it will be retrieved or that the LLM will strictly follow the retrieved context.

Check whether the content has been indexed correctly, whether the relevant chunks are retrieved, and whether the model uses the retrieved context appropriately when generating its answer.


How do I reduce my chat assistant's response latency?

To reduce response latency, consider the following:

  • Use a faster LLM with lower inference latency.
  • Reduce the number of retrieved chunks by adjusting retrieval parameters such as Top N.
  • Keep prompts concise and avoid unnecessary context.
  • Disable optional features that require additional model calls when they are not needed.
  • Use a reranker only when it provides a meaningful improvement in retrieval quality.
  • Make sure the deployed model service has sufficient computing resources and low network latency.

How do I reduce my Agent's response latency?

Agent response time depends on the number of components, model calls, and external services involved in the workflow.

To improve response speed:

  • Use faster models for components that do not require strong reasoning capabilities.
  • Reduce unnecessary LLM, retrieval, tool, and HTTP calls.
  • Simplify the Agent workflow and avoid excessively long execution paths.
  • Limit the number of iterations in loops or reasoning-intensive components.
  • Reduce the amount of context passed between components where possible.
  • Run independent operations in parallel when the workflow supports it.
  • Make sure external APIs and model services used by the Agent have low latency.

How do I use MinerU to parse PDF documents?

RAGFlow sends PDF documents to a remote MinerU service and polls for the parsed result. To use it:

  1. Prepare a reachable MinerU API service that provides the /file_parse and /tasks endpoints used by RAGFlow.
  2. In docker/.env or on the Model providers page in the UI, configure RAGFlow as a remote client to MinerU:
    • MINERU_APISERVER: The MinerU API endpoint (e.g., http://mineru-host:8886).
    • MINERU_API_KEY: An API key if the MinerU service requires one.
    • MINERU_BACKEND: The MinerU backend (matches current MinerU API / CLI -b values):
      • "pipeline" (default)
      • "vlm-engine"
      • "hybrid-engine"
      • "vlm-http-client"
      • "hybrid-http-client".
  3. In the web UI, navigate to your dataset's Configuration page and find the Ingestion pipeline section:
    • If you decide to use a chunking method from the Built-in dropdown, ensure it supports PDF parsing, then select MinerU from the PDF parser dropdown.
    • If you use a custom ingestion pipeline instead, select MinerU in the PDF parser section of the Parser component.
note

You can configure MinerU through the Model providers page instead of setting environment variables. Values configured for the parser take precedence over MINERU_APISERVER, MINERU_API_KEY, and MINERU_BACKEND.


How do I configure MinerU-specific settings?

Use the following environment variables when MinerU is not configured directly in the parser:

Environment variableDescriptionDefaultExample
MINERU_APISERVERURL of the MinerU API serviceunsetMINERU_APISERVER=http://your-mineru-server:8886
MINERU_API_KEYAPI key, when required by MinerUunsetMINERU_API_KEY=your-key
MINERU_BACKENDMinerU parsing backendpipelineMINERU_BACKEND=pipeline|vlm-engine|hybrid-engine|vlm-http-client|hybrid-http-client
  1. Set MINERU_APISERVER to point RAGFlow to your MinerU API server.
  2. Set MINERU_API_KEY if your MinerU service requires authentication.
  3. Set MINERU_BACKEND to specify the backend sent with the parse request.

For a custom Parser component, the numeric mineru_timeout_seconds setup field controls how long RAGFlow polls for the result. Its default is 30 seconds; zero or a negative value also falls back to 30 seconds.

NOTE

For other environment variables supported by MinerU itself, see the MinerU environment variable documentation.


How do I use MinerU with a vLLM server for document parsing?

Set MINERU_BACKEND to vlm-http-client or hybrid-http-client to use a downstream OpenAI-compatible server such as vLLM. Configure the downstream server URL on the MinerU service itself.

  1. Ensure a MinerU API service is reachable (for example http://mineru-host:8886).
  2. Configure the MinerU service to use a reachable OpenAI-compatible server.
  3. Configure the following in docker/.env (or your shell if running from source):
    • MINERU_APISERVER=http://mineru-host:8886
    • MINERU_BACKEND="vlm-http-client" (or "hybrid-http-client")
  4. Select MinerU as the PDF parser in the dataset's ingestion settings, then check the ingestor logs if the remote parse request fails.
NOTE

For these remote backends, the RAGFlow ingestor needs network access to MinerU. The downstream model service is configured and reached by MinerU.


How do I use an external Docling Serve server for document parsing?

The Go parser uses a remote Docling Serve endpoint. Set the endpoint in the parser configuration (docling_server_url) or provide it through DOCLING_SERVER_URL in docker/.env. If the service requires bearer authentication, also set docling_api_key in the parser configuration or provide DOCLING_API_KEY:

DOCLING_SERVER_URL=http://your-docling-serve-host:5001
DOCLING_API_KEY=your-api-key

The Go parser sends PDFs to Docling Serve using /v1/convert/source and also tries /v1alpha/convert/source for older servers. It sends DOCLING_API_KEY as a Bearer token. If neither the parser configuration nor the environment variable supplies a URL, Docling parsing fails with a configuration error. Leave the API key unset when the service does not require authentication.


How do I use PaddleOCR for document parsing?

RAGFlow includes PaddleOCR as an optional remote PDF parser. The Go implementation supports both the asynchronous PaddleOCR Job API and a synchronous self-hosted PaddleOCR.local service.

There are two main ways to configure and use PaddleOCR in RAGFlow:

1. Using the official PaddleOCR API

This method uses PaddleOCR's official API service with an access token.

Step 1: Configure RAGFlow

  • Via Environment Variables:

    # In your docker/.env file:
    PADDLEOCR_API_URL=https://paddleocr.aistudio-app.com/api
    PADDLEOCR_ALGORITHM=PaddleOCR-VL
    PADDLEOCR_ACCESS_TOKEN=your-access-token-here
  • Via UI:

    • Navigate to Model providers page
    • Add a new OCR model with factory type "PaddleOCR"
    • Configure the following fields:
      • PaddleOCR API URL: The API base URL, such as https://paddleocr.aistudio-app.com/api, without the /v2/ocr/jobs suffix
      • PaddleOCR Algorithm: Select the algorithm corresponding to the API endpoint
      • AI Studio Access Token: Your access token for the PaddleOCR API

Step 2: Usage in Dataset Configuration

  • In your dataset's Configuration page, find the Ingestion pipeline section
  • If using built-in chunking methods that support PDF parsing, select PaddleOCR from the PDF parser dropdown
  • If using custom ingestion pipeline, select PaddleOCR in the Parser component

Notes:

  • When configuring the asynchronous PaddleOCR model provider through PADDLEOCR_API_URL or the Model providers page, use an API base URL that includes /api, such as https://paddleocr.aistudio-app.com/api, but does not include /v2/ocr/jobs. The provider appends /v2/ocr/jobs.
  • When configuring the PDF parser directly with paddleocr_base_url in a custom Parser component or PADDLEOCR_BASE_URL in the environment, use the service origin without /api, such as https://paddleocr.aistudio-app.com. The direct parser appends /api/v2/ocr/jobs.
  • Access tokens can be obtained from the AI Studio platform.
  • This method requires internet connectivity to reach the official PaddleOCR API.

2. Using a self-hosted PaddleOCR service

For a synchronous self-hosted PaddleOCR service, add the PaddleOCR.local provider and configure its base URL. The current Go provider appends /layout-parsing to that URL.

A self-hosted service that implements the asynchronous PaddleOCR Job API can instead use the regular PaddleOCR provider described above.

Step 1: Deploy PaddleOCR Service

Provide a PaddleOCR service reachable from the RAGFlow container. For a synchronous PaddleOCR.local service running on the Docker host, an example base URL is:

http://host.docker.internal:8080

Enter the base URL without the /layout-parsing suffix because the Go provider appends it when sending the request.

Step 2: Configure RAGFlow

  • Via UI:

    • Navigate to Model providers page
    • Add a new OCR model with factory type PaddleOCR.local
    • Configure the following fields:
      • PaddleOCR API URL: The base URL of your self-hosted service
      • AI Studio Access Token: Set a bearer token if your service requires one; otherwise leave it empty

Step 3: Usage in Dataset Configuration

  • In your dataset's Configuration page, find the Ingestion pipeline section
  • If using built-in chunking methods that support PDF parsing, select PaddleOCR from the PDF parser dropdown
  • If using custom ingestion pipeline, select PaddleOCR in the Parser component

Asynchronous PaddleOCR environment variable summary

Environment VariableDescriptionDefaultRequired
PADDLEOCR_API_URLAsynchronous model-provider base URL, including /api but without /v2/ocr/jobsunsetYes, when configuring the model provider through environment variables
PADDLEOCR_BASE_URLDirect PDF-parser service origin, without /api/v2/ocr/jobsunsetYes, when configuring the parser directly through environment variables
PADDLEOCR_ALGORITHMAlgorithm to use for parsing"PaddleOCR-VL"No
PADDLEOCR_ACCESS_TOKENBearer token for the PaddleOCR serviceunsetWhen the service requires authentication

PADDLEOCR_API_URL configures and auto-provisions the asynchronous PaddleOCR model provider. PADDLEOCR_BASE_URL is the fallback used by the direct PDF parser. These variables are not required when configuring a provider through the UI and do not select PaddleOCR.local.


How do I use Ollama with RAGFlow for local LLM inference?

RAGFlow supports Ollama as a local model provider for private, offline inference.

Step 1: Start Ollama and pull a model

Start Ollama in one terminal:

export OLLAMA_HOST=0.0.0.0
ollama serve

Binding Ollama to 0.0.0.0 makes it reachable on every network interface. Do not expose this port directly to the public internet; restrict access with a firewall or private network.

While the service is running, pull the model in another terminal:

ollama pull llama3

Step 2: Add Ollama in RAGFlow

  1. Go to Settings > Model providers > Ollama.
  2. Set the Base URL to http://host.docker.internal:11434 for the provided Go Compose deployment, or http://localhost:11434 when RAGFlow and Ollama run directly on the same host. If you use a different container setup, use an address resolvable and reachable from the RAGFlow container.
  3. Enter the model name (e.g., llama3) and click Save.

Step 3: Use Ollama in your assistant

  • Open an assistant's Configuration page and select the Ollama model under Chat model.
On this page