One document, four destinations
OpenCrawling now ships an Apache SeaTunnel Output Connector, oc-seatunnel-output-connector, built on SeaTunnel 2.3.13 and its Zeta Engine . This connector solves one of the problems I’ve watched repeat across enterprise search engagements. A crawler used to have one job: extract, chunk, embed, and write to a single index. That pattern breaks the moment a RAG program grows up or any time there are multiple consumers.
In a real enterprise deployment, the same chunk has to land in several places at once:
- Vector stores such as Elastic, OpenSearch, Milvus, Qdrant, or Vespa, serving dense embeddings for sub-50ms kNN retrieval.
- Analytical databases such as ClickHouse or StarRocks, holding ingestion telemetry and token tracking for cost governance.
- Lakehouses on Apache Iceberg or Amazon S3, archiving raw content and chunks for long-term governance.
- Lexical indexes in Solr, OpenSearch, or Luxir, kept in sync for exact titles and keyword facets.
The default answer is a separate point-to-point sync job per destination. That means hitting the source repositories more than once, and if embedding, paying for that more than once.
Why SeaTunnel
Apache SeaTunnel is an open-source data integration platform. A job is a pipeline of sources, optional transforms, and sinks, defined in a config file and run in batch or streaming mode. The engine handles parallelism, checkpointing, and recovery.
The real benefit is SeaTunnel’s connector library: more than 100 production-grade destinations. Building and maintaining a writer for every vector store, warehouse, and lakehouse a customer might pick is a losing game for an ingestion project. SeaTunnel already did that work.
OpenCrawling does what it is good at: crawling, Tika extraction, chunking, embedding, and attaching security. It does that once per document. SeaTunnel takes the finished records and writes them everywhere they need to go.
How it works
The pipeline runs through Kafka end to end:
- The crawler pulls documents and publishes them to the ingestion topic.
- Apache Tika extracts text, and the core splits it into chunks.
- oc-embedding-service computes vectors in parallel and publishes enriched records to opencrawling-embedded.
- The SeaTunnel writer consumer streams those records to an intermediate topic.
- The SeaTunnel Zeta cluster reads that topic and writes to every configured sink: ClickHouse, Milvus, Iceberg, and so on.
The connector itself has two parts.
Control plane: SeaTunnelOutputConnector. It generates the SeaTunnel job definition and submits it to the Zeta cluster through REST API v2. A SeaTunnelRestClient built on the Java 25 HttpClient handles job submission, lists running and finished jobs, polls per-job metrics, and checks cluster health through /overview.
Data plane: SeaTunnelStoreWriterConsumer. It subscribes to opencrawling-embedded and streams records conforming to the Open Ingestion Standard (OIS) to the topic that SeaTunnel’s stock Kafka source consumes.
Security and deletes travel with the data
Fan-out is easy until you ask two questions. Does every destination enforce the same permissions? And when a document disappears from SharePoint, does it disappear everywhere?
On permissions, the OIS schema carries the full security model on every chunk, and the SeaTunnel catalog schema passes it to every sink:
| field | type | purpose |
|---|---|---|
| acl | array<string> | Security SIDs, e.g. group:engineers |
| security_allowed_read | array<string> | Explicit allowed readers |
| security_denied_read | array<string> | Explicit denied readers that override allows |
| security_inheritance | boolean | Inheritance flag preserved from the source hierarchy |
A vector store and a lakehouse fed from the same stream get the same ACLs. Security never detaches from content on the way out.
On deletes, re-embedding everything to clean up removed documents is too expensive to consider. When a document is deleted in a source such as Alfresco, SharePoint, or a web repository, OpenCrawling emits a lightweight OIS tombstone with action: DELETE. The connector maps it to a SeaTunnel CDC delete row kind, and ClickHouse, Iceberg, and Milvus purge the record. No embedding calls are made for a delete.
Why Kafka sits in the middle
The obvious design was a custom SeaTunnel connector JAR. We didn’t build one, and the reason is the JVM.
OpenCrawling runs on Java 25 with preview features on, for virtual threads and structured concurrency. SeaTunnel 2.3.13 engines typically run on Java 8, 11, or 17. A plugin compiled for our runtime dropped into /opt/seatunnel/connectors/ is a class version and classpath problem waiting to happen.
Putting Kafka between the two systems removes that coupling. SeaTunnel reads with its own stock Kafka source, so the cluster needs no custom JARs and stays a plain SeaTunnel install. Each side upgrades its runtime on its own schedule.
Try it
The repo includes an end-to-end test, scripts/test-seatunnel-decoupled.sh, backed by a Docker Compose file. It starts a SeaTunnel Zeta cluster, Kafka, Zookeeper, Redis, Ollama with mxbai-embed-large, and the OpenCrawling services. It then crawls test documents, computes 1024-dimension embeddings, deploys the streaming job, verifies it is running, and checks tombstone deletes before tearing everything down.
If you run RAG ingestion into more than one store, the pattern is worth copying even if you never touch OpenCrawling: embed once, attach security once, and let a dedicated integration engine handle the writes.