Back to News
RSS feedgithub.com

rsync.ai: A Self-Hosted AI Data Platform for Pipelines and Lineage

Summary

rsync.ai is a source-available, self-hosted data platform for batch pipelines, change-data capture (CDC), scheduled SQL models, data exploration, and table-level lineage. Users describe a job in plain English, after which an agent resolves the request into an explicit staged plan and pauses for approval when the source, schema, tables, keys, or sync mode is ambiguous. Nothing moves until the user answers those questions. Approved stages run as Temporal workflows, allowing long-running jobs to resume after restarts, redeployments, or worker failures; domain events expose stage state, row counts, and trace IDs in the UI. The platform ships 21 versioned containerized connectors, including relational databases, warehouses, object stores, APIs, and five CDC-capable databases: PostgreSQL, MySQL, SQL Server, Oracle, and MongoDB. It also includes a natural-language and SQL Data Explorer, saved queries, scheduled dependency-aware models, exports, and recent table-level lineage. The installer can use an OpenAI key, a bundled Ollama deployment, or no LLM; core pipelines, raw SQL, and connectors still work without a model, while model-dependent features remain unavailable until one is configured. Deployment is provided through Docker Compose or Helm/Kubernetes, with credentials encrypted using a user-held key. The current release described is v0.1.7. The project is early: there is no managed offering, the connector catalog is smaller than major managed ELT services, data-quality assertions are not yet built, and end-to-end managed-cluster Kubernetes deployments have not been verified. It is licensed under the source-available Elastic License 2.0, which permits internal use and modification but prohibits offering the software as a hosted service.