Projects with this topic
-
Java-based ETL template built with Flink, offering a modular, decoupled structure designed to quickly bootstrap scalable, testable, and reliable pipelines. Out of the box provides an easily adaptable PostgreSQL to MongoDB migration.
Updated -
Configuration and data workflows for an instance of Apache Airflow for the DDRplatform
Updated -
End-to-end ETL pipeline for consumer electronics analytics — Medallion Architecture (Bronze → Silver → Gold) with KNIME, Databricks, Delta Lake and Power BI
Updated -
High-throughput asynchronous ETL and stream processing pipeline in Python
Updated -
This project/library contains common elements related to ETL processes...
Updated -
Rustroid Sentinel is a high-performance backend system built to directly integrate with the NASA NeoWs API, continually extracting and analyzing data on potentially hazardous near-Earth objects (NEOs). Built entirely in Rust using the tokio ecosystem, the service delivers a lightning-fast, highly concurrent data pipeline designed for demanding operational environments.
UpdatedUpdated -
MetaBridge is an open-source, ontology-driven data integration and communication platform designed to enable secure, multi-party data sharing through controlled schema evolution, validated ETL pipelines, and strong governance. It provides a modular system for ingesting, transforming, and synchronizing data across heterogeneous systems while maintaining full auditability, spec hygiene, and access control. https://roxanneardary.com/metabridge/
Updated -
Python async video metadata processing pipeline with multi-source ingestion and ETL transforms.
Updated -
Public portfolio demo of a CRM analytics pipeline using Bitrix-style exports, Python ETL, DuckDB, PostgreSQL, and machine learning for deal probability scoring.
UpdatedUpdated -
Mapping-config driven data migration tool for CSV/JSON/XML/TSV with validation and reporting.
Updated -
ETL pipeline project using Python, pandas and SQLite to import and structure CSV, Excel and HTML data files.
Updated -
Automated Dataset Creation & Publishing Pipeline - Scrape, clean, transform, validate and publish datasets to HuggingFace Hub
Updated -
DataRider bloc for ETL Stream with Spark+Scala
Updated -
Dashboard Comercial desenvolvido do zero com modelagem relacional no DBeaver (SQL), pipeline de transformações no Power Query e visualizações interativas no Power BI. Projeto de estudo para portfólio.
Updated -
End-to-end AWS data lake pipeline for fleet telemetry data using S3, Spark, and Athena. Includes partitioned Parquet ETL, vehicle safety analytics, and SQL queries for overspeed and harsh braking detection.
Updated -
Plain text boilerplate removal using character n-gram frequency across a corpus. Builds a template model from a sample, filters files in a single linear pass, and validates automatically. Includes an obfuscated mode where the model is a set of integers and output filenames are hashed: the operator never reads the content. AWK for character processing, Bash for orchestration, Lisp layer planned for positional classification.
Updated -
This is a study project built using FastAPI to practice microservice architecture, data normalization techniques, and clean API design.
The service receives raw payloads from different simulated sources and transforms them into a standardized and validated structure.
It centralizes normalization logic and demonstrates how to build a scalable, maintainable, and test-friendly data processing layer.
Updated -
Solución end-to-end para la migración y análisis de datos utilizando Python, FastAPI, Kafka y PostgreSQL. Implementa un pipeline de datos asíncrono y una API RESTful para analíticas, todo completamente containerizado con Docker Compose para un despliegue fácil y reproducible.
Updated -
Unified project demonstrating both batch analytics and real-time streaming pipelines with Apache Spark:
Batch (PySpark/Jupyter): Processed S&P 500 stock data, applied transformations, and ran distributed computations.
Streaming (Spark + Kafka): Built a streaming pipeline to consume Kafka topics, process messages in real-time, and visualize outputs.
Deployed using Docker and Jupyter for reproducibility.
Updated