Open-source repositories gaining traction right now.
80 repositories

Original Apollo 11 Guidance Computer (AGC) source code for the command and lunar modules.

Turso is an in-process SQL database, compatible with SQLite.

1,324-exercise fitness dataset — animation GIFs, 180×180 thumbnails, muscle-group & equipment data, and step-by-step instructions in 6 languages. The exercise data layer behind the LogPress app.

小红书笔记 | 评论爬虫、抖音视频 | 评论爬虫、快手视频 | 评论爬虫、B 站视频 | 评论爬虫、微博帖子 | 评论爬虫、百度贴吧帖子 | 百度贴吧评论回复爬虫 | 知乎问答文章|评论爬虫

🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!

OpenMetadata is a unified metadata platform for data discovery, data observability, and data governance powered by a central metadata repository, in-depth column level lineage, and seamless team collaboration.

A curated list of awesome libraries, packages, strategies, books, blogs, tutorials for systematic trading.

Financial data platform for analysts, quants and AI agents.

FinceptTerminal is a modern finance application offering advanced market analytics, investment research, and economic data tools, designed for interactive exploration and data-driven decision-making in a user-friendly environment.

PostgreSQL in-database durable execution

Data Engineering Zoomcamp is a free 9-week course on building production-ready data pipelines. The next cohort starts in January 2026. Join the course here 👇🏼

Real-time global intelligence dashboard. AI-powered news aggregation, geopolitical monitoring, and infrastructure tracking in a unified situational awareness interface

GenBI (Generative BI) for AI agents, an open-source, governed text-to-SQL through an open context layer that turns natural-language questions into trusted dashboards, charts, and SQL across 20+ data sources, such as BigQuery, Snowflake, PostgreSQL, ClickHouse, Amazon Redshift, Databricks and more.

Free and Open Source, Distributed, RESTful Search Engine

Code for Machine Learning for Algorithmic Trading, 2nd edition.

Sponsor Star scikit-learn / scikit-learn scikit-learn: machine learning in Python

Prefect is a workflow orchestration framework for building resilient data pipelines in Python.

Dolt – Git for Data

Star influxdata / influxdb Scalable datastore for metrics, events, and real-time analytics

M3U Playlist for free TV channels

面向 A 股日K、分钟K与ETF分钟数据的本地量化引擎,集成增量同步、本地缓存、复权、批量查询、回测与指标计算。

Apache Superset is a Data Visualization and Data Exploration Platform

Apache Ossie, industry wide specification effort to standardize how we exchange semantic metadata across analytics, AI and BI platforms, providing a vendor neutral, single source of truth for semantic data

A self-hosted data logger for your Tesla 🚘 [main maintainer=@JakobLichterfeld]

The open and composable observability and data visualization platform. Visualize metrics, logs, and traces from multiple sources like Prometheus, Loki, Elasticsearch, InfluxDB, Postgres and many more.

Star dbeaver / dbeaver Free universal database tool and SQL client

Open source transactional distributed database. Linear scalability and proven fault-tolerance on commodity hardware or cloud infrastructure without compromising performance.

Sponsor Star paradedb / paradedb Simple, Elastic-quality search for Postgres

Star trinodb / trino Official repository of Trino, the distributed SQL query engine for big data, formerly known as PrestoSQL (https://trino.io)

Star apache / datafusion Apache DataFusion SQL Query Engine

Star debezium / debezium Change data capture for a variety of databases. Please log issues at https://github.com/debezium/dbz/issues.

Star apache / nifi Apache NiFi

Star prestodb / presto The official home of the Presto distributed SQL query engine for big data

Star apache / kafka Mirror of Apache Kafka

Star surrealdb / surrealdb A scalable, distributed, collaborative, document-graph database, for the realtime web

Star ranaroussi / yfinance Download market data from Yahoo! Finance's API

Star apache / druid Apache Druid: a high performance real-time analytics database.

Star apache / beam Apache Beam is a unified programming model for Batch and Streaming data processing.

Sponsor Star gosom / google-maps-scraper scrape data data from Google Maps. Extracts data such as the name, address, phone number, website URL, rating, reviews number, latitude and longitude, reviews,email and more for each place

Self-learning data agent that grounds its answers in 6 layers of context. Inspired by OpenAI's in-house implementation.

Archive a lifetime of email and chat. Offline search, analytics, and AI query over your full message history. Powered by DuckDB

Turn any collection of documents into a knowledge graph. Extract entities and relationships via LLM, deduplicate with your approval, and explore the result in your browser — all from the CLI.

A fast and soft pattern search for trillion-scale corpora.

A "Clawdbot" in every row with 400 lines of Postgres SQL




LLM-Driven Business Intelligence Engine

Monitoring the DOJ Epstein Files — 931,000 PDFs, 0 arrests

Full-stack LinkedIn OSINT toolkit. Four-phase funnel: discover companies by region, batch scrape employees, classify roles by hierarchy/department, and deep dive into profiles. Interactive D3.js org chart viewer, Groq AI enhancement, anti-detection stealth, proxy support, and graceful partial-save on interruption.

Self-hosted view.


MCP server for Brazilian agricultural data — connect LLMs to 19 public data sources via agrobr

SQLite-like embedded vector database

Neuroscope is a EEG data platform to help EEG researchers. Supports 11 different file formats including EDF/CSV/JSON and much more. There are TONS of adjustable parameters throughout the website, and ALL processing is ran locally, to get EASY visualizations for your research.


A Python library for storing and querying n-ary relationships with provenance tracking. SQLite-backed, zero configuration.

Geospatial Risk Engine

Euroleague basketball analytics platform with ML pipeline, interactive Streamlit dashboard, and live API integration

Orca is a production-ready, serverless and agentic-ready template for building a data warehouse.
Paste a github.com URL. Submissions are reviewed before they are tracked.