Open-source repositories gaining traction right now.
316 repositories

sora2 watermark cleaner pro

FireRedASR2S is a SOTA Industrial-Grade All-in-One ASR system with ASR, VAD, LID, and Punc modules. FireRedASR2 supports Chinese (Mandarin, 20+ dialects/accents), English, code-switching, and singing lyrics recognition. FireRedVAD supports speech/singing/music in 100+ langs. FireRedLID supports 100+ langs and 20+ zh dialects.


🦞 Claw Compactor — The 98% Crusher. Cut your AI agent token spend in half with 5 layered compression techniques.


🌋LavaSR: Fast Speech restoration and enhancement

Code for "SceneSmith: Agentic Generation of Simulation-Ready Indoor Scenes"

Skill Compose is an open-source agent builder and runtime platform for skill-powered agents. No workflow graphs. No CLI.

Your own personal AIGC Factory. Any picture. Any reel. The Comfy way. ©️

BitDance: Open-source autoregressive model with binary visual tokens. A research project for building powerful multimodal autoregressive model.

Open-source abilities for OpenHome agents.

Agent World Model: Infinity Synthetic Environments for Agentic Reinforcement Learning


Code for kai0, including training, inference and data collection.

Because `model.fit()` isn't an explanation

OpenClaw skill for cost-optimized model routing based on task complexity


An ARC-AGI solution using Agentica from Symbolica

Standalone Anima Lora trainer with GUI

A real-time and multilingual speech translation model

OpenClaw but it runs entirely on github actions

A RAG pipeline implementation built on the 'Epstein Files 20K' dataset from Hugging Face (Teyler).

Personal PyTorch implementation of "Generative Modeling via Drifting" with Claude

MultiModalWBC is a fully open-source, IsaacLab-based framework for multi-modal whole-body control, designed for motion imitation, motion tracking, and task-conditioned control in legged robots. The framework unifies robot proprioceptive states and multi-modal human motion conditions into a consistent interface

sora2 watermark remover

Official Codebase For paper "Continuous Denoising Enables One-step Language Modeling"

Total conversion for Claude Code. Use RAG and the RPG ruleset apis to play a persistent adventure in any book or world of your choosing.

Real-time speech-to-text caption appliance for a deaf user. Raspberry Pi + 10" touchscreen that transcribes phone calls and room conversation in near real-time.

Official Code Implementation of Translating Flow to Policy via Hindsight Online Imitation

缠中说禅博客,禅师的思维方式模拟


Intelligent audio analysis and automatic genre/mood tagging using Essentia ML models

HTML PPT Designer v5.2 - 智能演示文稿设计器,将任何内容转化为精致的 HTML 演示文稿

AI-powered financial analysis agent for Indian stock markets using AngelOne SmartAPI + Claude

量化投資研究 AI Agent — 透過 CLI 互動介面,自動搜尋財經新聞、分析市場情緒、產生風險評估報告。


An autonomous AI agent that plays Pokemon FireRed in real time using OpenAI's LLM, with a live web dashboard for monitoring.


Streaming Flux editor: live camera→ editing every frames at interactive FPS based on FLUX.2-Klein-4B. Runs on a single H100 at 15+ FPS

Music Recommendation Based on Mood


Deep learning image classifier using TensorFlow and MobileNetV2 to detect AI-generated images


DeepControl: Scaling Search-Augmented LLM Reasoning via Adaptive Information Control

Agent skill for fast, cheap market research using LLM synthetic surveys + Semantic Similarity Rating (SSR). No API keys needed.

A Benchmark and Evaluation Suite for Zero-shot Singing Voice Synthesis



[ICRA 26] C^2ROPE: Causal Continuous Rotary Positional Encoding for 3D Large Multimodal-Models Reasoning

[ICLR 2026] Advancing Universal Deep Learning for Electronic-Structure Hamiltonian Prediction of Materials

Wordpress for Voice AI.

Traffic Surveillance System using YOLOv8 and EasyOCR with document validation

단순한 구조의 숏츠 자동 생성 도구입니다.

Adaptive tokenization for proteins

Example for a Monty-enabled RLM in DSPy


AI-powered healthcare dashboard using Streamlit that integrates drug recommendation and symptom-based disease prediction with multi-model ML comparison and interactive visualizations.

RAIFE is a high-performance Rust pipeline for medical image analysis. Implements rapid NIfTI ingestion, preprocessing, and 2.5D stacking for BraTS-2018. Includes benchmarks demonstrating superior speed over MONAI workflows.

A ComfyUI custom node that turns plain English descriptions into fully structured, cinema-ready prompts for LTX-2 video generation — powered by a local, uncensored LLM with zero internet dependency after setup.

Paste a github.com URL. Submissions are reviewed before they are tracked.