The Postgres development platform. Supabase gives you a dedicated Postgres database to build your web, mobile, and AI applications.
Embeddings GitHub Repositories
Explore popular GitHub repositories tagged “embeddings”.
Compare stars, forks, and programming language using the same GitStar view as GitHub Trending.
Trending Repositories
Persistent Context Across Sessions for Every Agent – Captures everything your agent does during sessions, compresses it with AI, and injects relevant context back into future sessions. Works with Claude Code, OpenClaw, Codex, Gemini, Hermes, Copilot, OpenCode + More
This repository showcases various advanced techniques for Retrieval-Augmented Generation (RAG) systems. Each technique has a detailed notebook tutorial.
Open-source LLM knowledge platform: turn raw documents into a queryable RAG, an autonomous reasoning agent, and a self-maintaining Wiki.
A vector index built on TurboQuant, written in Rust with Python bindings
💡 All-in-one AI framework for semantic search, LLM orchestration and language model workflows
LangChain4j is an idiomatic, open-source Java library for building LLM-powered applications on the JVM. It offers a unified API over popular LLM providers and vector stores, and makes implementing tool calling (including MCP support), agents and RAG easy. It integrates seamlessly with enterprise Java frameworks like Quarkus and Spring Boot.
The all-in-one, open-source backend platform for agentic coding. InsForge gives your coding agent database, auth, storage, compute, hosting, and AI gateway to ship full-stack apps end-to-end.
100+ Chinese Word Vectors 上百种预训练中文词向量
Retrieval and Retrieval-augmented LLMs
SeaTunnel is a multimodal, high-performance, distributed, massive data integration tool.
Open Lakehouse Format for Multimodal AI. Convert from Parquet in 2 lines of code for 100x faster random access, vector index, and data versioning. Compatible with Pandas, DuckDB, Polars, Pyarrow, and PyTorch with more integrations coming..
Postgres with GPUs for ML/AI apps.
Memory library for building stateful agents
The easiest way to use deep metric learning in your application. Modular, flexible, and extensible. Written in PyTorch.
Fast and Accurate Code Search for Agents. Uses 99% fewer tokens than grep+read
High-performance data engine for AI and multimodal workloads. Process images, audio, video, and structured data at any scale
Find related notes and excerpts while writing. Your link building copilot displays relevant content in graph + list view. A local embedding model powers semantic search. Zero setup. No API key.
AutoRAG: Now your agent can find anything in your computer. It gets smarter if you are using it frequently.
A blazing fast inference solution for text embeddings models
text2vec, text to vector. 文本向量表征工具,把文本转化为向量矩阵,实现了Word2Vec、RankBM25、Sentence-BERT、CoSENT等文本表征、文本相似度计算模型,开箱即用。
Local persistent memory store for LLM applications including claude desktop, github copilot, codex, antigravity, etc.
One delightful Ruby framework for every major AI provider. Build AI agents, chatbots, RAG apps, and multimodal workflows in beautiful, expressive code.
A python library for self-supervised learning on images.
A library for transfer learning by reusing parts of TensorFlow models.
A curated list of Generative AI tools, works, models, and references
Towhee is a framework that is dedicated to making neural data processing pipelines simple and fast.
MTEB: State-of-the-art evaluation of embeddings across languages and modalities
Enterprise-grade (40m+ LOC) codebase intelligence, zero-setup, local & private Plugin/Skill/Extension or MCP: hybrid semantic search, polyglot dependency graphs, symbol-level impact analysis & call-flow, interactive HTML viewer, cross-project & branch-aware search, DB/API/infra knowledge. 61% less tokens, 84% fewer calls, 37x faster. Cloud in beta.
Fast, Accurate, Lightweight Python library to make State of the Art Embedding