This repository contains demos I made with the Transformers library by HuggingFace.
- Stars
- 11.7K
- Forks
- 1.7K
- Pushed
- 3mo ago
Current GitHub adoption and maintenance signals.
This repository contains demos I made with the Transformers library by HuggingFace.
This repository contains an overview of important follow-up works based on the original Vision Transformer (ViT) by Google.
A repository containing general tutorials I'd like to share with the world.
🤗Transformers: State-of-the-art Natural Language Processing for Pytorch and TensorFlow 2.0.
Repository containing awesome resources regarding Hugging Face tooling.
Transforming textual descriptions into process models using deep learning
A tiny package supporting distributed computation of COCO metrics for PyTorch models.
Short README about myself.
UniLM - Unified Language Model Pre-training / Pre-training for NLP and Beyond
No repository description provided.
A repository showcasing the entire workflow of putting computer vision models in production.
Some notes I took when learning about diffusion models.
a state-of-the-art-level open visual language model
Notebooks using the Hugging Face libraries 🤗
RF-DETR is a real-time object detection model architecture developed by Roboflow, released under the Apache 2.0 license.
Reference implementation of Mistral AI 7B v0.1 model.
Official codebase used to develop Vision Transformer, MLP-Mixer, LiT and more.
The official repository for MedSAM: Segment Anything in Medical Images.
This repository is meant for parsing evaluation results from Hugging Face models, and opening pull requests on the hub to display them at leaderboards.
No repository description provided.
Official repository for our work on micro-budget training of large-scale diffusion models.
Autoregressive Model Beats Diffusion: 🦙 Llama for Scalable Image Generation
No repository description provided.
This repository contains the assignments I made during the 2019 version of the Deep Learning for Computer Vision course taught at the University of Michigan.
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing?
Utilities to use the Hugging Face Hub API
🤗 ml-intern: an open-source ML engineer that reads papers, trains models, and ships ML models
A repository with various baselines for the agentic-document-ai project.
Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞
No repository description provided.
VisualQuality-R1 is the first open-sourced NR-IQA model can accurately describe and rate the image quality.
No repository description provided.
veRL: Volcano Engine Reinforcement Learning for LLM
DarkIR: Robust Low-Light Image Restoration [Official PyTorch Implementation]
Tips for releasing research code in Machine Learning (with official NeurIPS 2020 recommendations)
Simple, safe way to store and distribute tensors
OpenMMLab Computer Vision Foundation
Tools for extracting tables and results from Machine Learning papers
My solutions to the practical assignments of CS224n (Natural Language Processing with Deep Learning) [Stanford University-Winter 2019]
[DEIMv2] Real Time Object Detection Meets DINOv3
A demo on how to set up a simple agentic RAG backend and frontend.
EfficientSAM3 compresses SAM3 into lightweight, edge-friendly models via progressive knowledge distillation for fast promptable concept segmentation and tracking.
A benchmark for LLMs on complicated tasks in the terminal
No repository description provided.
A scalable generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and Speech AI (Automatic Speech Recognition and Text-to-Speech)
A lightweight, powerful framework for multi-agent workflows
ICLR2024 Spotlight: curation/training code, metadata, distribution and pre-trained models for MetaCLIP; CVPR 2024: MoDE: CLIP Data Experts via Clustering
Official code for the paper: Depth Anything At Any Condition
[CVPR 2025 Best Paper Award Candidate] VGGT: Visual Geometry Grounded Transformer
An MCP server to search for flights.
Code to repro PE and PLM
A blazing fast inference solution for text embeddings models
Implementation for SimDINO/SimDINOv2
Official code of DeepMesh: Auto-Regressive Artist-mesh Creation with Reinforcement Learning
OmniMamba: Efficient and Unified Multimodal Understanding and Generation via State Space Models
A jump start solution using GKE or Cloud Run with Cloud SQL and VertexAI
YOLOE: Real-Time Seeing Anything
:alarm_clock: AI conference deadline countdowns
YOLOv12: Attention-Centric Real-Time Object Detectors
This repository is an official implementation of the paper "LW-DETR: A Transformer Replacement to YOLO for Real-Time Detection".
User-friendly AI Interface (Supports Ollama, OpenAI API, ...)
This is a simple demonstration of more advanced, agentic patterns built on top of the Realtime API.
[ECCV 2022] ByteTrack: Multi-Object Tracking by Associating Every Detection Box
A Conversational Speech Generation Model
This repo contains the code for our paper An Image is Worth 32 Tokens for Reconstruction and Generation
Fast, efficient, battle-tested at Alibaba's scale. Hybrid architecture code review tool: deterministic pipelines + LLM Agent, precise line-level comments, built-in multi-language ruleset (NPE, thread-safety, XSS, SQL injection), OpenAI & Anthropic compatible.
OpenAI's Codex Security CLI and TypeScript SDK for finding, validating, and fixing security vulnerabilities. npm: https://www.npmjs.com/package/@openai/codex-security
A lightweight, local-first, and free experiment tracking library from Hugging Face 🤗
No repository description provided.
The repository provides code for running inference and finetuning with the Meta Segment Anything Model 3 (SAM 3), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.
Toolkit for linearizing PDFs for LLM datasets/training
AI-assisted academic posters.
⏰ AI conference deadline countdowns
Code & data for TaxCalcBench
[NeurIPS 2024] OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
A collection of tutorials on state-of-the-art computer vision models and techniques. Explore everything from foundational architectures like ResNet to cutting-edge models like RF-DETR, YOLO11, SAM 3, and Qwen3-VL.
Official code and models for Video Encoder-only Mask Transformer (VidEoMT).
SGLang is a fast serving framework for large language models and vision language models.
CVE cache of the official CVE List in CVE JSON 5 format
[CVPR 2025 Highlight] Official code and models for Encoder-only Mask Transformer (EoMT).
CUDA accelerated rasterization of gaussian splatting
Get started with building Fullstack Agents using Gemini 2.5 and LangGraph
[CVPR 2025 Best Paper Nomination] FoundationStereo: Zero-Shot Stereo Matching
[CVPR 2025] "DiffusionSfM: Predicting Structure and Motion via Ray Origin and Endpoint Diffusion" official implementation.
DSPy: The framework for programming—not prompting—language models
Official repository for "AnyCam: Learning to Recover Camera Poses and Intrinsics from Casual Videos" (CVPR 2025)
The simplest, fastest repository for training/finetuning small-sized VLMs.
Official Code for "LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models"
A TTS model capable of generating ultra-realistic dialogue in one pass.
Code for "Scaling Language-Free Visual Representation Learning" paper (Web-SSL).
Code for BLT research paper
[AAAI'25] DLF: Disentangled-Language-Focused Multimodal Sentiment Analysis
Official PyTorch Implementation of Opt-CWM: Self-Supervised Learning of Motion Concepts by Optimizing Counterfactuals.
The official implementation for the CVPR'2025 paper Dynamic Updates for Language Adaptation in Visual-Language Tracking
No repository description provided.
[CVPR'25] Official repository of Sonata: Self-Supervised Learning of Reliable Point Representations
[CVPR2025] HVI: A New Color Space for Low-light Image Enhancement && "You Only Need One Color Space: An Efficient Network for Low-light Image Enhancement"
PyTorch implementation of FractalGen https://arxiv.org/abs/2502.17437
Unified automatic quality assessment for speech, music, and sound.
Make websites accessible for AI agents