Ultralytics YOLO26, YOLO11, YOLOv8 — object detection, instance segmentation, semantic segmentation, image classification, pose estimation, object tracking
Image Classification GitHub Repositories
Explore popular GitHub repositories tagged “image-classification”.
Compare stars, forks, and programming language using the same GitStar view as GitHub Trending.
Trending Repositories
Ultralytics YOLOv5 in PyTorch for object detection, instance segmentation, classification, training, and export.
The largest collection of PyTorch image encoders / backbones. Including train, eval, inference, export scripts, and pretrained weights -- ResNet, ResNeXT, EfficientNet, NFNet, Vision Transformer (ViT), MobileNetV4, MobileNet-V3 & V2, RegNet, DPN, CSPNet, Swin Transformer, MaxViT, CoAtNet, ConvNeXt, and more
Label Studio is a multi-type data labeling and annotation tool with standardized output format
Implementation of Vision Transformer, a simple way to achieve SOTA in vision classification with only a single transformer encoder, in Pytorch
Computer Vision Annotation Tool (CVAT) is a leading platform for building high-quality visual datasets for vision AI. It offers open-source, cloud, and enterprise products, as well as labeling services, for image, video, and 3D annotation with AI-assisted labeling, quality assurance, team collaboration, analytics, and developer APIs.
This is an official implementation for "Swin Transformer: Hierarchical Vision Transformer using Shifted Windows".
Advanced AI Explainability for computer vision. Support for CNNs, Vision Transformers, Classification, Object detection, Segmentation, Image similarity and more.
PyTorch tutorials and fun projects including neural talk, neural style, poem writing, anime generation (《深度学习框架PyTorch:入门与实战》)
cvpr2024/cvpr2023/cvpr2022/cvpr2021/cvpr2020/cvpr2019/cvpr2018/cvpr2017 论文/代码/解读/直播合集,极市团队整理
Refine high-quality datasets and visual AI models
Techniques for deep learning with satellite & aerial imagery
[CVPR 2024 Oral] InternVL Family: A Pioneering Open-Source Alternative to GPT-4o. 接近GPT-4o表现的开源多模态对话模型
X-AnyLabeling: A lightweight, efficient, and unified cross-platform desktop application for annotating text, image, video, and multimodal data, combining versatile built-in tools with state-of-the-art AI models and flexible multi-format export.
Best Practices, code samples, and documentation for Computer Vision.
A collection of tutorials on state-of-the-art computer vision models and techniques. Explore everything from foundational architectures like ResNet to cutting-edge models like RF-DETR, YOLO11, SAM 3, and Qwen3-VL.
Curated list of Machine Learning, NLP, Vision, Recommender Systems Project Ideas
Experience, Learn and Code the latest breakthrough innovations with Microsoft AI
Gluon CV Toolkit
A treasure chest for visual classification and recognition powered by PaddlePaddle
An absolute beginner's guide to Machine Learning and Image Classification with Neural Networks
Practice on cifar100(ResNet, DenseNet, VGG, GoogleNet, InceptionV3, InceptionV4, Inception-ResNetv2, Xception, Resnet In Resnet, ResNext,ShuffleNet, ShuffleNetv2, MobileNet, MobileNetv2, SqueezeNet, NasNet, Residual Attention Network, SENet, WideResNet)
Differentiable architecture search for convolutional and recurrent networks
OpenMMLab Pre-training Toolbox and Benchmark
A library for transfer learning by reusing parts of TensorFlow models.
Accelerated deep learning R&D
A curated list of deep learning image classification papers and codes
Sandbox for training deep learning networks
AI Roadmap:机器学习(Machine Learning)、深度学习(Deep Learning)、对抗神经网络(GAN),图神经网络(GNN),NLP,大数据相关的发展路书(roadmap), 并附海量源码(python,pytorch)带大家消化基本知识点,突破面试,完成从新手到合格工程师的跨越,其中深度学习相关论文附有tensorflow caffe官方源码,应用部分含推荐算法和知识图谱
A cross-platform video structuring (video analysis) framework based on CV models & mLLM.