NCRF++, a Neural Sequence Labeling Toolkit. Easy use to any sequence labeling tasks (e.g. NER, POS, Segmentation). It includes character LSTM/CNN, word LSTM/CNN and softmax/CRF components.
Chunking GitHub Repositories
Explore popular GitHub repositories tagged “chunking”.
Compare stars, forks, and programming language using the same GitStar view as GitHub Trending.
Trending Repositories
Content-Addressable Data Synchronization Tool
A fast, lightweight and easy-to-use Python library for splitting text into semantically meaningful chunks.
Fully neural approach for text chunking
An extensible Java framework for building event-driven applications that break up XML and non-XML data into chunks for data integration
Alternative casync implementation
Adaptive Chunking: automatically select the best chunking method per document for RAG. Accepted at LREC 2026.
smart-llm-loader is a lightweight yet powerful Python package that transforms any document into LLM-ready chunks. Spend less time on preprocessing headaches and more time building what matters. From RAG systems to chatbots to document Q&A, SmartLLMLoader handles the heavy lifting so you can focus on creating exceptional AI applications.
The RAG Experiment Accelerator is a versatile tool designed to expedite and facilitate the process of conducting experiments and evaluations using Azure Cognitive Search and RAG pattern.
A package for parsing PDFs and analyzing their content using LLMs.
A new chunking strategy developed by ZeroEntropy for general semantic chunking using Llama-70B.
A TensorFlow implementation of Neural Sequence Labeling model, which is able to tackle sequence labeling tasks such as POS Tagging, Chunking, NER, Punctuation Restoration and etc.
Pure Rust PDF library for AI/RAG: structure-aware chunking, no ML, no C deps.
Open-source toolkit for reliable RAG pipelines: convert PDFs to Markdown, clean documents, inspect chunks, compare chunking strategies, and enrich metadata for LLM applications.
PDFStract - Extract, Chunking and Embedding Layer in Your RAG Pipeline - Available as CLI - WEBUI - API
🍱 Semantically create chunks from large document for passing to LLM workflows
A Python CLI to test, benchmark, and find the best RAG chunking strategy for your Markdown documents.
Live TS segmenter and HLS manifest creation in Go
An LLM GUI application; enables you to interact with your files, offering dynamic parameters that can modify response behavior during runtime.
An Overview of the Latest Document Chunking Research
An asynchronous event-driven HTTP client based on netty.
webpack 2, react hotloader 3, react router v4, code splitting and more
One library to split them all: Sentence, Code, Docs. Chunk smarter, not harder — built for LLMs, RAG pipelines, and beyond.
📑 Split Laravel jobs into multiple separate job chunks
Грамматический Словарь Русского Языка (+ английский, японский, etc)
Fast multi-threaded content-dependent chunking deduplication for Buffers in C++ with a reference implementation in Javascript. Ships with extensive tests, a fuzz test and a benchmark.
Incremental asset delivery library
FastCDC implementation in Python https://pypi.org/project/fastcdc/
Labelling Sequential Data in Natural Language Processing with R - using CRFsuite
Extract and align grammar patterns from English sentences.