Indexion Segment

by trkbt107ad5ad35a668No license2 starsListed Oct 8, 2026Updated Oct 8, 2026Repository updated 2 weeks ago

Split text into contextual chunks for RAG/embedding pipelines. Document segmentation and section extraction using window, tfidf, punctuation, or hybrid strategies chosen by intent.

Instructions onlyAI & Agents
AI-generated overview

Splits text into contextual chunks for RAG and embedding pipelines using window, TF-IDF, punctuation, or hybrid strategies.

What it does
The skill segments a text document into smaller contextual chunks intended for retrieval-augmented generation or embedding pipelines. It offers window divergence, TF-IDF, punctuation, and hybrid NCD+TF-IDF strategies, with tunable segment sizes, thresholds, and window sizes. Output is written to a chosen directory with a configurable filename prefix.
When to use it
Use it when text must be chunked for RAG or embedding pipelines, or when a document needs to be split into meaningful sections. It also suits preparing text for sub-document similarity analysis.
Requirements
Requires the indexion command-line tool with its segment subcommand; no scripts ship with the skill. Input is a text file and output is a directory.

indexion segment

Split text into contextual segments using divergence-based, TF-IDF, or punctuation strategies.

When to Use

  • User needs to chunk text for RAG or embedding pipelines
  • User wants to split a document into meaningful sections
  • User asks to segment text for processing
  • Preparing text for similarity analysis at sub-document level

Usage

bash
# Default window divergence strategyindexion segment <input-file> <output-dir>
# TF-IDF based segmentationindexion segment --strategy=tfidf <input-file> <output-dir>
# Punctuation-based segmentationindexion segment --strategy=punctuation <input-file> <output-dir>
# Custom segment sizesindexion segment --min-size=200 --max-size=3000 --target-size=800 document.txt output/
# Custom divergence thresholdindexion segment --threshold=0.5 document.txt output/
# Adaptive threshold mode (default)indexion segment --adaptive document.txt output/
# Hybrid NCD+TF-IDF modeindexion segment --hybrid --ncd-weight=0.6 --tfidf-weight=0.4 document.txt output/
# Custom window sizeindexion segment --window-size=5 document.txt output/
# Custom output prefixindexion segment --prefix=chunk document.txt output/

Options

OptionDefaultDescription
--strategy=NAMEwindowStrategy: window, tfidf, punctuation
--min-size=INT100Minimum segment characters
--max-size=INT2000Maximum segment characters
--target-size=INT500Target segment characters
--threshold=FLOAT0.42Divergence threshold
--window-size=INT3Window size
--adaptivetrueAdaptive threshold mode
--hybridfalseNCD+TF-IDF hybrid mode
--ncd-weight=FLOAT0.5NCD weight in hybrid mode
--tfidf-weight=FLOAT0.5TF-IDF weight in hybrid mode
--prefix=NAMEsegmentOutput file prefix

Strategies

StrategyDescription
window (default)Sliding window divergence detection
tfidfTF-IDF based topic change detection
punctuationPunctuation/sentence boundary based

Workflow

  1. Run indexion segment <input-file> <output-dir> to split text with defaults
  2. Adjust --threshold and --target-size to tune segmentation granularity
  3. Use --hybrid mode for better accuracy on mixed-content documents

Source and attribution

Source:trkbt10/indexion-skillsinskills/indexion-segmentat commit7ad5ad3

License: No license

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal