
glin
io.github.AkashChatterjeev0.1.4更新于 Oct 11, 2026
Exactly explainable classifiers for agents: train, tune, predict, explain on CSV or Excel.
概览
让助手基于 CSV 或 Excel 数据训练、调参、预测并解释表格分类模型。
- 功能
- 提供列出与查看已训练模型的工具,可从 CSV 文件或原始 CSV 文本训练新分类器,并预测类别及概率,同时给出完整的逐项得分审计(R50、R51、R54、R85)。训练支持分组留出集、自定义缺失代码和超参数,并针对未见输入、被丢弃的列和可能的标签泄漏给出警告(R68、R97、R98、R268)。还提供提示词,引导代理清洗数据、在防过拟合约束下调参,并在训练前准备 CSV(R102、R120、R123)。
- 适用场景
- 当你希望助手在表格数据上构建或应用可解释的分类模型(例如给商机打分)并说明每个预测的原因时使用(R133、R136)。适合能直接读取数据文件的本地桌面客户端(R92)。不适用于回归目标、自由文本或嵌套 JSON 列(R249、R253、R312)。
- 运行要求
- 需要本地 Python 环境,通过 uvx 运行或使用 uv、pipx、pip 安装 PyPI 上的 glin-ml 包(R1、R4)。MCP 客户端需配置为运行 glin 命令并带 serve 参数(R40)。读取 Excel 需安装可选的 excel 附加组件(R5、R6)。依赖 mcp Python SDK 1.21.1 及以上、低于 3(R28、R29)。未声明账号、API 密钥或环境变量。
安装
在 SourceWeft 中
- 打开 控制台中的 glin,将其添加到工作区。
- 为需要使用其工具的对话启用该服务。
Desktop only,通过 STDIO。 STDIO 服务会启动本地进程,因此需要 SourceWeft 桌面宿主。
其他 MCP 客户端
参照 仓库 中的启动说明。
README
glin
☕ If glin is useful to you and you want to help maintain it, you can buy me a coffee.
glin is a CLI tool and Python library. It gives AI agents a fast, statistical "gut feeling" (System 1 thinking).
glin trains an Explainable Boosting Machine (EBM) on a CSV. It then gives the model to LLM agents through MCP, locally over stdio. Each prediction comes with an exact breakdown of the features that drove it. The breakdown comes from the additive structure of the model. It is not a SHAP or LIME approximation.
glin runs locally only (stdio). It has no remote or HTTP mode. Version 0.1.3 is the last release that had one.
Changelog: CHANGELOG.md lists the changes in each release.
Case study: Child stunting across Indian districts (NFHS-5). A glin model flags districts with high child stunting. Claude calls the model through MCP.
Contents
- Install
- Train a model
- Use glin with an MCP client
- Tools exposed over MCP
- Prompts exposed over MCP
- Restricting
csv_path - Let an agent tune the model
- Example: ask Claude directly
- Case study
- Data requirements
- Releasing (maintainers)
Install
The PyPI package is named glin-ml. The name glin belongs to an unrelated project. The CLI command and the Python import are both glin.
Choose one command:
To read Excel files (.xlsx and .xls), install the optional extra: uv tool install 'glin-ml[excel]' or pip install 'glin-ml[excel]'. If you run the server with uvx, use uvx --from 'glin-ml[excel]' glin serve. See Excel input.
Then check the install:
Install as a Claude Code plugin
The plugin adds the glin MCP server and a skill. The skill tells the agent when and how to build, tune, predict with and explain a model. The plugin starts the server with uvx, so you must have uv installed. You do not need to install glin first. The plugin runs uvx --from 'glin-ml[excel]>=0.1.4' glin serve. This command includes Excel support, and it does not start a version older than 0.1.4.
In Claude Code, run:
The files are in .claude-plugin/marketplace.json and plugins/glin/.
MCP Registry
After each release, the maintainer publishes glin to the official MCP Registry as io.github.AkashChatterjee/glin. This is a manual step (see Publishing to the MCP Registry). The metadata is in server.json. The registry points to the glin-ml package on PyPI. The package also installs a glin-ml command, which is the same as glin, so uvx glin-ml serve works.
The registry entry installs glin-ml without the excel extra, because the registry format has no field for extras. To read Excel files, set your MCP client to run uvx --from 'glin-ml[excel]' glin serve (command uvx, arguments --from, glin-ml[excel], glin, serve).
Install for development
Install an editable copy from a clone of this repository:
glin works with the mcp 1.x and 2.x Python SDK lines (mcp>=1.21.1,<3). Versions before 1.21.1 do not work. mcp 1.21.0 and earlier fail on import with pydantic 2.14 and later. The earliest 1.x versions also do not have the API that glin uses. To test one line, add a constraint. For example: pip install -e ".[dev]" "mcp>=2,<3".
The Tests GitHub Actions workflow (.github/workflows/test.yml) runs the test suite on each pull request and on each push to main. It runs on Python 3.10 and 3.12 with the latest mcp 1.x and the latest mcp 2.x. It also runs on Python 3.10 with the lowest supported mcp (mcp==1.21.1).
Train a model
glin saves each model in ~/.glin/models/<name>/. During training, glin prints a live line that shows the boosting progress. To stop this line, add --quiet. The stage messages still print.
Model names
When you make a model (glin train --name, or model_name in train_model), the name must obey this rule:
- Use 1 to 64 characters.
- Use only letters, digits,
-,_and.. - Do not start the name with
.. Thus.,..and hidden names are not valid. - Do not use
/or\. Thus a path or an absolute path is not valid.
When you use or delete a model (inspect_model, predict, explain_model, glin delete), the name must be one folder name in the models folder. It must not contain / or \, and it must not be empty, . or ... Thus you can still use models that an earlier glin version saved with other names, for example a name with a space.
glin rejects a name that does not obey the applicable rule. The error message gives the rule. glin also makes sure that the model folder is directly in the models folder, also when a symbolic link is in the path. Thus a model name cannot write, read or delete files outside the models folder.
List and delete models
Excel input (.xlsx / .xls)
Spreadsheets from governments and analysts train directly. This includes banner rows and merged headers. Excel support is an optional extra, so the base install stays small. openpyxl reads .xlsx files. xlrd reads old .xls files.
- Sheet:
--sheettakes a name or a 1-based position. A workbook with one sheet needs no flag. A workbook with several sheets and no--sheetgives an error that lists the sheets. glin ignores case and spaces at the start and end of a name only if one sheet matches. If two or more sheets match (for exampleDataanddata), glin gives an error that lists them. Then use the exact name or the position. - Header detection: Without
--header-rows, glin finds the first row that is mostly filled and looks like labels. It skips title rows above it. If the choice is not clear, glin stops with an error. The error names the rows and asks for--header-rows. glin does not guess. Two cases are not clear. In the first, a sparse group-label row sits next to the header row. In the second, the first full row looks like data. - Footnotes: glin drops footnote rows at the end of the sheet. A footnote row has text in half of the columns or fewer, and no numbers. Its first filled cell starts with
Source,Note,Footnote,*,†,‡or©, or it is a sentence of six words or more. glin does the same for all column counts. A sparse last row that does not look like a note stays as data. glin gives a warning with the sheet row numbers of the rows that it dropped. --header-rows: Use 1-based row numbers as shown in Excel:4,3-4or3,4. glin joins a multi-row header into one name for each column with-, for example2023 - sales. A blank cell in an upper header row continues the label from the cell on its left (merged group headers). glin drops blanks and repeated parts. Duplicate names get the suffixes_2and_3. A header cell that holds a date becomes an ISO date, for example2024-01-01. Data starts below the last header row. glin prints the sheet, the header rows and the number of rows it read.- Formulas: glin uses the cached results of the cells. A spreadsheet app must calculate and save the formulas first.
- Empty rows and columns: glin drops fully empty rows and columns.
- CSV files: CSV behavior does not change.
--sheetand--header-rowsare errors for CSV files.
Use glin with an MCP client
Add glin to the config of your MCP client, for example Claude Desktop or Cursor. In Claude Desktop, the file is claude_desktop_config.json:
The server gives the glin version in serverInfo.version. glin serve writes only MCP messages to stdout. Log messages go to stderr. During serve, glin sets the log level of the interpret library to WARNING, so each training does not fill the client log.
Tools exposed over MCP
Every tool has MCP annotations (title, readOnlyHint, destructiveHint). All tools are read-only except train_model. train_model is marked destructive because overwrite=True replaces a model. A client can approve the read-only tools automatically and ask before it runs train_model.
list_models()lists all trained models.inspect_model(model_name)returns the feature schema and the target classes. It lists date columns underfeatures.date. Each date column shows its derived terms and thestart_datethatdays_since_startcounts from.predict(model_name, features, top_n=10)returns the predicted class and the probabilities, plus a full audit:base_logitis the intercept.contributionslists the contribution of every term, sorted by size.additivity_verifiedchecks thatreconstructed_probabilityequalsmodel_probability. Both are keyed by class name.- For binary targets,
score_classis the second entry oftarget_classes. glin sorts the classes by value, so this is not always the alphabetically second class. The scores andbase_logitare log-odds for this class. A positive score pushes toward it. A negative score pushes toward the other class. predicted_class,score_classand the entries oftarget_classeshave the type of the target values (for example1, ortruein JSON). The keys ofprobabilities,model_probabilityandreconstructed_probabilityare always text (for example"1"or"True"). To find the probability ofscore_classin Python, usemodel_probability[str(score_class)].- For multiclass targets,
base_logitis per class, and each term hasscore_by_class. warningslists inputs that the model never saw. Each warning hasfeature,valueandissue. The issue isunseen_category,out_of_range,unparseable_number,unparseable_date,ambiguous_date_formatordate_order_mismatch.out_of_rangewarnings also havetraining_range. Date format warnings also haveparsed_as, the date that glin used.unparseable_numberis for a numeric input that is not a finite number, for exampleabcorinf. glin then treats the input as missing. A blank input or a suppression code does not give this warning. Eachvalueis safe for JSON: glin sendsinfas the string"inf".missing_featureslists inputs that are absent or blank. These inputs are valid. glin reports them separately.- To send a date, pass the raw value, for example
"signup_date": "2024-03-06". glin derives the samesignup_date__*terms as at training time. Contributions name each term, for examplesignup_date__day_of_week=sat. A date that is absent, blank or not parseable counts as missing. A date that does not parse also adds a warning.
train_model(target_column, model_name, csv_path=None, csv_content=None, overwrite=False, hyperparameters=None, holdout_fraction=0.2, holdout_group_column=None, suppression_codes=None, sheet=None, header_rows=None)trains a new model. It can also retrain an existing model on new data withoverwrite=True. If the name exists andoverwriteis not set, the call fails. Nothing is replaced by accident.- Pass exactly one of
csv_pathandcsv_content. csv_pathis a.csv,.xlsxor.xlsfile that the server can read. See Restrictingcsv_path. Use it for local stdio setups. It works for large files (up to 200 MB by default), and the agent never copies the file contents.csv_contentis the raw CSV text. Use it only when the client and the server share no file system, and only for small data. Tens of thousands of rows in a tool argument are slow and use a lot of context.holdout_group_columnmakes the holdout a grouped holdout.suppression_codesreplaces the tokens that count as missing in numeric columns. The default is["SUPP", "NE", "NP", "z", "*", ".."]. See Feature columns: supported today.- For Excel files (they need
glin-ml[excel]), two more arguments apply.sheetis a name or a 1-based position. It is required for workbooks with several sheets.header_rowsis4,"3-4"or[3, 4]. glin detects it if you omit it. See Excel input. The result shows thesheetandheader_rowsthat glin used. - The result has the kept and dropped features, the resolved
hyperparameters,training_secondsand the term counts.date_columnsmaps each expanded date column to its derived terms. Ifholdout_fraction > 0, the result also has anevaluationblock of out-of-sample metrics. See Let an agent tune the model. warningshas the reason for each dropped column and any "Possible target leakage" entries. See Target leakage check.
- Pass exactly one of
explain_model(model_name, top_n=15, include_shapes=True)shows what the model learned. glin reads it exactly from the additive structure. It does not sample example predictions. The result has:- every feature, ranked by mean absolute log-odds contribution, with its main-effect part and its interaction part;
- the top terms, including pairwise interactions, with their share of the total;
- for binary targets, a shape for each feature that shows which categories or value ranges push toward which class. For a continuous feature, the shape gives the value range (bin edges, in the units of the feature) where the curve is highest and lowest. The edges have 4 significant digits. Values below 10,000,000 do not use exponent notation (for example
21140, not2.114e+04).
list_hyperparameters()lists every EBM hyperparameter. Each entry has its type, default, allowed range and a note on what a higher or lower value does.
Prompts exposed over MCP
-
build_classifier(csv_path="", target_column="", model_name="")runs the whole workflow for any dataset in one prompt:- Inspect and clean the CSV.
- Decide which columns are real predictors. Drop target leakage, such as text that describes the outcome, identifiers, free text and dates after the outcome.
- Tune the hyperparameters with these guardrails against overfitting:
- a fixed holdout;
- a budget of about 10 trials;
- a check of the gap between train and holdout;
- a check of seed noise;
- a preference for the simpler configuration;
- an early exit when the features carry no signal.
- Refit on all rows.
- Report the expected performance honestly.
Select it in the prompt picker of your MCP client. Claude Desktop lists server prompts in the attach (
+) menu. Give it a file path and a target. All other arguments are optional. The prompt is self-contained, because clients show prompts to the user and do not let the agent fetch them. -
tune_model_hyperparameters(model_name="", target_column="")runs the tuning loop in Let an agent tune the model. It includes the live hyperparameter table. glin reads the defaults and ranges from the code. -
prepare_csv_for_training(target_column="")is a checklist for an agent to follow before it callstrain_model. It covers:- parsing problems: banner rows, a wrong delimiter, ragged rows, and the choice of Excel sheet and header rows;
- how to choose a valid target column;
- what the preprocessor already handles: dirty numeric formats, date and timestamp expansion, missing values, and identifier-like, constant and high-cardinality columns;
- what the preprocessor does not handle: dates stored as bare numbers, embedded JSON or list cells, leakage columns and targets that look continuous.
glin reads this guidance from the actual defaults of the preprocessor, so it cannot drift from the code.
Restricting csv_path
train_model(csv_path=...) reads only a regular file that ends in .csv, .xlsx or .xls. The name csv_path stays for compatibility, and it also takes Excel files. The fully resolved path (symlinks and .. followed) must be inside an allowed directory. glin rejects directories, devices, other extensions, missing files, path traversal and symlinks that point outside. The error never includes file contents.
By default, the allowed directories are:
- the home directory;
- the working directory of the server, but only when it is not a filesystem root (
/) and not a parent of the home directory (for example/Users).
Some clients, for example Claude Desktop on macOS, start the server from /. Then only the home directory is allowed. Claude Code starts the server from the project directory, so that directory is also allowed. To set the directories yourself, use a repeatable flag:
If an allowed directory is /, glin writes a warning to stderr, because then csv_path can read any .csv, .xlsx or .xls file on the machine. In Python, build_server(allowed_dirs=[]) allows no directories. Then csv_path is disabled and only csv_content works.
glin rejects a file that is larger than 200 MB. It checks the size before it reads the file. To change the limit, use --max-file-mb, for example glin serve --max-file-mb 500.
csv_content is not affected by these rules. In the config of your MCP client, add the flags to args, for example ["serve", "--allow-dir", "/data/csvs"].
These checks have two known limits. Both need write access inside an allowed directory:
- A hard link inside an allowed directory to a file outside it is read. A hard link is the file itself, so glin cannot see where it came from.
- glin checks the path and then reads it. If the path changes between the check and the read, glin reads the new target.
Thus, do not allow a directory that untrusted users or processes can write to.
The older flag --mode stdio is still accepted so that existing client configs keep working. glin ignores it and writes one deprecation notice to stderr. You can remove it from args. --mode http and --mode sse stop with exit code 2 and this message: "HTTP mode was removed in v0.1.4; v0.1.3 is the last release with it. glin runs over stdio only."
Let an agent tune the model
An EBM is a sum of one learned curve for each feature, plus a few pairwise interaction tables. Cyclic boosting with tiny single-feature trees fits it. Most of its settings trade flexibility against smoothness. train_model(hyperparameters={...}) exposes them: max_bins, max_interaction_bins, interactions, learning_rate, max_rounds, early_stopping_rounds, validation_size, outer_bags, inner_bags, min_samples_leaf, max_leaves, reg_alpha, reg_lambda, smoothing_rounds, greedy_ratio and random_state. Settings that you omit keep their defaults. glin rejects unknown names, wrong types and values out of range. The error message lists every problem. glin/hyperparameters.py defines each setting once. Add or widen a setting there.
An agent needs a score to optimize. train_model holds out holdout_fraction of the rows and trains on the rest. The holdout is stratified. Its split uses a seed that does not depend on random_state, so every trial is scored on the same rows. The result has an evaluation block:
log_lossbaseline_log_loss(a guess from the class frequencies)train_log_loss(compare it withlog_lossto see overfitting or underfitting)accuracybalanced_accuracyroc_auc
glin saves the hyperparameters and the metrics in the metadata of the model.
The tune_model_hyperparameters prompt guides the agent through this loop:
- Train a baseline.
- Change one or two settings at a time.
- Check the seed noise.
- Refit on all rows with
holdout_fraction=0.
Trials must use a scratch model_name with overwrite=True. glin train on the command line trains on every row with the default hyperparameters.
Grouped holdout
Rows from the same state, region, customer, site or user look alike. A random holdout puts near-duplicates on both sides, so the metrics are too optimistic. In one real case, a model scored 82% accuracy and 0.89 AUC on a random holdout. It scored 63% and 0.77 on states that it never saw.
Pass holdout_group_column="state" (or any other column) to train_model. glin then holds out whole groups until about holdout_fraction of the rows are in the holdout. The other rows train the model.
-
The split uses a seed that does not depend on
random_state, so every trial uses the same rows. glin tries 25 seeded group orders. It keeps the order whose holdout class shares are closest to the overall class shares. -
The group column stays a normal feature. A held-out group therefore takes the unseen-category path at prediction time. glin fits the preprocessing on the training rows only.
-
glin groups text values in the same way as the preprocessor: it removes spaces at the start and end and changes letters to lowercase. Thus
"S06"and"s06 "are one group, and that group is on one side of the split only. Blank text,nanandnoneare missing values. If two or more spellings become one group, glin returns a warning. -
evaluationalso reportsholdout_group_column,holdout_groupsandtrain_groups. These counts use the normalised groups. Missing values count as one group. glin also saves the column name in the model metadata. -
glin rejects these cases with a clear error:
- the column does not exist, or it is the target;
- fewer than 5 distinct non-null groups exist;
- one group has more than half of the rows. If this group is the missing values, the error says so. Fill the missing values, or use a different column;
- every group is too large for the holdout target. The error gives the size of the smallest group. Use a larger
holdout_fraction; - the holdout has fewer than 2 classes, or the split leaves a class out of training.
If
holdout_fraction=0, glin ignores the column and returns a warning. -
The
build_classifierandtune_model_hyperparametersprompts tell the agent to use this option when such a column exists. They also tell the agent to report this holdout as the honest estimate. -
glin trainhas no holdout, so it has no equivalent.
Example: ask Claude directly
When a model is trained and the MCP server is attached, talk to Claude in plain language. You do not need to know the tool schema.
"I've got 3 deals to prioritize before quarter close: Northwind Systems ($42k, Opportunity stage, came from a referral, VP contact, owned by Carla Nguyen, 8 touches logged), Summit Retail Group ($8k, still a Lead, paid social source, owned by Brian Kessler), and Anchor Nonprofit ($3k, still a Lead, owned by Frank Suarez, no activity logged yet). Run them through deal_predictor_v1 and tell me which to prioritize."
Claude calls predict once for each deal. It returns a ranked answer with reasons, not only a number:
Note: a field that you do not mention reaches the model as missing. It does not reach the model as "average". Name a rep and a lifecycle stage, and use the categories that the model learned, for example Lead, Opportunity and Customer. This gives a better answer than extra detail in fields that the model already has.
Case study
Child stunting across Indian districts (NFHS-5) uses glin on real public-sector data. A glin model predicts whether a district has high child stunting. Claude calls the model through MCP. The model gives the exact reason for each prediction.
Data requirements
glin train checks your data before it does any work. The data is a CSV, or a sheet of an Excel workbook after header detection. A hard problem stops training with a clear error. An example is a target column with only one class. A soft problem prints a warning, and training continues. An example is a numeric target with many distinct values.
The rules are in glin/validation.py as a flat list. To support a new data shape, add one rule and one preprocessing case. The rest of the pipeline does not change.
In the Python API, a DataFrame can have column names that are not text, for example 0 and 1. glin changes them to text ("0", "1"), and the results use these names. predict accepts either form as a key. Two names that are the same as text, such as 0 and "0", give an error.
Feature columns: supported today
Dates and timestamps
These strings work: 2024-03-06, 2024-03-06 14:30:00, 2024-03-06T14:30:00Z, 03/06/2024, 6 Mar 2024 and Mar 6, 2024. Real datetimes work in the Python API.
- Detection: At least 80% of the non-null values must be full calendar dates. You can change the threshold with
EBMTabularPreprocessor(date_parse_threshold=...). glin checks dates before the identifier and cardinality rules, so a date that is unique for each row stays. - Terms: glin makes
<col>__year,<col>__month(jantodec),<col>__day_of_week(montosun),<col>__hourand<col>__days_since_start.days_since_startcounts whole days since the earliest training date. glin stores that date with the model, andinspect_modelshows it. - Reading explanations: Month and day of week are categorical, so an explanation reads
signup_date__day_of_week = sat. The other terms are ordered numbers. - Dropped parts: glin drops a part that never changes in training. Single-year data has no
__year. Date-only data has no__hour. A column with one distinct date is dropped. - No sin/cos terms: An EBM learns a free-form curve for each feature, so it can already show the December-to-January wrap-around. Sin and cos terms are hard to read.
- Time zones: glin converts values that have a time zone to UTC.
- Day and month order: glin reads
dd/mmormm/ddfrom the values. A field above 12 decides. If nothing decides, glin assumes month first, andtrain_modelreturns a warning. Use ISOYYYY-MM-DDto avoid the warning. - Values that do not parse: They become missing, with a warning.
- Bare numbers are never dates. Years,
20240306, epoch seconds and IDs stay numeric. Month-only and quarter strings (Jan,Q1 2020) and version strings (1.2.3) stay categorical. If the date matters, convert these values to date strings first. - At predict time: Send the raw date. A date outside the training range returns an
out_of_rangewarning. A date that does not parse returnsunparseable_date. A month or weekday that training never saw returnsunseen_category. A numeric date such as05/03/2024on a column with no numeric dates in training (ISO or month names) returnsambiguous_date_formatif the day and the month can change places. glin reads it month first. A numeric date that cannot be in the training order returnsdate_order_mismatch, for example02/13/2024on add/mmcolumn. glin reads it in the other order. Both warnings show the date that glin used inparsed_as. An ISO date never gets these warnings.
Feature columns: not yet supported
Target leakage check
After preprocessing, train_model and glin train look for single feature columns that separate the target almost perfectly on their own. This is the usual sign that someone derived the target from that column with a threshold, a band or a category mapping. The model then only reads the target back.
glin flags such a column in warnings as Possible target leakage: .... The warning shows the evidence: a numeric threshold or band such as > 62.5 -> 'yes', or a category mapping such as {p, q} -> 'ok'. It suggests that you drop the column and retrain if the column is not known at prediction time. The check is only a warning. Training continues. The Python result also has leakage_suspects. The build_classifier, prepare_csv_for_training and tuning prompts tell the agent to review these warnings with the user.
- Method: glin fits a single-feature stump for each feature. It scores the stump by balanced accuracy on the training rows. The holdout rows are not used. For numeric features, the stump is a decision tree with at most
n_classes + 1leaves (one threshold, or a band with two cut points). For categorical features, the stump maps each category to its majority class. glin also uses this category map for a numeric feature with 20 or fewer distinct values, because such a column is often a category stored as integer codes (for example survey answer codes). Missing values are one more category. A target such asq12_code in {2, 5, 7}is not monotone, so a threshold cannot show it, but the category map can. glin flags a feature at balanced accuracy of 0.99 or more. For a binary target, this also means a one-vs-rest AUC of 0.99 or more. A strong but imperfect predictor does not trigger the check. - As the model sees it: glin checks percent, currency and suppression-coded columns after numeric conversion. Missingness counts. A column that is blank exactly when the target is positive is flagged as
missing -> .... glin checks date columns through their derived terms and reports them under the raw column, for example'signup' (via derived term 'signup__days_since_start'). - Guards against chance and memorization:
- Each numeric leaf needs at least 5 rows (more on bigger data).
- Each category needs at least 5 rows. This floor does not increase with the data size. Thus glin can also flag a column with 100 to 250 categories (the preprocessor keeps at most 250).
- A categorical column must have at least 90% of its rows in categories with 5 or more rows. A near-unique column therefore cannot "separate" the target by memorization. In a test, 200 noise columns with 250 categories and 20,000 rows gave no flags.
- glin leaves out classes with fewer than 10 rows.
- glin skips the check below 30 rows.
- The cost is a few milliseconds for each column. glin samples at most 20,000 rows. 500 columns take about 2 seconds.
- glin reports at most 10 columns, the best first.
Target column requirements
- The target needs at least 2 distinct non-null values (binary or multiclass). A single-class target is a hard error.
- glin drops rows with a missing target value and returns a warning. The other rows are not affected.
- A numeric target with more than 20 distinct values triggers a warning that it looks like a regression target. glin trains classifiers, not regressors, so each value would become its own unrelated class. Band the target into a few classes (quantiles or domain thresholds), or pick another target.
- A numeric target with 3 to 20 distinct values (for example a score or an ordinal level) triggers a warning. The multiclass model ignores the order of the values. The warning gives an example with the real values of the target. Use fewer bands or a binary threshold. The
prepare_csv_for_trainingandbuild_classifierprompts give the same guidance. - These two checks also apply to a text target whose values are numbers, for example
"85%"or"1"to"5". glin parses the values with the same rules as a numeric feature column. The model still uses each text value as a class. - A text target with more than 20 distinct values triggers a warning. Each value becomes its own class, and each class needs many rows.
- Full regression (the upstream
ExplainableBoostingRegressor) is planned separately. It is not yet supported.
Releasing (maintainers)
Only the Publish to PyPI GitHub Actions workflow (.github/workflows/publish.yml) publishes releases to PyPI. Nobody uploads from a local machine. Do not use twine upload, uv publish or personal API tokens. To release, push a tag of the form vX.Y.Z:
The workflow runs the Tests workflow first (the same matrix as for pull requests). If a test fails, the workflow stops and does not publish. Then it builds the sdist and the wheel, and checks that both carry the version of the tag. Then it publishes them to PyPI as glin-ml at that exact version, and creates the GitHub Release from CHANGELOG.md.
hatch-vcsderives the package version from the git tag.pyproject.tomlhas no fixed version, so you do not bump the package version before a release. The plugin and registry files have their own versions (see Release steps).- An untagged local build gets a dev version such as
0.1.4.dev1+g3fa91a0. A source tree with no git metadata reports0.0.0. glin.__version__reads the installed package metadata.- The workflow ignores tags that do not match
vX.Y.Z. - PyPI rejects an upload of a version number that already exists. To fix a bad release, tag a new version.
Release steps
The owner does these steps for each release:
- In
CHANGELOG.md, move the entries under[Unreleased]to a new heading## [X.Y.Z] - YYYY-MM-DD. Add the compare link at the bottom. Compare the entries withgh issue list --milestone vX.Y.Z --state closed. - Set the release version in
server.json(versionandpackages[0].version) and inplugins/glin/.claude-plugin/plugin.json(version). If the plugin needs the new release, also change the minimum version inplugins/glin/.mcp.json. Merge these changes intomain. - Tag the release and push the tag, for example
git tag v0.1.4andgit push origin v0.1.4. - In GitHub Actions, open the
Publish to PyPIrun and approve thepypienvironment. The upload step starts only after this approval. - The workflow publishes to PyPI. Then it creates the GitHub Release with the
CHANGELOG.mdsection of the version as the notes. IfCHANGELOG.mdhas no section for the version, the workflow stops before it publishes. - Check that the new version is on PyPI and that the GitHub Release exists.
- Publish to the MCP Registry with
mcp-publisher login githubandmcp-publisher publish. See Publishing to the MCP Registry.
To compare the changelog with the merged pull requests, open a draft release on GitHub and click "Generate release notes". .github/release.yml groups the pull requests by label.
One-time setup that makes this the only path
The workflow publishes through PyPI trusted publishing (OIDC). This also makes local uploads impossible.
- PyPI → the
glin-mlproject → Publishing: add a trusted publisher for ownerAkashChatterjee, repositoryglin, workflowpublish.ymland environmentpypi. - PyPI → Account settings → API tokens: delete every API token that can upload
glin-ml. Do not create new tokens. With no tokens and 2FA on, the only credential that can publish is the short-lived OIDC identity of the workflow. Also review the collaborators of the project. Remove anyone who does not need upload rights. - GitHub → Settings → Rules → Rulesets: add a tag ruleset for
v*. Restrict who can create tags, and block deleting or moving them. A new matching tag starts a release. - GitHub → Settings → Environments →
pypi: add required reviewers. Every release then needs a manual approval before the upload step runs.
Publishing to the MCP Registry
The registry holds metadata only. It checks that the PyPI package has the line mcp-name: io.github.AkashChatterjee/glin in its README. The README has this line as a hidden comment. Publish to PyPI first, then to the registry:
- Release the new version on PyPI (see above).
- Check that
versionandpackages[0].versioninserver.jsonare the released version. - Run these commands from the repository root:
- Check the result:
curl "https://registry.modelcontextprotocol.io/v0.1/servers?search=io.github.AkashChatterjee/glin".
The registry rejects a version that already exists. Update server.json for each release.
Checking a release
To check a build without publishing it, run uv build locally and inspect dist/. Never upload it. After a release, confirm that the public package works with a clean install:
来源:README.md,提交 022f702
工具
0版本历史
1- v0.1.4最新Oct 11, 2026


