Cnki Paper Detail

cookjohn/cnki-skills/skills/cnki-paper-detail

by cookjohn20d65f660456daf53ad0f7c74494ac3b829b925fNo license983 starsListed Oct 9, 2026Updated Oct 9, 2026Repository updated 7 months ago

Extract full paper details from a CNKI paper page including title, authors, affiliations, abstract, keywords, fund, classification. Use when the user needs detailed information about a specific paper.

Instructions onlyResearch & Analysis
AI-generated overview

Extracts full metadata from a CNKI paper detail page, including title, authors, abstract, keywords and citation counts.

What it does
The skill drives a Chrome DevTools browser session to a CNKI paper detail page and reads its metadata. It returns the title, authors with affiliation numbers, affiliations, abstract, keywords, fund information, classification code, journal, publication info, table of contents and citation network counts. If script-based extraction fails, it falls back to parsing the page accessibility snapshot.
When to use it
Use it when a user needs detailed information about one specific CNKI paper and can supply the paper detail URL or is already on that page. It suits requests for a paper's authors, abstract, keywords, funding or citation counts.
Requirements
Requires a Chrome DevTools MCP browser session with navigation, snapshot and script evaluation tools, plus network access to CNKI. A manual slider captcha may need to be solved by the user. No scripts are shipped; it is instructions only.

CNKI Paper Detail Extraction

Extract complete metadata from a CNKI paper detail page.

Arguments

$ARGUMENTS is optionally a CNKI paper detail URL (containing kcms2/article/abstract). If not provided, assumes the current page is already a paper detail page.

Steps

1. Navigate to the paper page (if URL provided)

If $ARGUMENTS contains a URL:

  • Use mcp__chrome-devtools__navigate_page with the URL.
  • Use mcp__chrome-devtools__wait_for with text ["摘要"] and timeout 15000.

2. Check for captcha

Use mcp__chrome-devtools__take_snapshot. If "拖动下方拼图完成验证" found, notify user:

CNKI 正在显示滑块验证码。请在 Chrome 浏览器中手动完成拼图验证,完成后告诉我继续。

3. Extract paper metadata via JavaScript

Use mcp__chrome-devtools__evaluate_script with this function:

javascript
() => {  const brief = document.querySelector('.brief');  if (!brief) return { error: 'Paper detail section (.brief) not found' };
  // Title  const title = brief.querySelector('h1')?.innerText?.trim()    ?.replace(/\s*附视频\s*$/, '')  // remove "附视频" suffix    ?.replace(/\s*网络首发\s*$/, ''); // remove "网络首发" suffix
  // Authors - first h3.author contains author links with sup tags  const authorH3s = brief.querySelectorAll('h3.author');  const authorSection = authorH3s[0];  const authors = [];  if (authorSection) {    const authorLinks = authorSection.querySelectorAll('a');    authorLinks.forEach(a => {      const name = a.innerText?.replace(/\d+$/, '').trim();      const supMatch = a.innerText?.match(/(\d+)$/);      const affiliationNum = supMatch ? supMatch[1] : '';      authors.push({ name, affiliationNum });    });  }
  // Affiliations - second h3.author contains org links  const affiliations = [];  if (authorH3s.length > 1) {    const orgLinks = authorH3s[1].querySelectorAll('a');    orgLinks.forEach(a => {      affiliations.push(a.innerText?.trim());    });  }
  // Abstract  const abstractEl = document.querySelector('.abstract-text');  const abstract = abstractEl?.innerText?.trim() || '';
  // Keywords  const keywordsP = document.querySelector('p.keywords');  const keywords = keywordsP    ? Array.from(keywordsP.querySelectorAll('a')).map(a => a.innerText?.replace(/;$/, '').trim())    : [];
  // Fund  const fundsP = document.querySelector('p.funds');  const fund = fundsP?.innerText?.trim() || '';
  // Classification code  const clcCode = document.querySelector('.clc-code');  const classification = clcCode?.innerText?.trim() || '';
  // Journal/source  const docTop = document.querySelector('.doc-top');  const journal = docTop?.querySelector('a')?.innerText?.trim() || '';
  // Online first / publication info  const headTime = document.querySelector('.head-time');  const pubInfo = headTime?.innerText?.trim() || '';
  // Is online first?  const isOnlineFirst = !!brief.querySelector('.icon-shoufa');
  // Article outline/TOC  const catalogList = document.querySelector('.catalog-list, .catalog-listDiv');  const toc = catalogList?.innerText?.trim() || '';
  // Citation network counts  const citationTabs = document.querySelectorAll('ul.module-tab.tpl_lieteratures li');  const citationInfo = {};  citationTabs.forEach(li => {    const id = li.getAttribute('data-id');    const text = li.innerText?.trim();    const countMatch = text.match(/(\d+)/);    if (id) {      citationInfo[id] = {        label: text.replace(/\d+/, '').trim(),        count: countMatch ? parseInt(countMatch[1]) : 0      };    }  });
  return {    title,    authors,    affiliations,    abstract,    keywords,    fund,    classification,    journal,    pubInfo,    isOnlineFirst,    toc,    citationInfo  };}

4. Format and present the output

## {title} {isOnlineFirst ? "[网络首发]" : ""}
**Authors:**{For each author: "- {name} ({affiliation})"}
**Affiliations:**{For each affiliation: "- {affiliation}"}
**Journal:** {journal}**Publication Info:** {pubInfo}
**Abstract:**{abstract}
**Keywords:** {keywords joined by ", "}
**Fund:** {fund}**Classification:** {classification}
**Citation Network:**{For each citation type: "- {label}: {count}"}

5. Fallback: snapshot-based parsing

If JS extraction fails, use mcp__chrome-devtools__take_snapshot and parse the accessibility tree:

  • Title: heading level 1 element
  • Authors: link elements whose URLs contain kcms2/author/detail
  • Affiliations: link elements whose URLs contain kcms2/organ/detail
  • Abstract: StaticText following "摘要:"
  • Keywords: link elements whose URLs contain kcms2/keyword/detail
  • Fund: link elements following "基金资助:"
  • Classification: StaticText following "分类号:"

Verified DOM Selectors

DataSelectorNotes
Paper section.briefMain paper info container
Title.brief h1May contain icons, clean text needed
Authors.brief h3.author:first-of-type aText has superscript numbers (e.g., "张三1")
Affiliations.brief h3.author:nth-of-type(2) aText starts with "N." (e.g., "1.北京大学")
Abstract.abstract-textFull abstract text
Keywordsp.keywords aSemicolon-separated keyword links
Fundp.fundsFund information text
Classification.clc-codeCLC classification codes
Journal.doc-top aSource journal link
Online first.brief .icon-shoufaPresent if paper is online first
Citation tabsul.module-tab.tpl_lieteratures lidata-id attr identifies type

Source and attribution

Source:cookjohn/cnki-skillsinskills/cnki-paper-detailat commit20d65f6

License: No license

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal