← All creators

awslew

User

Generic vision core + pluggable Claude Code scene skills for text-only AI agents (DeepSeek/GLM). OCR / table->markdown / chart reading / UI review. Fork of Anionex/agent-vision-toolkit.

5 indexed · 0 Featured · 1 stars · avg score 74
Prolific

Categories

Indexed Skills (5)

Data & Documents Listed

chart-reading

从图表图片提取结构化数据:柱状图/折线图/饼图/散点图等。读取数值、坐标轴、 图例,输出 JSON + Markdown 摘要,标注哪些值是精确标注、哪些是刻度线估读。 用户说"读这个图的数据""图表提取""这个柱状图/折线图/饼图的数值是多少" 并给图表图时自动触发。

1 Updated 4 weeks ago
awslew
Web & Frontend Listed

ui-feedback

给无视觉底座(DeepSeek 等)提供 UI 设计反馈的闭环:截图 → 视觉"眼睛" (MiMo-V2.5 等视觉模型)最大化信息量分析 → 结构化反馈报告,直接喂回给 设计 UI 的模型用于迭代优化。用户说"评审这个UI/设计反馈/界面怎么改/还原 这个页面/这个布局配色怎么优化"并给出 UI 截图时自动触发。

1 Updated 4 weeks ago
awslew
Data & Documents Listed

vision-core

Scene-agnostic vision core for ds-vision-kit: glance (describe/ask/OCR an image), ground (locate a target, pixel box), detect (element inventory), trace (image to SVG geometry), crop (cut a pixel box to a file), plus local pixel tools palette / pixel-diff / extract-fg / html-shot / long-ocr. Use for ANY image task — the scene skills (ui-feedback, ocr-extract, chart-reading, image-qa) build on this. See ../references/scenes.md to pick a scene.

1 Updated 4 weeks ago
awslew
Testing & QA Listed

image-qa

通用图像问答与元素定位:描述一张图、回答关于图片的问题、定位图中的元素、 对比两张图、判断照片/文档/示意图内容。任何"看图说话"类请求的兜底场景。 用户说"这是什么图/图里有什么/描述一下/找一下图中X/对比这两张图/判断这 个截图"并给图时自动触发;不属于 UI 评审 / OCR 提取 / 图表读取的看图请求 都落这里。

1 Updated 4 weeks ago
awslew
Data & Documents Listed

ocr-extract

把图片里的文字逐字提取为结构化 Markdown:正文/代码/表格。OCR 文字提取、 表格转 Markdown、扫描件转录、长截图/聊天记录转文本。用户说"把这张图的 文字/表格提取出来""OCR""转录这段截图""这个扫描件转文字""表格转markdown" 并给图时自动触发。

1 Updated 4 weeks ago
awslew

Bio shown is the top-scored skill's repo description as a fallback — real GitHub bios land in a future update.