Code & Scripts

DeepSeek V4 vs Claude vs Local LLM Benchmark Kit

A standardized Python benchmark suite with a simple results dashboard that lets Mac Studio and local LLM users compare DeepSeek V4, Ollama models, and Claude API on real-world coding, reasoning, and agentic tasks.

$ tar -tzf deepseek-v4-benchmark-kit.tar.gz
  • benchmark/run.py — CLI entrypoint; run all suites or target a single model in one command
  • benchmark/models.py — provider adapters for DeepSeek API, Anthropic API, and Ollama (local)
  • benchmark/cost.py — pricing table and token-to-USD calculator, updatable when rates change
  • tasks/coding/ — 4 tasks: debug, refactor, write unit tests, explain a traceback
  • tasks/reasoning/ — 4 tasks: logic puzzle, multi-step math, constraint scheduling, causal chain
  • tasks/agentic/ — 4 tasks: plan a script, decompose a feature, write a bash pipeline, generate a PR description
  • dashboard/ — terminal leaderboard + static HTML report (no CDN, fully offline)
  • docs/task-authoring-guide.md — add your own tasks in minutes
  • docs/model-pricing-table.md — current pricing reference for all supported models
$29

one-time purchase

Instant download after purchase
📧Download link sent to your email
🔄7-day download access
14-day money-back guarantee
View refund policy

## goes-well-with

From the same shelf

## not-ready-to-buy

Take the field notes instead

One practical write-up a week from the same workbench these kits come from — plus reader pricing when new kits ship.