autoresearch
The self-improving lab notebook
You named it autoresearch, which is either a bold research agenda or the moment the experiment started assigning itself homework. Either way, the premise has excellent mad-scientist energy.
Profiles, repos, scores & roasts
Ship machine
silver medal
Overall
0.0/ 100
autoresearch
You named it autoresearch, which is either a bold research agenda or the moment the experiment started assigning itself homework. Either way, the premise has excellent mad-scientist energy.
llm.c
llm.c takes the grand language-model conversation and translates it into the language of kernels, memory, and compile errors: apparently even giant ideas must eventually fit through a pointer.
llm-council
llm-council turns model comparison into a parliamentary system, because one language model was apparently not enough bureaucracy for a single prompt.
hn-time-capsule
hn-time-capsule preserves old discussion as though online discourse were a fine wine. A charmingly ambitious attempt to make the comment section historically significant.
llm101n
llm101n sounds like a complete curriculum, while its archived status gives it the energy of a semester that ended with an impressive final project and no sequel syllabus.
Impact
25% weight
The portfolio centers on ambitious, distinctive educational and experimental machine-learning projects, including implementations, training-oriented work, and research-oriented tools.
Consistency
15% weight
The supplied repositories show repeated substantial activity from 2024 through 2026 across multiple related projects, though the evidence does not establish a complete contribution history.
Quality
25% weight
Non-fork ownership, substantial project scope, clear specialization, and continued updates provide strong maturity signals; source-level engineering quality cannot be assessed from the available metadata.
Depth
20% weight
Multiple major repositories show sustained ownership and meaningful scope, with activity distributed across foundational neural-network education, language-model systems, and research tooling.
Breadth
10% weight
The work spans Python, CUDA, and Jupyter Notebook, with educational, implementation, experimentation, community-oriented, and information-retrieval project types, though the domain focus is concentrated.
Community
5% weight
The supplied follower and star counts indicate substantial public attention, but community evidence is used only as a weak positive tie-breaker and does not raise the other categories.
Based on the bounded repository sample saved with this analysis.
Non-fork sample
12
Last 12 months
367
Accepted fact
215796
Profile date
Apr 2010
Scores marked Profile rating use bounded public repository metadata. Full repo analysis appears only when a separate accepted repository analysis exists.
karpathy /
92/100
An ambitious Python research and automation project with very recent activity and unusually broad apparent experimentation scope.
karpathy /
89/100
A substantial CUDA project targeting language-model work at a lower systems level, giving the portfolio notable computational depth.
karpathy /
88/100
A focused Python implementation project centered on compact language-model experimentation, with sustained ownership and recent updates.
karpathy /
87/100
A large, active Python project focused on accessible language-model experimentation and tooling, with strong apparent ambition.
karpathy /
84/100
A substantial notebook-centered neural-network learning curriculum with clear educational scope and continued activity.
karpathy /
83/100
A compact foundational autodiff and neural-network project presented through notebooks, combining conceptual clarity with meaningful implementation scope.
karpathy /
82/100
An active Python project exploring multi-model coordination or evaluation, with an unusual and distinctive premise.
karpathy /
80/100
A maintained Python language-model implementation project that reinforces the portfolio’s sustained focus on understandable, compact systems.
karpathy /
76/100
A Python project centered on generative modeling and language-oriented experimentation, contributing foundational depth to the portfolio.
karpathy /
73/100
A long-running Python research-information project that broadens the portfolio into literature discovery and preservation.
karpathy /
72/100
An archived language-model education project with substantial apparent scope, though its archived status limits evidence of ongoing maintenance.
karpathy /
66/100
A smaller Python project built around preserving or exploring historical online discussion, showing useful thematic breadth beyond model implementation.
How this score was produced
Overall = Σ(category × weight) + deterministic top-end curve
| Category | Weight | Score | Contribution |
|---|---|---|---|
| Impact | 25% | 92 | 23.00 |
| Quality | 25% | 78 | 19.50 |
| Depth | 20% | 89 | 17.80 |
| Consistency | 15% | 82 | 12.30 |
| Breadth | 10% | 68 | 6.80 |
| Community | 5% | 50 | 2.50 |
What it measures. Project usefulness, ambition, originality, and coherence. Popularity is not impact.
Evidence for karpathy. The portfolio centers on ambitious, distinctive educational and experimental machine-learning projects, including implementations, training-oriented work, and research-oriented tools.
What it measures. Engineering and project-maturity signals supported by the available public metadata.
Evidence for karpathy. Non-fork ownership, substantial project scope, clear specialization, and continued updates provide strong maturity signals; source-level engineering quality cannot be assessed from the available metadata.
What it measures. Sustained ownership, meaningful scope, and continued maintenance rather than one-shots.
Evidence for karpathy. Multiple major repositories show sustained ownership and meaningful scope, with activity distributed across foundational neural-network education, language-model systems, and research tooling.
What it measures. Recency and contribution patterns across the supplied activity window.
Evidence for karpathy. The supplied repositories show repeated substantial activity from 2024 through 2026 across multiple related projects, though the evidence does not establish a complete contribution history.
What it measures. Language entropy and project-type diversity across owned repos.
Evidence for karpathy. The work spans Python, CUDA, and Jupyter Notebook, with educational, implementation, experimentation, community-oriented, and information-retrieval project types, though the domain focus is concentrated.
What it measures. A weak positive tie-breaker for supplied community signals, never a popularity penalty.
Evidence for karpathy. The supplied follower and star counts indicate substantial public attention, but community evidence is used only as a weak positive tie-breaker and does not raise the other categories.
Validate. The server validates the login, signed browser session, attempts, cooldown, daily ceiling, and the single provider permit.
Collect. The VPS collector reads bounded public profile metadata and up to 12 recent repository metadata rows in volatile memory.
Rate. One pinned provider returned a closed profile-rating object with no tools, browsing, shell, or repository-content access.
Verify. Server checks bound the subject, rejected unsafe or echoed text, validated the schema, and recomputed rubric-v3 final arithmetic.
Save. Only a valid completion and allowlisted post-completion facts commit atomically; the public projection is then rebuilt.
This saved profile analysis is metadata-only and did not read repository trees, README files, or source code. Availability of new analysis is separate from this historical result.
Rated 2026-08-10 · rating-rubric/3