{"v":2,"api":"nishi-compare","generated_unix":1784545533,"domain":"autograde","kind":"dashboard","title":"Is our system standing on its own? -- live scoreboard","plain_english":true,"measured_by":"the system itself (nx_autograde runs the tests; this page only translates the results)","scoreboard":[{"question":"Can we automatically check whether generated code is correct?","answer":"yes"},{"question":"Can the system find a bug, ask permission, fix it, and undo the fix?","answer":"yes"},{"question":"How much of the work runs with zero Claude cost?","answer":"70 percent, goal 100"},{"question":"Does the code test suite itself work?","answer":"yes"},{"question":"Is our own compiler as fast as the industry standard (gcc)?","answer":"not yet"},{"question":"Can the local AI produce working code by itself?","answer":"yes for small fixes: it finds the failing function itself, writes the fix, and the machine verifies it -- zero Claude"},{"question":"Graded by the industry bug-fixing benchmark's own rules (SWE-bench contract), what's our resolve rate?","answer":"71 percent of our foreign-bug instances resolved (fix passes, nothing else breaks); sovereign analog, not the Python set"}],"claude_free_trend_percent":[{"local_model":"0.5B","percent":38},{"local_model":"1.5B","percent":54},{"local_model":"1.5B","percent":61},{"local_model":"1.5B","percent":70}]}
