OnpremBench.aiby Understand.tech
← All deployment packages

Coding / by Understand Tech Lab

Code with a local assistant

Test useful code changes on a fixed repository, with quality and completion time measured together.

Untested evaluation templateVersion 0.1.0No lab-verified results

Before running

Make the setup exact.

Model: Select a coding model; pin model revision, precision and context settings.

Stack: Local inference runtime + editor or agent client. Pin the client and tool configuration.

Record your hardware SKU, memory allocation, topology, dataset revision and thresholds. The downloaded manifest is a starting specification; it does not install software.

Platform documentation

Evaluation plan

  1. Choose a shareable repository at a fixed commit and five small issues with tests. Keep a clean copy for each attempt.
  2. Select a compatible local model and client using the platform documentation. Record every version and endpoint setting.
  3. Run each issue with the same prompt and tool permissions. Save diffs, tool calls, token counts and elapsed time.
  4. Run the repository tests and inspect changes manually. Repeat each task three times from the same starting commit.
  5. Increase concurrency only after single-request quality is acceptable. Record failures and memory usage.

What counts as useful?

  • Count tasks accepted after tests and human review.
  • Measure time and energy per accepted task, including retries.
  • Record context length, active requests, errors and model-loading time.

Known limits

Passing repository tests alone does not prove code correctness.

Results depend on client tools and prompts as well as the model.