Failed - cancelled for unusable speed

Qwen3.6-35B-A3B-MTP-IQ2_M-Unsloth Hermes Benchmark

The benchmark was started correctly, but the model did not produce a usable first response or any completed tool call before the run was cancelled for speed.

Run Summary

Model label
Qwen3.6-35B-A3B-MTP-IQ2_M-Unsloth
LM Studio model id
qwen3.6-35b-a3b-mtp
Hermes session
20260527_044555_47e006
Reasoning
High
Hermes duration
13m 54s
Outcome
No benchmark artifact produced by Hermes

Timing

Event Observed time Detail
Prompt submitted 2026-05-27 04:46:20 +0600 Hermes was pointed at /Users/armanshawon/Documents/Benchmark/test-prompt.md and instructed to write the benchmark output inside /Users/armanshawon/Documents/Benchmark.
Prompt processing 04:46:20 to 04:50:42 LM Studio progressed from 0.0% to 100.0% in about 4m 22s. A duplicate 100.0% entry appeared at 04:50:55.
Waiting for first usable model response 04:50:55 to interruption Hermes remained in the API call waiting for model response. At interruption Hermes reported 803.0s elapsed.
Cancellation About 04:59:44 +0600 User called the run for speed. Hermes was interrupted and exited cleanly.
Post-run model state 04:59:57 +0600 lms ps showed qwen3.6-35b-a3b-mtp back to IDLE.

Tool Call Detail

Tool Status Observation
Hermes model switch Completed Hermes was switched from the previous model to qwen3.6-35b-a3b-mtp.
Hermes reasoning setting Completed /reasoning high was accepted before the benchmark prompt was submitted.
Benchmark task tools None completed Hermes session summary reported 0 tool calls. No read, search, browser, write, or extraction tool call completed before cancellation.
Error count 0 No actual tool errors were observed. The run failed on speed, not on repeated tool-call errors.

Failure Reason

The model was unusably slow for this Hermes benchmark setup. Prompt ingestion completed, but after more than 13 minutes of total Hermes runtime and approximately 803 seconds waiting on the API call, the model had not returned a usable first response or invoked any tool. Since the benchmark requires a practical agent run with file creation and tool use, this run is marked failed.

Result classification: speed failure before first agent action. This is different from a reasoning/tool-use failure because the model never reached a completed benchmark step.