Failed - cancelled for unusable speed
Qwen3.6-35B-A3B-MTP-IQ2_M-Unsloth Hermes Benchmark
The benchmark was started correctly, but the model did not produce a usable first response or any completed tool call before the run was cancelled for speed.
Run Summary
Model label
Qwen3.6-35B-A3B-MTP-IQ2_M-Unsloth
LM Studio model id
qwen3.6-35b-a3b-mtpHermes session
20260527_044555_47e006Reasoning
High
Hermes duration
13m 54s
Outcome
No benchmark artifact produced by Hermes
Timing
| Event | Observed time | Detail |
|---|---|---|
| Prompt submitted | 2026-05-27 04:46:20 +0600 | Hermes was pointed at /Users/armanshawon/Documents/Benchmark/test-prompt.md and instructed to write the benchmark output inside /Users/armanshawon/Documents/Benchmark. |
| Prompt processing | 04:46:20 to 04:50:42 | LM Studio progressed from 0.0% to 100.0% in about 4m 22s. A duplicate 100.0% entry appeared at 04:50:55. |
| Waiting for first usable model response | 04:50:55 to interruption | Hermes remained in the API call waiting for model response. At interruption Hermes reported 803.0s elapsed. |
| Cancellation | About 04:59:44 +0600 | User called the run for speed. Hermes was interrupted and exited cleanly. |
| Post-run model state | 04:59:57 +0600 | lms ps showed qwen3.6-35b-a3b-mtp back to IDLE. |
Tool Call Detail
| Tool | Status | Observation |
|---|---|---|
| Hermes model switch | Completed | Hermes was switched from the previous model to qwen3.6-35b-a3b-mtp. |
| Hermes reasoning setting | Completed | /reasoning high was accepted before the benchmark prompt was submitted. |
| Benchmark task tools | None completed | Hermes session summary reported 0 tool calls. No read, search, browser, write, or extraction tool call completed before cancellation. |
| Error count | 0 | No actual tool errors were observed. The run failed on speed, not on repeated tool-call errors. |
Failure Reason
The model was unusably slow for this Hermes benchmark setup. Prompt ingestion completed, but after more than 13 minutes of total Hermes runtime and approximately 803 seconds waiting on the API call, the model had not returned a usable first response or invoked any tool. Since the benchmark requires a practical agent run with file creation and tool use, this run is marked failed.
Result classification: speed failure before first agent action. This is different from a reasoning/tool-use failure because the model never reached a completed benchmark step.