GLM-4.7-Flash-UD-Q3_K_XL Unsloth - High Reasoning Benchmark
Cancelled throughput run
This file records the Hermes benchmark attempt for the local LM Studio model GLM-4.7-Flash-UD-Q3_K_XL from Unsloth. The run was intentionally cancelled because the model was too slow on this machine to complete the benchmark in a useful amount of time.
Run Summary
| Benchmark file | test-prompt.md |
|---|---|
| Evaluation criteria | evaluation-criteria.md |
| Hermes active model id | glm-4.7-flash |
| Model label for result | GLM-4.7-Flash-UD-Q3_K_XL, Unsloth variant |
| Reasoning effort | high, confirmed by Hermes before the prompt was submitted |
| Provider/backend | LM Studio local server at 127.0.0.1:1234 |
| Started | 2026-05-27 03:36:57 +0600 |
| Cancelled | 2026-05-27 04:11:05 +0600 |
| Total elapsed after prompt submission | 34m 08s |
| Hermes session duration | 34m 28s |
| Hermes session id | 20260527_033642_aebcff |
| Final status | Cancelled / no completed benchmark artifact from Hermes |
Observed Timing
34m 08sTotal elapsed after benchmark prompt submission
9m 02sFirst LM Studio prompt-processing pass, 03:37:00 to 03:46:02
15m 15sSecond LM Studio prompt-processing pass, 03:48:56 to 04:04:11
1329.3sHermes-reported wait during the interrupted API call
LM Studio logs showed prompt processing was still advancing, so the run was not interrupted during the long prompt-processing phase. It was cancelled only after processing had completed and generation remained impractically slow. The observed token generation speed was below 1 token/second on this rig.
Tool Calls And Execution Detail
| Hermes command | Started a fresh Hermes CLI session, confirmed active model glm-4.7-flash, and set /reasoning high. |
|---|---|
| Prompt submitted | Hermes was instructed to read /Users/armanshawon/Documents/Benchmark/test-prompt.md, follow it exactly, and create the required benchmark output file in the benchmark folder. |
| Observed Hermes tool call | read_file on /Users/armanshawon/Documents/Benchmark/test-prompt.md. Hermes displayed the read as completed in about 1.1s. |
| Observed file output | No GLM result file was created before cancellation. The folder still only contained the prior GPT-5.5 result files and benchmark inputs. |
| Interrupt detail | Hermes was interrupted during an API call. Hermes reported: Operation interrupted: waiting for model response (1329.3s elapsed). |
| Error count | 0 model/tool errors observed. This was a throughput cancellation, not a correctness failure after repeated errors. |
Cancellation Rationale
- The model completed prompt processing slowly but steadily, so it was not classified as stuck during processing.
- After prompt processing reached 100%, generation did not progress to a usable result within the available time.
- Throughput was observed below 1 token/second, making this model impractical for this Hermes benchmark on the current GPU setup.
- Because Hermes never produced the required result file, capability scoring from
evaluation-criteria.mdwas not applied.
Benchmark Outcome
| Capability score | Not scored |
|---|---|
| Tool JSON score | Not scored |
| Coding score | Not scored |
| Writing score | Not scored |
| Context score | Not scored |
| Practicality note | Failed practical usability for this benchmark because local generation speed was below 1 token/second. |