GLM-4.7-Flash-UD-Q3_K_XL Unsloth - High Reasoning Benchmark

Cancelled throughput run

This file records the Hermes benchmark attempt for the local LM Studio model GLM-4.7-Flash-UD-Q3_K_XL from Unsloth. The run was intentionally cancelled because the model was too slow on this machine to complete the benchmark in a useful amount of time.

Run Summary

Benchmark filetest-prompt.md
Evaluation criteriaevaluation-criteria.md
Hermes active model idglm-4.7-flash
Model label for resultGLM-4.7-Flash-UD-Q3_K_XL, Unsloth variant
Reasoning efforthigh, confirmed by Hermes before the prompt was submitted
Provider/backendLM Studio local server at 127.0.0.1:1234
Started2026-05-27 03:36:57 +0600
Cancelled2026-05-27 04:11:05 +0600
Total elapsed after prompt submission34m 08s
Hermes session duration34m 28s
Hermes session id20260527_033642_aebcff
Final statusCancelled / no completed benchmark artifact from Hermes

Observed Timing

34m 08sTotal elapsed after benchmark prompt submission
9m 02sFirst LM Studio prompt-processing pass, 03:37:00 to 03:46:02
15m 15sSecond LM Studio prompt-processing pass, 03:48:56 to 04:04:11
1329.3sHermes-reported wait during the interrupted API call

LM Studio logs showed prompt processing was still advancing, so the run was not interrupted during the long prompt-processing phase. It was cancelled only after processing had completed and generation remained impractically slow. The observed token generation speed was below 1 token/second on this rig.

Tool Calls And Execution Detail

Hermes command Started a fresh Hermes CLI session, confirmed active model glm-4.7-flash, and set /reasoning high.
Prompt submitted Hermes was instructed to read /Users/armanshawon/Documents/Benchmark/test-prompt.md, follow it exactly, and create the required benchmark output file in the benchmark folder.
Observed Hermes tool call read_file on /Users/armanshawon/Documents/Benchmark/test-prompt.md. Hermes displayed the read as completed in about 1.1s.
Observed file output No GLM result file was created before cancellation. The folder still only contained the prior GPT-5.5 result files and benchmark inputs.
Interrupt detail Hermes was interrupted during an API call. Hermes reported: Operation interrupted: waiting for model response (1329.3s elapsed).
Error count 0 model/tool errors observed. This was a throughput cancellation, not a correctness failure after repeated errors.

Cancellation Rationale

Benchmark Outcome

Capability scoreNot scored
Tool JSON scoreNot scored
Coding scoreNot scored
Writing scoreNot scored
Context scoreNot scored
Practicality noteFailed practical usability for this benchmark because local generation speed was below 1 token/second.