Rendered at 11:19:31 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
trashymctrash 48 minutes ago [-]
Opus 5.5 has reduced this to an acceptable level for me. Are you also using this plugin with that model?
5 hours ago [-]
citizenfishy 40 minutes ago [-]
Claude Mods seem to answer this pain better without affecting the model reasoning
TZubiri 44 minutes ago [-]
Read up on Chain of Thought
The model is essentially thinking out loud, when you ask it to be more concise, you make it think less, therefore producing more erroneous answers.
Some models have an internal chain of thought (claude being one of them), which sometimes isn't even published to avoid reverse engineering, but it seems that this might still be a problem.
What you'd want actually is a layer that summarizes the actual answer, but that's actually an internal prompt by claude that you are not seeing, the model just doesn't expose the necessary bits for you to hack this together.
Try another model that exposes the raw llm output instead of exposing a CoT result directly.
Of course the real hack is learning to read diagonally without reading every single word, this is a skill that is useful in general. It's also less effort in general, instead of making plugins and super customizing the thing, you just consume the default settings, which are hyperoptimized, and require no time spent in configuration.
cyanydeez 37 minutes ago [-]
I think you're humanizing too much.
The <think> blocks are an attempt to explore the gradient descent space to escape local minimums and find a better global minimum to continue the descent.
While verbosity _might_ do this better, you could easily consider things like "but wait am I forgetting ...." as just one token. So if you actually do it right, you could replace all that with a "hold on" or something of a terse variety.
The model is essentially thinking out loud, when you ask it to be more concise, you make it think less, therefore producing more erroneous answers.
Some models have an internal chain of thought (claude being one of them), which sometimes isn't even published to avoid reverse engineering, but it seems that this might still be a problem.
What you'd want actually is a layer that summarizes the actual answer, but that's actually an internal prompt by claude that you are not seeing, the model just doesn't expose the necessary bits for you to hack this together.
Try another model that exposes the raw llm output instead of exposing a CoT result directly.
Of course the real hack is learning to read diagonally without reading every single word, this is a skill that is useful in general. It's also less effort in general, instead of making plugins and super customizing the thing, you just consume the default settings, which are hyperoptimized, and require no time spent in configuration.
The <think> blocks are an attempt to explore the gradient descent space to escape local minimums and find a better global minimum to continue the descent.
While verbosity _might_ do this better, you could easily consider things like "but wait am I forgetting ...." as just one token. So if you actually do it right, you could replace all that with a "hold on" or something of a terse variety.