I’ve been getting this compaction error in a conversation that has now grown to 626k and is beyond my Qwen 3.8 27B model. So I pointed the compaction to Sol 6.0 that has a context of 1M and still get the above.
This was not an issue for versions 0.25.14 and earlier, but I started this conversation in 0.25.15 and this issue is persisting in 0.25.16 and 0.25.17.
Previously, I noticed that capturing the context limit for API endpoints does not get an accurate size from the model as I believe the value pair returns a different value than what Osaurus expects.
For example, I’ll point to OpenAI GPT Sol 6 and it should allow for 1m context, but Osaurus thinks this is unknown and defaults to the “Unknown-Model Metadata Fallback (tokens)” and I’ve left the “Context Window Cap (tokens)” blank.
It would be great if I could use a model in conversation that uses 128k-262k and a compaction model that uses a larger size (1m) to ensure I’m not trimming my conversation, but actually compacting it, so I can keep my conversation going.
