Currently listed through these providers:
Model details
Qwen 3.8 27B Fable
Qwen 3.8 27B Fable is presented as a creative finetune aimed at expressive dialogue, long-form storytelling, character work, and roleplay. The NanoGPT listing frames it as an open-weight multimodal model, with a generous 262.1K-token context window and a 32.8K-token maximum output that support sustained narrative and multi-turn roleplay sessions. The 27B parameter scale suggested by the model name positions it as a mid-sized dense transformer built for interactive, persona-rich writing rather than compact single-turn tasks. A second provider row indicates a Korea-region routing option with lower output pricing and slightly better latency, giving deployers a choice between the default auto-routing tier and an alternative regional endpoint.
Because the model has no published benchmark results on the NanoGPT page, practical evaluation rests on observed provider metrics and use-case fit. Under auto routing, measured latency sits around 8.7 seconds with throughput near 61 tokens per second, while the alternative regional endpoint shows roughly 7.4-second latency and about 45 tokens per second, useful numbers for sizing interactive storytelling experiences. The combination of open weights, a large context window, multimodal input, and a finetune geared toward character voice makes this variant a strong match for worldbuilding, scripted interactive fiction, and assistant personas that need to hold tone across very long conversations. It is less suited to workflows that require published benchmark evidence or strict latency budgets below the observed single-digit-second response times.
Quick Info
Powered by- Provider
- NanoGPT
- Model key
- qwen/qwen3.8-27b-fable
- Release date
- Jul 29, 2026
- Last updated
- Aug 28, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.25
- Output token cost
- $1.50
Limits
- Input tokens
- 262,144 tokens
- Output tokens
- 32,768 tokens
- Context window
- 262,144 tokens