Amazon Bedrock
Writer announces Palmyra X5 LLM with 1M-token context window to power AI agents - SiliconANGLE
Model details
Palmyra X5 is engineered as the runtime for enterprise AI agents that need to process large volumes of domain data, chain multiple tools together, and maintain fast response times. Its hybrid transformer architecture supports a one-million-token context window, enabling teams to feed entire regulatory filings, product catalogs, or internal playbooks into a single call without relying on brittle chunking strategies. The model also delivers sub-second tool-calling and built-in connectors for retrieval-augmented generation and knowledge graphs, positioning it as a foundation for multi-agent orchestration rather than a simple completion engine.
The training approach is notable for its efficiency: Palmyra X5 was developed using synthetic data at a reported compute cost of one million dollars in GPUs, and Writer has emphasized that its models never undergo post-training quantization or distillation. That distinction means the behavior teams validate during evaluation remains intact in production. Benchmarks show the model processing a full million-token prompt in roughly twenty-two seconds and issuing multi-turn function calls in about three hundred milliseconds, while costing three to four times less per token than comparable frontier models. For organizations moving from experimental AI pilots to auditable, agent-driven workflows, that combination of speed, cost efficiency, and behavioral fidelity shapes a practical path toward scalable enterprise deployment.
Transparent token rates
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Amazon Bedrock
Writer announces Palmyra X5 LLM with 1M-token context window to power AI agents - SiliconANGLE