Which area does this feature relate to?
Building Blocks (existing)
Describe the feature you'd like to request
Currently the bb-agent wrapper over Strands does not support prompt caching. Especially for long-running agents with a lot of context this leads to increased bills since users pay the uncached prices for their models (i.e. $0.30 vs $3.00 per 1M input tokens for cache reads vs. uncached input with Claude Sonnet 4.6).
Strands supports prompt caching on Bedrock (https://strandsagents.com/docs/user-guide/concepts/model-providers/amazon-bedrock/#caching) - but the existing building block does not expose this as a configurable option.
Use case
Agents with long system prompts, many tools, or multi-turn conversations re-send a large stable prefix on every request. Without prompt caching that prefix is billed at full input price each time. On Claude Sonnet 4.6 that's $3.00/1M instead of $0.30/1M for a cache read (10x), plus higher latency.
Strands already supports Bedrock prompt caching on its BedrockModel, but bb-agent's ModelConfig doesn't expose it, so there's no easy way to turn it on from an AWS Blocks app without modifying the local code after installing the package.
Proposed solution
Add an optional cacheConfig field to bb-agent's ModelConfig (bedrock only) and pass it into the Strands BedrockModel. Off by default, so no breaking change.
export interface ModelConfig {
// ...existing fields
/** Prompt caching for the bedrock provider. Optional so off by default. */
cacheConfig?: { strategy: 'auto' | 'anthropic' };
}
In createStrandsModel's bedrock branch:
return new BedrockModel({
modelId: config.modelId,
...(config.inferenceConfig && { /* ... */ }),
...(config.cacheConfig && { cacheConfig: config.cacheConfig }),
});
Usage:
model: { deployed: { ...BedrockModels.BALANCED, cacheConfig: { strategy: 'auto' } } }
'auto' lets Strands place cache points for known Bedrock model IDs (incl. the BedrockModels presets); 'anthropic' force-enables it for ARN application inference profiles that 'auto' can't detect. This mirrors the implementation in the strands typescript SDK. Ignored by the openai-api and canned providers.
Reference: Strands' own caching support — https://strandsagents.com/docs/user-guide/concepts/model-providers/amazon-bedrock/#caching
Alternatives considered
The feature can be enabled by manually/with an agent modifying the source code in bb-agent package, but this is messy and doesn't survive updates.
Additional context
This script runs a two-turn conversation over a long, stable system prompt and prints the cache token counts. Because bb-agent has no way to enable caching today, cacheWrite/cacheRead stay 0 on every turn — the whole prefix is billed at full input price each turn. It builds the model through bb-agent's own createStrandsModel to reflect what is currently in AWS Blocks without the overhead of provisioning a full AWS Blocks app.
import { Agent } from '@strands-agents/sdk';
// Assumes this script is run from the root of the repo.
import { createStrandsModel } from './packages/bb-agent/dist/model-factory.js';
process.env.AWS_REGION ??= 'us-east-1';
const MODEL_ID = 'global.anthropic.claude-sonnet-4-6';
// Long, stable prefix so it clears the ~1,024-token cache minimum and is identical across turns
const SYSTEM_PROMPT = [
'You are "Atlas", a meticulous senior support engineer for the ACME Cloud platform.',
'Follow these operating rules on every response, without exception:',
...Array.from({ length: 40 }, (_, i) =>
`Rule ${i + 1}: Be precise, cite the relevant ACME Cloud service by name, prefer the ` +
'least-privilege remediation, never invent API names, and always end with a one-line summary.'),
].join('\n');
const TURNS = [
'A user reports 503s from EdgeCDN after a deploy. First triage step?',
'Now they also see elevated latency on ObjectStore. What next?',
];
// Current implementation: ModelConfig has no way to enable caching.
const model = await createStrandsModel({ provider: 'bedrock', modelId: MODEL_ID });
const agent = new Agent({ model, systemPrompt: SYSTEM_PROMPT, printer: false });
for (let i = 0; i < TURNS.length; i++) {
const u = (await agent.invoke(TURNS[i])).metrics?.accumulatedUsage ?? {};
console.log(`turn ${i + 1}: input=${u.inputTokens ?? 0} ` +
`cacheWrite=${u.cacheWriteInputTokens ?? 0} cacheRead=${u.cacheReadInputTokens ?? 0}`);
}
Expected output: cacheWrite=0 cacheRead=0 on both turns, and input on turn 2 still includes the full system prompt.
Running this against an AWS account with DEFAULT_REGION=us-east-1 gives the following output:
$ node prompt-caching-issue.mjs
turn 1: input=1780 cacheWrite=0 cacheRead=0
turn 2: input=3969 cacheWrite=0 cacheRead=0
Is this something you'd be interested in working on?
Which area does this feature relate to?
Building Blocks (existing)
Describe the feature you'd like to request
Currently the bb-agent wrapper over Strands does not support prompt caching. Especially for long-running agents with a lot of context this leads to increased bills since users pay the uncached prices for their models (i.e. $0.30 vs $3.00 per 1M input tokens for cache reads vs. uncached input with Claude Sonnet 4.6).
Strands supports prompt caching on Bedrock (https://strandsagents.com/docs/user-guide/concepts/model-providers/amazon-bedrock/#caching) - but the existing building block does not expose this as a configurable option.
Use case
Agents with long system prompts, many tools, or multi-turn conversations re-send a large stable prefix on every request. Without prompt caching that prefix is billed at full input price each time. On Claude Sonnet 4.6 that's $3.00/1M instead of $0.30/1M for a cache read (10x), plus higher latency.
Strands already supports Bedrock prompt caching on its
BedrockModel, but bb-agent'sModelConfigdoesn't expose it, so there's no easy way to turn it on from an AWS Blocks app without modifying the local code after installing the package.Proposed solution
Add an optional
cacheConfigfield to bb-agent'sModelConfig(bedrock only) and pass it into the StrandsBedrockModel. Off by default, so no breaking change.In
createStrandsModel's bedrock branch:Usage:
'auto'lets Strands place cache points for known Bedrock model IDs (incl. theBedrockModelspresets);'anthropic'force-enables it for ARN application inference profiles that'auto'can't detect. This mirrors the implementation in the strands typescript SDK. Ignored by theopenai-apiandcannedproviders.Reference: Strands' own caching support — https://strandsagents.com/docs/user-guide/concepts/model-providers/amazon-bedrock/#caching
Alternatives considered
The feature can be enabled by manually/with an agent modifying the source code in
bb-agentpackage, but this is messy and doesn't survive updates.Additional context
This script runs a two-turn conversation over a long, stable system prompt and prints the cache token counts. Because bb-agent has no way to enable caching today,
cacheWrite/cacheReadstay 0 on every turn — the whole prefix is billed at full input price each turn. It builds the model through bb-agent's owncreateStrandsModelto reflect what is currently in AWS Blocks without the overhead of provisioning a full AWS Blocks app.Expected output:
cacheWrite=0 cacheRead=0on both turns, andinputon turn 2 still includes the full system prompt.Running this against an AWS account with
DEFAULT_REGION=us-east-1gives the following output:Is this something you'd be interested in working on?