DSPy changelog: what's new each month
Every stable DSPy release summarised by month: the highlights, new features, improvements, fixes and anything you need to act on. 3 months covered; the current month updates daily.
September 2026
1 release: 3.4.0
September brought major updates to DSPy's language model interface, introducing OpenAI-style message calls and a new custom engine system that gives users fine-grained control over model connections and behavior.
- Client configuration (apikey, apibase, timeout) must now be set on custom engine objects, not on LM construction or calls.
- Custom engines require importable classes and JSON-compatible serialization; unsafe loading requires allowunsafelmstate=True flag.
- Native n answers behavior changed: separate sequential requests are now used, potentially affecting billing and retry logic.
Highlights
- OpenAI-style lm(messages=[...]) calls now supported, including SDK message objects.
- Custom engine interface allows you to implement your own BaseLM.forward() and aforward() integrations.
- Native n answers now use separate sequential requests, with explicit control over billing and retry behavior.
- Timeout handling improved with native bounds on individual transport waits rather than total generation time.
- Streaming and caching behavior clarified: streams are not retried after emitting chunks, and incomplete responses are not cached.
New
- OpenAI-style lm(messages=[...]) calls with SDK message objects.
- Custom engine objects that own their own connections and client settings.
- LegacyEngine and AsyncLegacyEngine as transition tools for custom integrations.
- Native n answers via separate sequential requests.
- Stream chunk emission prevents retry behavior.
- Rate-limit error preservation of retry hints, request IDs, and headers.
Improved
- Client settings (apikey, apibase, timeout) now live on engine objects rather than LM construction.
- Timeout behavior now bounds individual transport waits instead of total generation time.
- Incomplete responses are no longer incorrectly cached as successful results.
- Rate-limit errors preserve metadata from the original request for better debugging.
Fixed
- Streams no longer attempt retries after emitting a chunk.
- Incomplete responses no longer cached as successful completions.
- Custom engines can properly save and restore state via JSON-compatible dumpstate() and loadstate() methods.
August 2026
2 releases: 3.3.0 → 3.3.1
August brought major improvements to DSPy's tool calling and message handling architecture, with native multi-turn support and better structured history management. Security and reliability also received attention with sandbox isolation fixes.
Highlights
- Native multi-turn tool call support now replays prior calls and results as structured messages instead of flattening them into text.
- Parallel tool calls are preserved with IDs in both native and non-native modes for clearer tracking.
- Each conversation turn is now stored as structured messages in dspy.History, enabling better prompt caching across stable prefixes.
- Custom LM implementations can now work with a single typed request-response path instead of guessing provider-specific formats.
- Sandbox isolation was significantly strengthened to prevent desynchronization, file collisions, and host-tool identity mutations.
New
- Parallel tool calls support with per-call tracking by ID.
- Multi-turn native tool call support with message replay.
- Structured message-based turn storage in dspy.History.
- Image.fromurl() now downloads resources and returns embedded data URIs.
- Diagnostic event tracking for sandbox-to-host tool calls and interpreter lifecycle.
Improved
- LiteLLM transitions to optional compatibility fallback instead of required core dependency.
- Custom LM authors now implement one typed LMRequest to LMResponse path.
- Custom LMs can translate between DSPy types and their own provider or inference stack.
- Adapters can now depend on DSPy's representation of messages, multimodal content, tool calls, reasoning, citations, and metadata.
- Prompt caching effectiveness improved through stable message prefixes in structured history.
Fixed
- Unsolicited sandbox diagnostics no longer desynchronize JSON-RPC replies.
- Request IDs are now unpredictable to prevent execution manipulation.
- Mounted files with distinct host paths no longer silently collide at the same guest path.
- Guest code cannot change host-tool identity through JavaScript global mutation.
- Sandbox runtime files are protected and Deno cache access is revoked after execution.
May 2026
1 release: 3.2.1
May brought fixes for streaming and caching behavior, along with documentation updates to better support production deployments and community use cases.
Highlights
- Fixed custom headers forwarding in async streaming LM calls to LiteLLM
- Fixed per-call caching=False behavior for both sync and async embedding calls
- Removed upper bound constraint on litellm dependency
- Reorganized documentation to better highlight production use cases and deployment guidance
Improved
- Usecase page updated with clearer guidelines for community contributions
- Production use-cases documentation copy revised
- Deployment moved into technical documentation tabs for better visibility
Fixed
- Async streaming LM calls now correctly forward custom headers to LiteLLM
- Embedder per-call caching=False now honored for sync and async calls
- MkDocs admonition rendering in Deployment and Observability documentation
- Duplicate-word typos in documentation, source code, and tests
Summaries are written automatically from the official release notes (full changelog ↗); check the original notes before relying on a detail. DSPy: pricing, features and alternatives · All changelogs
