# DBOS Documentation This file contains all documentation content in a single document following the llmstxt.org standard. ## AI Quickstart You can integrate DBOS durable workflows with your AI agents (or other AI systems) to make them reliable, observable, and resilient to failures. Rather than bolting on ad-hoc retry logic, DBOS workflows give you one consistent model for ensuring your agents can recover from any failure from exactly where they left off. In particular, integrating DBOS to your agents gives you: - [**Resilience to failure**](../python/tutorials/workflow-tutorial.md): Automatically recover your agents from server restarts, process crashes, network hiccups or outages, and other unexpected events. - [**Observability and reproducibility**](./debugging.md): Monitor your agentic workflows in real time. If they exhibit unexpected behavior, use saved workflow progress to reproduce it in a development environment to identify and fix the root cause. - [**Support for long-running agents and human-in-the-loop**](./hitl.md): Build agents that run for hours, days, or weeks (potentially waiting for human responses) and seamlessly recover from any interruption. - [**Durable streaming**](./streaming.md): Stream output from your agents as it's generated to build interactive or conversational flows that recover from any failure. - [**Parallel, scalable, and distributed agents**](./distributing-agents.md): Use durable queues to build agents with parallel tool calls or tasks, potentially distributing it across many servers with managed flow control. ### Get Started You can integrate DBOS into an agent built in regular Python or TypeScript, or use native integrations with popular agentic frameworks like [Pydantic AI](https://ai.pydantic.dev/durable_execution/dbos), [LlamaIndex](https://developers.llamaindex.ai/python/llamaagents/workflows/dbos/), [OpenAI Agents SDK](https://openai.github.io/openai-agents-python/running_agents/#dbos), [Google ADK](https://adk.dev/integrations/dbos/), and the [Vercel AI SDK](https://ai-sdk.dev/). #### 1. Install DBOS `pip install` DBOS into your application. ```shell pip install dbos ``` #### 2. Configure and Launch DBOS Add these lines of code to your agent's main function. They initialize DBOS when your agentic application starts. ```python import os from dbos import DBOS, DBOSConfig config: DBOSConfig = { "name": "my-app", "application_version": "0.1.0", "system_database_url": os.environ.get("DBOS_SYSTEM_DATABASE_URL"), } DBOS(config=config) DBOS.launch() ``` :::info DBOS uses a database to durably store workflow and step state. By default, it uses SQLite, which requires no configuration. For production use, we recommend connecting your DBOS application to a Postgres database. When you're ready for production, you can connect this initialization code to Postgres by setting the `DBOS_SYSTEM_DATABASE_URL` environment variable to a connection string to your Postgres database. ::: #### 3. Annotate Workflows and Steps Next, annotate your main agentic loop as a durable workflow and each LLM and tool call it makes as a step. This causes DBOS to checkpoint the progress of your agent in your database so it can recover from any failure. For instance, in the [deep research agent example](../python/examples/hacker-news-agent.md), here is the main agentic loop: ```python @DBOS.workflow() def agentic_research_workflow(topic: str, max_iterations: int = 3): """ This agent starts with a research topic then: 1. Searches Hacker News for information on that topic. 2. Iteratively searches related topics, collecting information. 3. Makes decisions about when to continue. 4. Synthesizes findings into a final report. """ ... ``` And here is an example step, an LLM call to evaluate results: ```python @DBOS.step() def evaluate_results_step( topic: str, query: str, stories: List[Dict[str, Any]], comments: Optional[List[Dict[str, Any]]] = None, ) -> EvaluationResult: """LLM evaluates search results and extracts insights.""" ... ``` To learn more about how to build with DBOS Python, check out the [Python docs](../python/programming-guide.md). #### 1. Install DBOS `npm install` DBOS into your application. ```shell npm install @dbos-inc/dbos-sdk@latest ``` #### 2. Configure and Launch DBOS Add these lines of code to your agent's main function. They initialize DBOS when your agentic application starts. ```typescript import { DBOS } from "@dbos-inc/dbos-sdk"; DBOS.setConfig({ "name": "my-app", "applicationVersion": "0.1.0", "systemDatabaseUrl": process.env.DBOS_SYSTEM_DATABASE_URL, }); await DBOS.launch(); ``` :::info DBOS uses a database to durably store workflow and step state. By default, it uses a Postgres database. You can start Postgres locally with `npx dbos postgres start`, or set the `DBOS_SYSTEM_DATABASE_URL` environment variable to a connection string to an existing Postgres database. ::: #### 3. Register Workflows and Steps Next, register your main agentic loop as a durable workflow and run each LLM and tool call as a step. This causes DBOS to checkpoint the progress of your agent in your database so it can recover from any failure. For instance, in the [deep research agent example](../typescript/examples/hacker-news-agent.md), here is the main agentic loop, registered as a workflow: ```typescript async function agenticResearchWorkflowFunction( topic: string, maxIterations: number, ): Promise { ... } export const agenticResearchWorkflow = DBOS.registerWorkflow( agenticResearchWorkflowFunction, ); ``` And here is an example step, an LLM call to evaluate results: ```typescript const evaluation = await DBOS.runStep( () => evaluateResults(topic, query, stories, comments), { name: "evaluateResults" }, ); ``` To learn more about how to build with DBOS TypeScript, check out the [TypeScript docs](../typescript/programming-guide.md). #### 1. Install Pydantic AI with DBOS Install Pydantic AI with the DBOS optional dependency. ```shell pip install pydantic-ai[dbos] ``` #### 2. Configure DBOS and Wrap Your Agent Import and configure DBOS, then wrap your Pydantic AI agent in a `DBOSAgent` for durable execution. `DBOSAgent` automatically wraps your agent's run loop as a DBOS workflow and model requests and MCP communication as DBOS steps. ```python import asyncio # highlight-next-line from dbos import DBOS, DBOSConfig from pydantic_ai import Agent # highlight-next-line from pydantic_ai.durable_exec.dbos import DBOSAgent # highlight-start dbos_config: DBOSConfig = { 'name': 'pydantic_dbos_agent', 'application_version': '0.1.0', 'system_database_url': 'sqlite:///dbostest.sqlite', } DBOS(config=dbos_config) #highlight-end agent = Agent( 'gpt-5', instructions="You're an expert in geography.", name='geography', ) # highlight-next-line dbos_agent = DBOSAgent(agent) async def main(): # highlight-next-line DBOS.launch() result = await dbos_agent.run('What is the capital of Mexico?') print(result.output) if __name__ == "__main__": asyncio.run(main()) ``` Custom tool functions can optionally be decorated with `@DBOS.step` if they involve non-determinism or I/O. To learn more, check out the [Pydantic AI integration guide](../integrations/pydantic-ai.md) and the [Pydantic AI docs](https://ai.pydantic.dev/durable_execution/dbos). #### 1. Install LlamaIndex with DBOS Install the [`llama-agents-dbos`](https://github.com/run-llama/workflows-py/tree/main/packages/llama-agents-dbos) package. ```shell pip install llama-agents-dbos ``` #### 2. Configure DBOS and Use the DBOS Runtime Import and configure DBOS, then create a `DBOSRuntime` and pass it to your LlamaIndex workflow. The DBOS runtime automatically persists every workflow transition so your workflow can resume exactly where it left off after any failure. ```python import asyncio # highlight-next-line from dbos import DBOS, DBOSConfig # highlight-next-line from llama_agents.dbos import DBOSRuntime from pydantic import Field from workflows import Context, Workflow, step from workflows.events import Event, StartEvent, StopEvent # highlight-start config: DBOSConfig = { "name": "llamaindex-example", "application_version": "0.1.0", "system_database_url": "sqlite:///example.sqlite", } DBOS(config=config) # highlight-end class MyResult(StopEvent): output: str = Field(description="Result") class MyWorkflow(Workflow): @step async def start(self, ctx: Context, ev: StartEvent) -> MyResult: return MyResult(output="Hello from a durable workflow!") # highlight-next-line runtime = DBOSRuntime() workflow = MyWorkflow(runtime=runtime) async def main() -> None: # highlight-next-line await runtime.launch() result = await workflow.run(run_id="my-run-1") print(result.output) asyncio.run(main()) ``` To learn more, check out the [LlamaIndex integration guide](../integrations/llamaindex.md) and the [LlamaIndex docs](https://developers.llamaindex.ai/python/llamaagents/workflows/dbos/). #### 1. Install DBOS and the OpenAI Agents Integration Install DBOS and the [durable OpenAI agents integration](https://github.com/dbos-inc/dbos-openai-agents). ```shell pip install dbos dbos-openai-agents ``` #### 2. Configure DBOS and Wrap Your Agent Use `DBOSRunner` as a drop-in replacement for `Runner` and annotate your agent's workflow and tool calls with DBOS decorators. ```python import asyncio from agents import Agent, function_tool # highlight-start from dbos import DBOS, DBOSConfig from dbos_openai_agents import DBOSRunner #highlight-end @function_tool # highlight-next-line @DBOS.step() async def get_weather(city: str) -> str: """Get the weather for a city.""" return f"Sunny in {city}" agent = Agent(name="weather", tools=[get_weather]) # highlight-start @DBOS.workflow() async def run_agent(user_input: str) -> str: result = await DBOSRunner.run(agent, user_input) return str(result.final_output) # highlight-end async def main(): # highlight-start config: DBOSConfig = { "name": "my-agent", "application_version": "0.1.0", "system_database_url": 'sqlite:///my_agent.sqlite', } DBOS(config=config) DBOS.launch() # highlight-end output = await run_agent("How is the weather in San Francisco") print(output) if __name__ == "__main__": asyncio.run(main()) ``` To learn more, check out the [OpenAI Agents SDK integration guide](../integrations/openai-agents.md) and the [OpenAI Agents SDK documentation](https://openai.github.io/openai-agents-python/running_agents/#dbos). #### 1. Install DBOS and the Google ADK Integration Install DBOS and the [durable Google ADK agents integration](https://github.com/dbos-inc/dbos-google-adk). ```shell pip install dbos dbos-google-adk ``` #### 2. Configure DBOS and Wrap Your Agent Add `DBOSPlugin` to your `Runner` and annotate your agent's workflow and tool calls with DBOS decorators. ```python import asyncio import logging # highlight-start from dbos import DBOS, DBOSConfig from dbos_google_adk import DBOSPlugin # highlight-end from google.adk.agents import LlmAgent from google.adk.runners import Runner from google.adk.sessions import InMemorySessionService from google.genai import types # Decorate tool calls with @DBOS.step() for durable execution # highlight-next-line @DBOS.step() async def get_weather(city: str) -> str: """Get the weather for a city.""" return f"Sunny in {city}" agent = LlmAgent(name="weather", model="gemini-flash-latest", tools=[get_weather]) runner = Runner( app_name="my-agent", agent=agent, # highlight-next-line plugins=[DBOSPlugin()], session_service=InMemorySessionService(), ) # Drive the agent from a DBOS workflow for durable execution # highlight-next-line @DBOS.workflow() async def run_agent(user_id: str, session_id: str, message: str) -> str: new_message = types.Content(role="user", parts=[types.Part.from_text(text=message)]) async for event in runner.run_async( user_id=user_id, session_id=session_id, new_message=new_message ): if event.is_final_response(): return event.content.parts[0].text return "" async def main(): # highlight-start # DBOS checkpoints to SQLite by default. Postgres is recommended for production. config: DBOSConfig = {"name": "my-agent", "application_version": "0.1.0", "system_database_url": "sqlite:///dbostest.sqlite"} DBOS(config=config) DBOS.launch() # highlight-end await runner.session_service.create_session( app_name="my-agent", user_id="u", session_id="s" ) print(await run_agent("u", "s", "How is the weather in San Francisco?")) if __name__ == "__main__": asyncio.run(main()) ``` To learn more, check out the [Google ADK integration guide](../integrations/google-adk.md) and the [Google ADK documentation](https://adk.dev/integrations/dbos). #### 1. Install DBOS and the Vercel AI Integration Install DBOS and the [Vercel AI SDK integration](https://www.npmjs.com/package/@dbos-inc/vercel-ai). ```shell npm install @dbos-inc/vercel-ai @dbos-inc/dbos-sdk ai ``` #### 2. Wrap Your Model and Run Your Agent in a Workflow Wrap your model with `durableCalls` middleware so every model call is checkpointed, then register your agent's generation loop as a DBOS workflow. DBOS checkpoints the progress of your agent in your database so it can recover from any failure, replaying completed model calls from their checkpoints instead of re-contacting the provider. ```typescript // highlight-start import { DBOS } from '@dbos-inc/dbos-sdk'; import { durableCalls } from '@dbos-inc/vercel-ai'; // highlight-end import { generateText, wrapLanguageModel } from 'ai'; import { openai } from '@ai-sdk/openai'; // Wrap your model so every model call is checkpointed in Postgres // highlight-next-line const model = wrapLanguageModel({ model: openai('gpt-5'), // highlight-next-line middleware: durableCalls({ retriesAllowed: true, maxAttempts: 5 }), }); // Register your agent's generation loop as a durable workflow // highlight-next-line const researchAgent = DBOS.registerWorkflow( async (question: string) => { const { text } = await generateText({ model, prompt: question, system: 'You are a helpful research assistant.', }); return text; }, { name: 'researchAgent' }, ); async function main() { // highlight-start DBOS.setConfig({ name: 'my-agent', systemDatabaseUrl: process.env.DBOS_SYSTEM_DATABASE_URL }); await DBOS.launch(); // highlight-end console.log(await researchAgent('Why did the agent cross the road?')); } main(); ``` Wrap your tools with [`durableTools`](../integrations/vercel-ai.md#durable-tools) to make their side effects durable too. The integration also supports [durable streams](../integrations/vercel-ai.md#durable-streams), [MCP tools](../integrations/vercel-ai.md#durable-mcp-tools), [subagents](../integrations/vercel-ai.md#durable-subagents), [embeddings](../integrations/vercel-ai.md#durable-embedding-models), and [image generation](../integrations/vercel-ai.md#durable-image-models). To learn more, check out the [Vercel AI SDK integration guide](../integrations/vercel-ai.md) and the [Vercel AI SDK documentation](https://ai-sdk.dev/). ### Using Coding Agents DBOS provides skills and prompts to help you use coding agents to add DBOS to your AI applications. Learn more about them here: - [AI-assisted development in Python](../python/prompting.md) - [AI-assisted development in TypeScript](../typescript/prompting.md) - [AI-assisted development in Go](../golang/prompting.md) - [AI-assisted development in Java](../java/prompting.md) Additionally, DBOS provides an MCP server so your agents can observe and monitor your workflows and help you find and catch issues. Learn more about it [here](../integrations/mcp.md). --- ## Observability & Reproducibility One of the most common problems you encounter building and operating agents is **debugging failures**, particularly those caused by unexpected agent behavior. For example, an agent might: - Return a malformed structured output, causing a tool call to fail. - Invoke the wrong tool or the right tool with the wrong inputs, causing the tool to fail. - Generate an undesirable or inappropriate text output, with potentially business-critical consequences. These behaviors are especially hard to diagnose in a complex or long-running agent—if an agent runs for two hours then fails unexpectedly, it's difficult to reproduce the exact set of conditions that caused the failure and test a fix. Durable workflows help by making it easier to **observe** the root cause of the failure, deterministically **reproduce** the failure, and **test or apply** fixes. Because workflows checkpoint the outcome of each step of your workflow, you can review these checkpoints to see the cause of the failure and audit every step that led to it. For example, using the [DBOS Console dashboard](../conductor/workflow-management.md), you might see that your agent failed because of a validation error caused by a malformed structured output: Once you've identified the cause of a failure, you can use the [**workflow fork**](../python/tutorials/workflow-management.md#forking-workflows) operation to reproduce it. Fork restarts a workflow from a completed step, using checkpointed information to deterministically reproduce the state of the workflow up to that step. Thus, you can rerun the misbehaving step under the exact conditions that originally caused the misbehavior. Once you can reproduce a failure in a development environment, it becomes much easier to fix. You can add additional logging or telemetry to the misbehaving step to identify the root cause. Then, when you have a fix, you can reproduce the failure with the fix in place to test if it works. For example, if you hypothesize that the malformed output was caused by an error in the prompt, you can fix the prompt, rerun the failed step, and watch it complete successfully: --- ## Parallelizing & Scaling Agents AI agents and applications often need to **run many tasks in parallel**. A single step of an agentic loop might invoke several tools at once based on an LLM response. A document ingestion pipeline using Retrieval-Augmented Generation (RAG) might index tens of thousands of documents concurrently. A deep research agent might scrape hundreds of websites at the same time. DBOS workflows make these parallel patterns durable and scalable through a **durable queue** abstraction. A workflow can enqueue any number of tasks for concurrent processing, then wait for their results. Because every task is checkpointed, your agent can recover from any failure mid-flight without re-running work that already succeeded. ### Parallel Tool Calls When an LLM returns multiple tool calls in a single response, you can execute them in parallel by enqueuing each one as a workflow and waiting for all of them to complete: ```python DBOS.register_queue("tool_queue") @DBOS.workflow() def run_tool_calls(tool_calls): handles: List[WorkflowHandle] = [] # Enqueue each tool call to run in parallel for call in tool_calls: handle = DBOS.enqueue_workflow("tool_queue", execute_tool, call.name, call.arguments) handles.append(handle) # Wait for all tool calls to finish and collect their outputs return [handle.get_result() for handle in handles] ``` Because each tool call runs as a durable workflow, an agent that crashes partway through a fan-out resumes without re-issuing tools that already completed—an important property when tools have side effects or call expensive APIs. ### Distributing Work Across Servers The same pattern scales to data pipelines and other batch workloads. For example, a document ingestion pipeline can enqueue a workflow to index each document in a batch: ```python DBOS.register_queue("indexing_queue") @DBOS.workflow() def index_documents(urls): handles: List[WorkflowHandle] = [] # Enqueue each document for indexing for url in urls: handle = DBOS.enqueue_workflow("indexing_queue", index_document, url) handles.append(handle) # Wait for all documents to finish indexing, count the total number of indexed pages outputs = [] for handle in handles: outputs.append(handle.get_result()) return outputs ``` Enqueued workflows can be dequeued and executed by any of your application's servers, distributing the work across your fleet. If your application is resource intensive or uses rate-limited APIs, you can use queues to rate-limit or control the concurrency of your workflows. For example, you can specify that no more than 10 workflows should run concurrently on a single server: ```python DBOS.register_queue("indexing_queue", worker_concurrency=10) ``` Because queues are backed by durable workflows, they automatically recover from any failure: if a server restarts or has a network hiccup partway through a multi-hour run of your pipeline on a batch of 10K documents, your pipeline will recover from the last indexed document instead of restarting from the beginning and redoing expensive work. If you're interested in building distributed AI agents or data pipelines, check out the [document ingestion example](../python/examples/document-detective.md), which shows best practices for building durable distributed applications. To learn more about how to scale applications with durable queues, check out the [queues tutorial](../python/tutorials/queue-tutorial.md). --- ## Reliable Human-in-the-Loop Many agents need a **human in the loop** for decisions that are too important to fully trust an LLM. However, it's not easy to design an agent that waits for human feedback. The key issue is **time**: a human might take hours or days to respond to an agent, so the agent must be able to reliably wait for a long time (during which the server might be restarted, software might be upgraded, etc.). Durable workflows help because they provide tools like **durable messaging** and **workflow events** that let agents durably communicate with the outside world. ### Waiting for Human Input You can use [`DBOS.recv`](../python/tutorials/workflow-communication.md#recv) inside a workflow to durably wait for a message. For example, you can add a line of code to your agent that tells it to wait hours or days for a notification: ```python approval: Optional[HumanResponseRequest] = DBOS.recv(timeout_seconds=TIMEOUT) ``` Because the workflow's progress is checkpointed and both the deadline and notification are stored in your database, this can safely wait for a long time. Anything can happen while your agent is waiting (its server can restart, its code can be upgraded, etc.) and it will recover and keep waiting until the notification arrives or the deadline is reached. To send that notification (for example, from an HTTP endpoint), use [`DBOS.send`](../python/tutorials/workflow-communication.md#send): ```python @app.post("/agents/{agent_id}/respond") def respond_to_agent(agent_id: str, response: HumanResponseRequest): DBOS.send(agent_id, response) return {"ok": True} ``` All messages are persisted to the database, so if `send` completes successfully, the destination workflow is guaranteed to receive it. ### Publishing Agent Status Agents can publish their current status using [`DBOS.set_event`](../python/tutorials/workflow-communication.md#workflow-events), and external code can read it with [`DBOS.get_event`](../python/tutorials/workflow-communication.md#get_event). This is useful for letting a frontend know what an agent is doing, whether it's working, waiting for approval, or finished: ```python @DBOS.workflow() def durable_agent(request: AgentStartRequest): agent_status = AgentStatus(status="working", ...) DBOS.set_event(AGENT_STATUS, agent_status) # Do some work, then request approval agent_status.status = "pending_approval" DBOS.set_event(AGENT_STATUS, agent_status) approval = DBOS.recv(timeout_seconds=TIMEOUT) if approval is not None and approval.response == "approve": agent_status.status = "working" DBOS.set_event(AGENT_STATUS, agent_status) # Continue execution... else: agent_status.status = "denied" DBOS.set_event(AGENT_STATUS, agent_status) raise Exception("Agent denied or timed out") ``` You can then use the [workflow introspection API](../python/reference/contexts.md#list_workflows) and [`DBOS.get_event`](../python/tutorials/workflow-communication.md#get_event) to monitor and display your agents, for example to build an "inbox" of all agents currently waiting for human input: ```python @app.get("/agents/waiting") async def list_waiting_agents(): agent_workflows = await DBOS.list_workflows_async( status="PENDING", name=durable_agent.__qualname__ ) statuses = await asyncio.gather( *[DBOS.get_event_async(w.workflow_id, AGENT_STATUS) for w in agent_workflows] ) return [s for s in statuses if s.status == "pending_approval"] ``` For a complete working example, check out the [agent inbox application](../python/examples/agent-inbox.md). To learn more about `send`/`recv`, `set_event`/`get_event`, and streaming, see the [workflow communication docs](../python/tutorials/workflow-communication.md). --- ## Streaming Responses AI agents often need to **stream output to clients in real time**, for example, to display LLM output as it is generated, surface intermediate tool results, or report the progress of a long-running task. DBOS workflows provide **durable streams**: append-only channels you can write to from inside a workflow and read from anywhere in your application. Every write is persisted, so if a server restarts mid-response the workflow recovers from where it left off and the reader keeps receiving values without dropping output. ### Writing to a Stream Inside a workflow or step, write values to a stream identified by a string key. When you're done producing values, close the stream so readers know they've received everything; otherwise streams are automatically closed when the workflow terminates. **Python** This example streams an LLM response as it's generated: ```python from openai import OpenAI client = OpenAI() @DBOS.step() def stream_completion(prompt: str, stream_key: str) -> str: full_response = "" response = client.chat.completions.create( model="gpt-5", messages=[{"role": "user", "content": prompt}], stream=True, ) for chunk in response: token = chunk.choices[0].delta.content if token: DBOS.write_stream(stream_key, token) full_response += token return full_response @DBOS.workflow() def chat_workflow(prompt: str) -> str: answer = stream_completion(prompt, "tokens") DBOS.close_stream("tokens") return answer ``` **TypeScript** This example streams an LLM response as it's generated: ```typescript import OpenAI from "openai"; const client = new OpenAI(); async function streamCompletion( prompt: string, streamKey: string, ): Promise { let fullResponse = ""; const response = await client.chat.completions.create({ model: "gpt-5", messages: [{ role: "user", content: prompt }], stream: true, }); for await (const chunk of response) { const token = chunk.choices[0].delta.content; if (token) { await DBOS.writeStream(streamKey, token); fullResponse += token; } } return fullResponse; } async function chatWorkflowFunction(prompt: string): Promise { const answer = await DBOS.runStep( () => streamCompletion(prompt, "tokens"), { name: "streamCompletion" }, ); await DBOS.closeStream("tokens"); return answer; } export const chatWorkflow = DBOS.registerWorkflow(chatWorkflowFunction); ``` ### Reading from a Stream You can read from a stream using its workflow ID and key from anywhere in your application. The reader yields values in order until the stream is closed or the workflow terminates. For example, start an agentic workflow and print its output as it's written: **Python** ```python handle = DBOS.start_workflow(chat_workflow, "Tell me a joke") for token in DBOS.read_stream(handle.workflow_id, "tokens"): print(token, end="", flush=True) ``` **TypeScript** ```typescript const handle = await DBOS.startWorkflow(chatWorkflow)("Tell me a joke"); for await (const token of DBOS.readStream(handle.workflowID, "tokens")) { process.stdout.write(token); } ``` You can also read streams from outside your application using a [DBOS Client](../python/reference/client.md#read_stream). To learn more, see the workflow streaming tutorial ([Python](../python/tutorials/workflow-communication.md#workflow-streaming), [TypeScript](../typescript/tutorials/workflow-communication.md#workflow-streaming)). --- ## DBOS Architecture DBOS provides a high-performance, easy-to-use library for durable workflows built on top of Postgres. You use DBOS by installing the open-source library into your application and annotating workflows and steps. While your application runs, DBOS checkpoints those workflows and steps to a Postgres database. When failures occur, whether from crashes, interruptions, or restarts, DBOS uses those checkpoints to recover each of your workflows from the last completed step. Architecturally, an application built with DBOS looks like the below diagram. The open-source DBOS library uses Postgres to orchestrate durable workflows and queues. There's no separate orchestration server and no infrastructure required besides Postgres. When running in production, we also recommend connecting your DBOS applications to [Conductor](#operating-dbos-in-production-with-conductor), a "control plane" for your durable workflows that coordinates workflow recovery to guarantee high availability and provides operational tooling such as an admin UI and dashboard, observability integrations, and managed workflow retention policies. To learn more about how to add DBOS to your application, check out the language-specific integration guides ([Python](./python/integrating-dbos.md), [TypeScript](./typescript/integrating-dbos.md), [Go](./golang/integrating-dbos.md), [Java](./java/integrating-dbos.md)). ### Using DBOS in a Distributed Setting You can create a distributed DBOS application by launching multiple server processes (sometimes called "workers" or "executors") on a variety of platforms, such as a Kubernetes cluster, a fleet of EC2 instances, or a serverless platform like Google Cloud Run. Within an application, each server must connect to the same Postgres database, called the system database. This database stores all workflow checkpoints, step outputs, and schedule and queue state. To distribute work across many servers in a cluster, you should use [durable queues](#durable-queues). Distributed applications should also connect to [DBOS Conductor](#operating-dbos-in-production-with-conductor), the control plane for cluster-wide observability and management. For example, if one of your workers crashes or fails, Conductor detects the failure and automatically recovers its workflows to a compatible live worker. When using DBOS in a distributed setting, you often want to implement durable workflows in one service, but manage them from another service. For example, you may want your API server to enqueue and monitor durable jobs on your data processing service. You can use the DBOS Client ([Python](./python/reference/client.md), [TypeScript](./typescript/reference/client.md), [Go](./golang/reference/dbos-context.md#newclient), [Java](./java/reference/client.md)) to programmatically interact with your application from external code. Your API server can create a client connected to your data processing service's system database and use it to enqueue a job, monitor the job's status, and retrieve its result when complete. Here's a diagram of what that might look like: You may also have multiple applications or services that need durable workflows. For example, you might have a service that runs business workflows, a service that handles data ingestion, and a service that runs an AI agent. You can separately add DBOS to each of these applications. Each application must have a unique name. You can give each application its own system database (this doesn't require multiple Postgres servers: a single physical Postgres server can host multiple logical system databases), or multiple applications can [share a single system database](./explanations/sharing-a-system-database.md). Applications sharing a system database are isolated from one another, each running only its own workflows, queues, and schedules, but can interoperate: for example, one application can enqueue another's workflows and wait for their results. Within an application, all servers must use the same programming language. However, cross-language interaction is possible via the DBOS Client, or by sharing a system database between applications in different languages. For example, a TypeScript application can enqueue workflows onto a separate Python application, monitor their progress, and gather results. Cross-language operations are documented [here](./explanations/portable-workflows.md). ### How DBOS Scales You can easily scale a DBOS application by adding more servers to it, so the scalability of DBOS is fundamentally determined by the database it is connected to. The only overhead DBOS adds is database writes: one database write per step (to checkpoint the step's outcome) plus two additional database writes per workflow (one at the beginning to checkpoint workflow inputs, one at the end to checkpoint the workflow outcome). In [benchmarks](https://www.dbos.dev/blog/benchmarking-workflow-execution-scalability-on-postgres), a DBOS application using a single Postgres database can sustain a throughput of >40K workflows or steps per second. Scaling beyond that is possible by sharding workflows across multiple Postgres databases. It is worth noting that since DBOS checkpoints workflow inputs and outputs and step outputs, the sizes of its writes are determined by the sizes of your inputs and outputs. If your steps return small objects, the write sizes are negligible, but if they return large files, the write sizes are large. Thus, we recommend architecting steps to avoid large output sizes (for example, store large files in cloud blob storage like S3 and have steps return pointers to those files). ### How Workflow Recovery Works DBOS achieves fault tolerance by checkpointing workflows and steps. Every workflow input and step output is durably stored in the system database. When workflow execution fails, whether from crashes, network issues, or server restarts, DBOS leverages these checkpoints to recover workflows from their last completed step. Workflow recovery occurs in three steps: 1. First, DBOS detects interrupted workflows. In single-node deployments, this happens automatically at startup when DBOS scans for incomplete (PENDING) workflows. In a distributed deployment, some coordination is required, either automatically through services like [DBOS Conductor](#operating-dbos-in-production-with-conductor) or [manually](./production/workflow-recovery.md). 2. Next, DBOS restarts each interrupted workflow by calling it with its checkpointed inputs. As the workflow re-executes, it checks before each step if that step's output is checkpointed in Postgres. If there is a checkpoint, the step returns the checkpointed output instead of executing. 3. Eventually, the recovered workflow reaches a step with **no checkpoint**. This marks the point where the original execution failed. The recovered workflow executes that step normally and proceeds from there, thus **resuming from the last completed step.** For DBOS to be able to safely recover a workflow, your code must satisfy two requirements: 1. The workflow function must be **deterministic**: if executed multiple times, with the same arguments and step return values, the workflow should invoke the same steps with the same inputs in the same order. If you need to perform any non-deterministic operation like accessing the database, calling a third-party API, generating a random number, or getting the local time, you should do it in a step instead of directly in the workflow function. 2. Steps should be **idempotent**, meaning it should be safe to retry them multiple times. If a workflow fails while executing a step, it retries the step during recovery. However, once a step completes and is checkpointed, it is never re-executed. ### Upgrading Workflow Code One challenge you may encounter when operating long-running durable workflows in production is **how to deploy breaking changes without disrupting in-progress workflows.** A breaking change to a workflow is any change in what steps run or the order in which steps run. The issue is that if a breaking change was made to a workflow, the checkpoints created by a workflow that started on the previous version of the code may not match the steps called by the workflow in the new version of the code, which makes the workflow difficult to recover. DBOS supports two strategies for safely upgrading workflow code: **patching** and **versioning**. When using patching, you add DBOS patch statements to your code to make a breaking change in a conditional so old workflows can safely recover. When using versioning, DBOS versions applications and workflows so workflows only recover to processes running compatible code. Learn more about both strategies in the workflow upgrade tutorial ([Python](./python/tutorials/upgrading-workflows.md), [TypeScript](./typescript/tutorials/upgrading-workflows.md), [Go](./golang/tutorials/upgrading-workflows.md), [Java](./java/tutorials/upgrading-workflows.md)). ### Durable Queues One powerful feature of DBOS is that you can **enqueue** workflows for distributed execution with flow control. You can enqueue a workflow from within your DBOS application using the DBOS library or from another application using a DBOS Client ([Python](./python/reference/client.md), [TypeScript](./typescript/reference/client.md), [Go](./golang/reference/dbos-context.md#newclient), [Java](./java/reference/client.md)). An enqueued workflow may be dequeued and executed by your application's servers. All processes running DBOS periodically poll queues to find and execute new work. You can configure which processes listen to which queues. To help you operate at scale, DBOS queues provide **flow control**. You can customize the rate and concurrency at which workflows are dequeued and executed. For example, you can set a **worker concurrency** for each of your queues on each of your servers, limiting how many workflows from that queue may execute concurrently on that server. For more information on queues, see the docs ([Python](./python/tutorials/queue-tutorial.md), [TypeScript](./typescript/tutorials/queue-tutorial.md), [Go](./golang/tutorials/queue-tutorial.md), [Java](./java/tutorials/queue-tutorial.md)). ### Operating DBOS in Production with Conductor When operating DBOS durable workflows in production, we strongly recommend connecting your application to Conductor. Conductor is the control plane for your durable workflows, providing: - [**High availability**](./production/workflow-recovery.md): In a distributed environment with many executors running durable workflows, Conductor automatically detects when the execution of a durable workflow is interrupted (for example, if its executor is restarted, interrupted, or crashes) and recovers the workflow to another healthy executor. - [**Workflow and queue observability**](./conductor/workflow-management.md): Conductor provides dashboards of all active and past workflows and all queued tasks as well as real-time workflow visualization. - [**Workflow and queue management**](./conductor/workflow-management.md): From the Conductor dashboard, you can pause any workflow execution, start any stopped or enqueued workflow, or restart any workflow from a specific step. This is useful for rapidly responding to incidents or debugging. - [**Managed Retention Policies**](./conductor/retention.md): From the Conductor dashboard, manage how much workflow history each of your applications should retain and for how long to retain it. - [**Autoscaling and version management**](./conductor/autoscaling.md): Conductor computes how many executors each version of your application needs from queue utilization, so autoscalers like KEDA can size a deployment per application version, drain old versions down to zero, and drive rollouts. - [**Observability Integrations**](./conductor/metrics.md): Conductor exposes metrics about your applications' workflows, steps, and executors from a Prometheus-compatible endpoint, so you can monitor your DBOS applications in Datadog, Grafana, or any other tool that understands the OpenMetrics format. Architecturally, Conductor looks like this: Each of your application servers opens a secure websocket connection to Conductor. All of Conductor's capabilities are powered by these websocket connections. When you open a Conductor dashboard in your browser, your request is sent over websocket to one of your application servers, which serves the request (for example, retrieving a list of recent workflows) and sends the result back through the websocket. If one of your application servers fails, Conductor detects the failure through the closed websocket connection and, after a grace period, directs another server to recover its workflows. This architecture has two useful implications: 1. Conductor is **secure** and **privacy-preserving**. It does not have access to your database, nor does it need direct access to your application servers. Instead, your servers open outbound websocket connections to it and communicate exclusively through its websocket protocol. 2. Conductor is **off your workflows orchestration path**. Conductor drives observability, recovery, and retention policies, and is never involved in workflow execution (unlike the external orchestrators of other workflow systems). If your application's connection to Conductor is interrupted, it will continue to operate normally, and any failed workflows will automatically be recovered as soon as the connection is restored. For more information on Conductor, see [its docs](./conductor/overview.md). --- ## Alerting If you are using [Conductor](./overview.md), you can configure automatic alerts when certain failure conditions are met. You can configure alerts either in Conductor directly or on [Conductor-exported metrics](#metrics-based-alerts) using your existing observability stack. :::info Alerts require at least a [DBOS Teams](https://www.dbos.dev/dbos-pricing) plan. ::: #### Creating Alerts You can create new alerts (or view or update your existing alerts) from your application's "Alerting" page on the DBOS Console. Currently, you can create alerts for the following failure conditions: - If a certain number of workflows (parameterizable by workflow type) fail in a set period of time. - If a workflow remains enqueued for more than a certain period of time (parameterizable by queue name), indicating the queue is overwhelmed or stuck. - If an application is unresponsive (no connected executors, or connected but unresponsive executors). If multiple applications [share a system database](../explanations/sharing-a-system-database.md), each application's alerts consider only the workflows it owns. You may also specify an application to receive the alert—this does not need to be the same as the application that generated the alert. For some failure conditions (e.g., unresponsive application), the application receiving the alert is required to be different from the one generating it. #### Receiving Alerts You can register an alert handler in your application to receive alerts from Conductor. Your handler can log the alerts or forward them to another system, such as Slack or PagerDuty. Only one alert handler may be registered per application, and it must be registered before launching DBOS. If no handler is registered, alerts are logged automatically. The handler receives three arguments: - **rule_type**: The type of alert rule. One of `WorkflowFailure`, `SlowQueue`, or `UnresponsiveApplication`. - **message**: The alert message. - **metadata**: Additional key-value string metadata about the alert. The metadata keys depend on the rule type: **`WorkflowFailure`**: - `workflow_name`: The workflow name filter, or `*` for all workflows. - `failed_workflow_count`: The number of failed workflows detected in the time window. - `threshold`: The configured failure count threshold. - `period_secs`: The time window in seconds. **`SlowQueue`**: - `queue_name`: The queue name filter, or `*` for all queues. - `stuck_workflow_count`: The number of workflows that have been in the queue for longer than the time threshold. - `threshold_secs`: The enqueue time threshold in seconds. **`UnresponsiveApplication`**: - `application_name`: The application name. - `connected_executor_count`: The number of executors currently connected to this application. **Python** Example logging alerts: ```python from dbos import DBOS @DBOS.alert_handler def handle_alert(rule_type: str, message: str, metadata: dict[str, str]) -> None: DBOS.logger.warning(f"Alert received: {rule_type} - {message}") for key, value in metadata.items(): DBOS.logger.warning(f" {key}: {value}") ``` Example forwarding alerts to Slack using [incoming webhooks](https://docs.slack.dev/messaging/sending-messages-using-incoming-webhooks/) ```python @DBOS.alert_handler def handle_alert(rule_type: str, message: str, metadata: dict[str, str]) -> None: webhook_url = os.environ.get("SLACK_WEBHOOK_URL") slack_text = f"*Alert: {rule_type}*\n{message}\n" + "\n".join(f"• {k}: {v}" for k, v in metadata.items()) try: resp = requests.post(webhook_url, json={"text": slack_text}, timeout=10) resp.raise_for_status() except requests.RequestException as e: DBOS.logger.error(f"Failed to send Slack alert: {e}") ``` Example forwarding alerts to PagerDuty using the [Events API](https://developer.pagerduty.com/docs/events-api-v2-overview): ```python @DBOS.alert_handler def handle_alert(rule_type: str, message: str, metadata: dict[str, str]) -> None: routing_key = os.environ.get("PAGERDUTY_ROUTING_KEY") payload = { "routing_key": routing_key, "event_action": "trigger", "payload": { "summary": f"{rule_type}: {message}", "severity": "error", "source": "my-app", "custom_details": metadata, }, } try: resp = requests.post( "https://events.pagerduty.com/v2/enqueue", json=payload, timeout=10, ) resp.raise_for_status() except requests.RequestException as e: DBOS.logger.error(f"Failed to send PagerDuty alert: {e}") ``` See the [Python reference](../python/reference/contexts.md#alert_handler) for more details. **Go** Example logging alerts: ```go dbos.SetAlertHandler(dbosContext, func(ruleType string, message string, metadata map[string]string) { slog.Warn(fmt.Sprintf("Alert received: %s - %s", ruleType, message)) for key, value := range metadata { slog.Warn(fmt.Sprintf(" %s: %s", key, value)) } }) ``` Example forwarding alerts to Slack using [incoming webhooks](https://docs.slack.dev/messaging/sending-messages-using-incoming-webhooks/) ```go dbos.SetAlertHandler(dbosContext, func(ruleType string, message string, metadata map[string]string) { webhookURL := os.Getenv("SLACK_WEBHOOK_URL") var metaParts []string for k, v := range metadata { metaParts = append(metaParts, fmt.Sprintf("• %s: %s", k, v)) } slackText := fmt.Sprintf("*Alert: %s*\n%s\n%s", ruleType, message, strings.Join(metaParts, "\n")) body, _ := json.Marshal(map[string]string{"text": slackText}) resp, err := http.Post(webhookURL, "application/json", bytes.NewReader(body)) if err != nil { slog.Error(fmt.Sprintf("Failed to send Slack alert: %v", err)) return } defer resp.Body.Close() }) ``` Example forwarding alerts to PagerDuty using the [Events API](https://developer.pagerduty.com/docs/events-api-v2-overview): ```go dbos.SetAlertHandler(dbosContext, func(ruleType string, message string, metadata map[string]string) { routingKey := os.Getenv("PAGERDUTY_ROUTING_KEY") payload := map[string]any{ "routing_key": routingKey, "event_action": "trigger", "payload": map[string]any{ "summary": fmt.Sprintf("%s: %s", ruleType, message), "severity": "error", "source": "my-app", "custom_details": metadata, }, } body, _ := json.Marshal(payload) resp, err := http.Post("https://events.pagerduty.com/v2/enqueue", "application/json", bytes.NewReader(body)) if err != nil { slog.Error(fmt.Sprintf("Failed to send PagerDuty alert: %v", err)) return } defer resp.Body.Close() }) ``` See the [Go reference](../golang/reference/methods.md#alerting) for more details. **Typescript** Example logging alerts: ```typescript DBOS.setAlertHandler(async (ruleType: string, message: string, metadata: Record) => { DBOS.logger.warn(`Alert received: ${ruleType} - ${message}`); for (const [key, value] of Object.entries(metadata)) { DBOS.logger.warn(` ${key}: ${value}`); } }); ``` Example forwarding alerts to Slack using [incoming webhooks](https://docs.slack.dev/messaging/sending-messages-using-incoming-webhooks/) ```typescript DBOS.setAlertHandler(async (ruleType: string, message: string, metadata: Record) => { const webhookUrl = process.env.SLACK_WEBHOOK_URL!; const slackText = `*Alert: ${ruleType}*\n${message}\n` + Object.entries(metadata).map(([k, v]) => `• ${k}: ${v}`).join("\n"); const resp = await fetch(webhookUrl, { method: "POST", headers: { "Content-Type": "application/json" }, body: JSON.stringify({ text: slackText }), }); if (!resp.ok) { DBOS.logger.error(`Failed to send Slack alert: ${resp.status}`); } }); ``` Example forwarding alerts to PagerDuty using the [Events API](https://developer.pagerduty.com/docs/events-api-v2-overview): ```typescript DBOS.setAlertHandler(async (ruleType: string, message: string, metadata: Record) => { const routingKey = process.env.PAGERDUTY_ROUTING_KEY!; const payload = { routing_key: routingKey, event_action: "trigger", payload: { summary: `${ruleType}: ${message}`, severity: "error", source: "my-app", custom_details: metadata, }, }; const resp = await fetch("https://events.pagerduty.com/v2/enqueue", { method: "POST", headers: { "Content-Type": "application/json" }, body: JSON.stringify(payload), }); if (!resp.ok) { DBOS.logger.error(`Failed to send PagerDuty alert: ${resp.status}`); } }); ``` See the [TypeScript reference](../typescript/reference/methods.md#dbossetalerthandler) for more details. **Java** Example logging alerts: ```java import dev.dbos.transact.DBOS; import org.slf4j.Logger; import org.slf4j.LoggerFactory; Logger logger = LoggerFactory.getLogger("AlertHandler"); dbos.registerAlertHandler((ruleType, message, metadata) -> { logger.warn("Alert received: {} - {}", ruleType, message); metadata.forEach((key, value) -> logger.warn(" {}: {}", key, value)); }); ``` Example forwarding alerts to Slack using [incoming webhooks](https://docs.slack.dev/messaging/sending-messages-using-incoming-webhooks/) ```java dbos.registerAlertHandler((ruleType, message, metadata) -> { String webhookUrl = System.getenv("SLACK_WEBHOOK_URL"); String metaLines = metadata.entrySet().stream() .map(e -> "• " + e.getKey() + ": " + e.getValue()) .collect(Collectors.joining("\n")); String slackText = String.format("*Alert: %s*\n%s\n%s", ruleType, message, metaLines); HttpClient client = HttpClient.newHttpClient(); HttpRequest request = HttpRequest.newBuilder() .uri(URI.create(webhookUrl)) .header("Content-Type", "application/json") .POST(HttpRequest.BodyPublishers.ofString("{\"text\":\"" + slackText + "\"}")) .build(); try { client.send(request, HttpResponse.BodyHandlers.ofString()); } catch (Exception e) { logger.error("Failed to send Slack alert: {}", e.getMessage()); } }); ``` Example forwarding alerts to PagerDuty using the [Events API](https://developer.pagerduty.com/docs/events-api-v2-overview): ```java dbos.registerAlertHandler((ruleType, message, metadata) -> { String routingKey = System.getenv("PAGERDUTY_ROUTING_KEY"); String payload = String.format(""" {"routing_key":"%s","event_action":"trigger","payload":{ "summary":"%s: %s","severity":"error","source":"my-app", "custom_details":%s}}""", routingKey, ruleType, message, new ObjectMapper().writeValueAsString(metadata)); HttpClient client = HttpClient.newHttpClient(); HttpRequest request = HttpRequest.newBuilder() .uri(URI.create("https://events.pagerduty.com/v2/enqueue")) .header("Content-Type", "application/json") .POST(HttpRequest.BodyPublishers.ofString(payload)) .build(); try { client.send(request, HttpResponse.BodyHandlers.ofString()); } catch (Exception e) { logger.error("Failed to send PagerDuty alert: {}", e.getMessage()); } }); ``` #### Metrics-Based Alerts If you scrape [Conductor metrics](./metrics.md) into a monitoring system such as Prometheus (with [Alertmanager](https://prometheus.io/docs/alerting/latest/alertmanager/)), Datadog, or Grafana, you can define alerts directly on DBOS metrics. This is an alternative to the Conductor-managed alerts above, useful if you already run a monitoring and alerting stack. Here are some example alerting conditions written in [PromQL](https://prometheus.io/docs/prometheus/latest/querying/basics/): **Elevated workflow failures:** more than 10 workflows failing per minute for an application: ```promql sum by (application) (dbos_conductor_v1_workflow_failed_rate) * 60 > 10 ``` **Stuck queue:** a workflow has been waiting in a queue for more than 5 minutes: ```promql time() - dbos_conductor_v1_workflow_oldest_enqueued_timestamp_seconds > 300 ``` **Application offline:** no healthy executors are connected for an application: ```promql absent(dbos_conductor_v1_executor_count{application="my-app", status="HEALTHY"}) ``` Detecting an offline application is a "missing data" condition: a fully disconnected application emits no executor series at all. In PromQL, use `absent()` with an explicit `application` label for each app you want to monitor. --- ## Audit Logs If you are using [Conductor](./overview.md), you can retrieve an **audit log** of the mutating operations performed against your organization: registering and deleting applications, managing workflows and schedules, creating and revoking API keys, changing roles and membership, and updating organization settings. The audit log is append-only and records who did what, when, from where, and whether the operation succeeded. :::info Audit logs require a [DBOS Enterprise](https://www.dbos.dev/dbos-pricing) plan. ::: ### The Audit Logs Endpoint Conductor exposes an organization's audit log through the [Conductor API](./reference/conductor-api.md) at: ``` GET https://cloud.dbos.dev/conductor/v2/orgs/{orgName}/audit-logs ``` `{orgName}` is your DBOS organization name. The endpoint is authenticated with a Conductor API key, passed as a bearer token in the `Authorization` header. You can generate an API key from the [key settings page](https://console.dbos.dev/settings/apikey) of the DBOS Console. **Make sure the key has the [`organization.read`](./permissions.md) permission.** A read is a simple authenticated `GET`: ```bash curl -G https://cloud.dbos.dev/conductor/v2/orgs/my_org/audit-logs \ -H "Authorization: Bearer $DBOS_API_KEY" \ --data-urlencode "operation=workflow.cancel" \ --data-urlencode "limit=50" ``` :::note Audit logs are an organization-level concept, so a [self-hosted Conductor](./self-hosting/hosting-conductor.md) running with authentication disabled does not register this operation and responds `404`. See [Self-hosted differences](./reference/conductor-api.md#self-hosted-differences). ::: Entries are returned newest first (by emit time). #### Filtering and pagination By default the endpoint returns the most recent entries for your organization. You can narrow the results with these query parameters, all optional: | Parameter | Description | | --- | --- | | `startTime` | Only return entries at or after this time. [RFC 3339](https://www.rfc-editor.org/rfc/rfc3339) timestamp (e.g. `2026-07-01T00:00:00Z`), inclusive. | | `endTime` | Only return entries before this time. RFC 3339 timestamp, exclusive. | | `operation` | Only return entries for this [operation](#operations), matched exactly (e.g. `application.delete`). | | `subject` | Only return entries whose actor matches this value, compared against both the subject's display name (email or API-key name) and its id. | | `target` | Only return entries whose target resource id matches this value exactly (e.g. an application name or workflow id). | | `limit` | Maximum number of entries to return. Defaults to `100`; the maximum is `1000`, and a larger value is rejected as an error. | | `offset` | Number of matching entries to skip. Defaults to `0`. | Pagination is offset-based over the filtered, time-ordered results. To page through the log, hold the filters constant and advance `offset` by `limit` on each request. A page with fewer than `limit` entries means you have reached the end. For example, to fetch the second page of application deletions in June: ``` https://cloud.dbos.dev/conductor/v2/orgs/my_org/audit-logs?operation=application.delete&startTime=2026-06-01T00:00:00Z&endTime=2026-07-01T00:00:00Z&limit=100&offset=100 ``` ### The Response The endpoint returns a JSON array of audit entries: ```json [ { "id": "3f9a1c2e-6b0d-4f8a-9c1e-2a7b5d4c8e10", "emitTime": "2026-07-06T18:22:41.512Z", "operation": "workflow.cancel", "status": "success", "subject": { "type": "user", "id": "user_123", "display": "alice@example.com" }, "target": { "type": "workflow", "id": "e1b2c3d4-a5b6-7c8d-9e0f-1a2b3c4d5e6f" }, "sourceIp": "203.0.113.7", "details": { "application_name": "dbos-node-toolbox" } } ] ``` Each entry has the following fields: | Field | Description | | --- | --- | | `id` | Unique identifier of the audit entry. | | `emitTime` | When the operation was recorded, as an RFC 3339 timestamp. | | `operation` | The operation performed (see [Operations](#operations)). | | `status` | `success`, or `failure` if the operation was rejected or errored (for example, a denied attempt or invalid request). | | `subject` | Who performed the operation. | | `subject.type` | `user` or `api_key`. | | `subject.id` | Stable identifier of the user or API key. | | `subject.display` | Human-readable actor: the user's email or the API key's name. Preserved even if the user or key is later deleted. | | `target` | The resource the operation acted on. Omitted when no specific target applies. | | `target.type` | The [type](#target-types) of the target resource. | | `target.id` | Identifier or name of the target resource. | | `sourceIp` | IP address the request originated from. | | `details` | Operation-specific context, as a JSON object (see [Details](#details)). `null` when there is none. | #### Operations The `operation` field, and the `operation` query filter, use these values: | Category | Operations | | --- | --- | | Applications | `application.create`, `application.update`, `application.delete`, `application.set_latest_version` | | Workflows | `workflow.cancel`, `workflow.resume`, `workflow.restart`, `workflow.fork`, `workflow.fork_from_failure`, `workflow.delete`, `workflow.import`, `workflow.bulk_cancel`, `workflow.bulk_delete`, `workflow.bulk_resume` | | Schedules | `schedule.pause`, `schedule.resume`, `schedule.backfill`, `schedule.trigger` | | Alerting rules | `alerting_rule.create`, `alerting_rule.delete` | | API keys | `token.create`, `token.revoke` | | Roles | `role.create`, `role.delete`, `role.grant` | | Organization | `organization.update`, `user.join`, `user.remove` | #### Target types The `target.type` field is one of: `application`, `workflow`, `schedule`, `alerting_rule`, `token`, `role`, `user`, `organization`. #### Details `details` carries additional, operation-specific context. Keys include: | Key | Appears on | | --- | --- | | `application_name` | Any operation scoped to an application. | | `workflow_ids` | Bulk workflow operations (`workflow.bulk_cancel`, `workflow.bulk_delete`, `workflow.bulk_resume`). | | `private_mode` | `application.create` | | `permissions`, `applications` | `token.create` | | `permissions` | `role.create` | | `role_name` | `role.grant` | | `new_name`, `audit_log_retention_days` | `organization.update` | ### Retention Audit entries are retained per-organization for a configurable window; entries older than the window are automatically deleted. Retention defaults to **90 days** and can be set to any value between **7 and 3650 days**. The retention period can be set through the DBOS Console. --- ## Autoscaling and Version Management [Conductor](./overview.md) lets you attach autoscaling policies to your applications. An autoscaling policy computes how many executors your application needs, per application version, to drain one of your application's queues. A common example is configuring a [KEDA](https://keda.sh/) ScaledObject to size your application deployments based on queue utilization. :::info Autoscaling requires a [DBOS Teams](https://www.dbos.dev/dbos-pricing) plan. ::: To use policies: 1. **Attach an autoscaling policy** to an application, naming the queue whose backlog drives the executor count. 2. **Poll the desired executor count**, either one version at a time or for all active versions at once. All endpoints on this page are part of the [Conductor API](./reference/conductor-api.md); see that page for the base URL and authentication. The examples below use `$CONDUCTOR` for the base URL and `$CONDUCTOR_KEY` for an [API key](./permissions.md). ### How It Works Conductor counts the workflows that are `ENQUEUED` or `PENDING` on the policy queue, grouped by application version, and sizes each version to its own backlog: ``` desiredExecutors = ceil(queueDepth / workerConcurrency) ``` If the queue also has a global concurrency limit, the recommendation is additionally capped at `ceil(concurrency / workerConcurrency)`. Versions matter because a queued workflow is executed only by executors running the version it was enqueued under. (Note that the latest version of your application can also dequeue workflows that have not been assigned a version yet.) When you roll out a new version, its executors pick up new work while old version's executors must stay available until the old version's backlog drains. Conductor therefore reports the latest version as needing at least one executor, and reports an old version at zero once nothing is left for it on the queue. The policy queue must have a **worker concurrency** set. Work outside the policy queue, such as workflows started directly or enqueued on another queue, is not visible to the policy. Recommendations are computed from your application's [system database](../explanations/system-tables.md) through one of its healthy executors. :::info [Partitioned queues](../python/tutorials/queue-tutorial.md#partitioning-queues) are also sized by their queue-wide concurrency parameters, their per-partition limits are not considered. ::: ### Autoscaling From the Console The **Executors** tab of your application's page on the [DBOS Console](https://console.dbos.dev) lets you install the policy: In this view: - **Autoscaling policy**: pick the queue whose utilization should drive the executor count. Only eligible queues are listed, each with its worker concurrency. The two optional fields, maximum old versions and maximum executors per old version, are rollout caps that let you orchestrate deployments from an operator. Editing requires the `application.write` permission. - **Desired executors**: Conductor's live recommendation for the latest version, with the version it covers, when the backlog was observed, and the queue's backlog and per-worker limit behind the number. - **Connected executors**: the application's executors grouped by version, latest first, each panel showing how many are healthy and how many Conductor wants for that version. ### Attaching a Policy with the API Attach a policy with a `PUT`: ```shell curl -X PUT "$CONDUCTOR/v2/orgs/{orgName}/apps/{appName}/autoscaling-policy" \ -H "Authorization: Bearer $CONDUCTOR_KEY" \ -H "Content-Type: application/json" \ -d '{"queue": "orders"}' ``` Conductor validates the queue against a running executor before storing the policy and echoes the policy back as stored: ```json { "policy": { "queue": "orders" } } ``` The policy has one required field and an optional `rollout` section governing how old versions are sized: | Field | Description | | --- | --- | | `queue` | The queue whose utilization should drive the desired executor count. It must exist and have a worker concurrency set. | | `rollout.maxOldApplicationVersions` | How many old application versions the [all-versions endpoint](#all-versions-at-once) may include, newest first. Defaults to `0`, which means that only the latest version is reported. | | `rollout.maxExecutorsForOldApplicationVersions` | Cap every old version's recommendation at this many executors, regardless of its backlog. `0` is valid and reports old versions at zero. Omit to size old versions from their own backlog, uncapped. | For example, this policy keeps at most two old versions running, with at most one executor each, so most capacity goes to the latest version during a rollout: ```json { "queue": "orders", "rollout": { "maxOldApplicationVersions": 2, "maxExecutorsForOldApplicationVersions": 1 } } ``` `GET` the same path to read the stored policy (`404` when none is set), and `DELETE` it to turn autoscaling off. Setting and deleting a policy require the `application.write` permission and are recorded in the [audit log](./audit-logs.md). ### Reading the Desired Executor Count Conductor exposes two endpoints to read scaling recommendation: one version at a time, which suits an autoscaler like KEDA, or all versions at once, which suits an operator managing deployments. Both read endpoints return the same recommendation object per version: | Field | Description | | --- | --- | | `applicationVersion` | The application version this recommendation covers. | | `isLatest` | `true` for the application's latest registered version. | | `desiredExecutors` | How many executors of this version are needed to satisfy the queue load at the time of the reading. | | `queueName` | The queue the stored policy scales on. | | `queueDepth` | The `ENQUEUED` and `PENDING` backlog counted for this version on the policy queue. | | `observedAt` | When the backlog was measured, in epoch milliseconds. | #### One Version at a Time ```shell curl -H "Authorization: Bearer $CONDUCTOR_KEY" \ "$CONDUCTOR/v2/orgs/{orgName}/apps/{appName}/autoscale/versions/latest" ``` ```json { "applicationVersion": "1787155000092755696-000b07ce8f114cc4", "isLatest": true, "desiredExecutors": 4, "queueName": "orders", "queueDepth": 12, "observedAt": 1787155105672 } ``` The `{version}` path parameter is either `latest` or any version the application has registered. `latest` always resolves to whichever version is currently latest, so a single poller pointed at it keeps working across rollouts with no reconfiguration. The latest version is always reported as needing at least one executor. An old version is reported at zero once it has no work left on the queue, which signals that its executors could be torn down. The policy's `maxExecutorsForOldApplicationVersions` cap applies to old versions, and `maxOldApplicationVersions` does not apply to this endpoint, since it only ever reports the version you asked for. This endpoint is made to be polled per deployment, for example by one KEDA ScaledObject per version's Deployment. The response is an absolute executor count, so configure your autoscaler to map it one-to-one to replicas. #### All Versions at Once ```shell curl -H "Authorization: Bearer $CONDUCTOR_KEY" \ "$CONDUCTOR/v2/orgs/{orgName}/apps/{appName}/autoscale" ``` ```json [ { "applicationVersion": "1787155000092755696-000b07ce8f114cc4", "isLatest": true, "desiredExecutors": 4, "queueName": "orders", "queueDepth": 12, "observedAt": 1787155105672 }, { "applicationVersion": "1787140000012345678-9f0e1d2c3b4a5968", "isLatest": false, "desiredExecutors": 1, "queueName": "orders", "queueDepth": 2, "observedAt": 1787155105672 } ] ``` This endpoint returns one entry per version that should be running: - The latest version comes first and is always present, at one executor when it has no work. - It is followed by at most `maxOldApplicationVersions` old versions that still have work on the queue, most recently registered first. - An old version with no remaining work is omitted. Its absence is the signal that its deployment can be deleted. - `maxExecutorsForOldApplicationVersions` caps every old version's desired executors count. This shape suits a controller that owns the full set of deployments: it can create a deployment for each version in the response, size each to its `desiredExecutors`, and delete any deployment whose version is no longer listed. ### Errors | Status | Meaning | | --- | --- | | `400` | The policy names no queue, a queue the application does not define, or a queue that has no worker concurrency. | | `404` | The application has no autoscaling policy, or the requested version was never registered. | | `502` / `503` | No healthy executor of the application is connected, or every executor failed to answer. | :::warning A stored policy can become invalid if you later change the queue's definition. ::: --- ## Distributed Recovery If your application is connected to [DBOS Conductor](./overview.md), workflow recovery is automatic. When Conductor detects that an executor is unhealthy, it automatically signals another executor to recover its workflows. When an executor disconnects from Conductor, its status is changed to `DISCONNECTED` while Conductor waits for it to reconnect. If it has not reconnected after a certain period of time, its status is changed to `DEAD` and Conductor signals another executor to recover its workflows. After recovery is confirmed, Conductor deletes its record of the executor. By default, the executor timeout is 60 seconds, so Conductor waits 60 seconds after an executor disconnects before recovering its workflows. You can configure the executor timeout per application from the DBOS Console. --- ## Metrics If you are using [Conductor](./overview.md), you can scrape metrics about your applications' workflows, steps, and executors from a [Prometheus](https://prometheus.io/)-compatible endpoint. This lets you monitor your DBOS applications in Prometheus, Grafana, or any other tool that understands the [OpenMetrics](https://prometheus.io/docs/specs/om/open_metrics_spec/) format. :::info Metrics require at least a [DBOS Teams](https://www.dbos.dev/dbos-pricing) plan. ::: :::info Metrics require DBOS Python >=2.23.0 or DBOS TypeScript >=4.19. ::: ### The Metrics Endpoint Conductor exposes metrics for all of your applications at a single Prometheus-compatible OpenMetrics scrape endpoint: ``` https://cloud.dbos.dev/v1/metrics ``` The endpoint is authenticated with a Conductor API key, passed as a bearer token in the `Authorization` header. You can generate an API key from the [key settings page](https://console.dbos.dev/settings/apikey) of the DBOS Console. **Make sure to enable the metrics read permission for the key.** A scrape is a simple authenticated `GET`: ```bash curl https://cloud.dbos.dev/v1/metrics \ -H "Authorization: Bearer $DBOS_API_KEY" \ -H "Accept: application/openmetrics-text" ``` #### Endpoint Integrations The endpoint works with any tool that can scrape the OpenMetrics or Prometheus exposition format. **Prometheus** To scrape the endpoint from Prometheus, add a job like the following to your `prometheus.yml`. Store your API key in a file and reference it with `authorization.credentials_file` (or use `credentials` directly): ```yaml scrape_configs: - job_name: dbos scheme: https metrics_path: /v1/metrics scrape_interval: 60s honor_timestamps: true static_configs: - targets: ["cloud.dbos.dev"] authorization: type: Bearer credentials_file: /etc/prometheus/dbos_api_key ``` Set `honor_timestamps: true` so the window timestamps the endpoint emits are preserved. **Datadog** To collect the metrics with Datadog, use the [OpenMetrics integration](https://docs.datadoghq.com/integrations/openmetrics/) built into the Datadog Agent. Add an instance like the following to `conf.d/openmetrics.d/conf.yaml`, then restart the Agent: ```yaml instances: - openmetrics_endpoint: https://cloud.dbos.dev/v1/metrics namespace: dbos # Collect no more than once per minute (see "Aggregation Window" below). min_collection_interval: 60 metrics: - "dbos_conductor_v1_.*" headers: Authorization: "Bearer " Accept: application/openmetrics-text ``` **OpenTelemetry Collector** The [OpenTelemetry Collector](https://opentelemetry.io/docs/collector/)'s [`prometheus` receiver](https://github.com/open-telemetry/opentelemetry-collector-contrib/tree/main/receiver/prometheusreceiver) scrapes the endpoint and forwards the metrics to any backend you configure an exporter for. It takes a standard Prometheus `scrape_configs` block, so store your API key in a file and reference it with `authorization.credentials_file`: ```yaml receivers: prometheus: config: scrape_configs: - job_name: dbos scheme: https metrics_path: /v1/metrics # Scrape no more than once per minute (see "Aggregation Window" below). scrape_interval: 60s honor_timestamps: true static_configs: - targets: ["cloud.dbos.dev"] authorization: type: Bearer credentials_file: /etc/otelcol/dbos_api_key exporters: # Configure an exporter for your observability backend. otlphttp: endpoint: https://your-backend.example.com service: pipelines: metrics: receivers: [prometheus] exporters: [otlphttp] ``` Set `honor_timestamps: true` so the window timestamps the endpoint emits are preserved. #### Filtering metrics By default the endpoint returns every metric for every application in your organization. You can narrow a scrape with these repeatable query parameters: | Parameter | Description | | --- | --- | | `applications` | Only report metrics for the named application(s). Matched exactly. | | `workflow_names` | Only report metrics for the named workflow(s). | | `metrics` | Only emit the named metric families (e.g. `dbos_conductor_v1_workflow_success_rate`). | Each parameter may be repeated to select multiple values, for example: ``` https://cloud.dbos.dev/v1/metrics?applications=my-app&applications=my-other-app ``` ### Available Metrics Every metric this endpoint emits is an OpenMetrics **gauge**. All metric names are prefixed with `dbos_conductor_v1_`, and every series carries an `application` label. If multiple applications [share a system database](../explanations/sharing-a-system-database.md), each application's metrics count only the workflows and steps it owns. #### Aggregation window Each scrape reports data for the **most recently completed clock-aligned minute**. For example, a scrape at any time during `12:34` reports data aggregated over the window `[12:33:00, 12:34:00)`. Although every metric is a gauge, the value a gauge carries falls into one of three flavors, noted in the **Measurement** column below: - **Rate** — a per-second average over the window. For example, if 120 workflows succeeded in the window, `workflow_success_rate` reports `2`; multiply by 60 to recover the count over the minute. Because these are already-averaged gauges (not counters), do **not** wrap them in PromQL `rate()`. - **Point-in-time** — the value at scrape time, not tied to the window (for example, the number of workflows currently enqueued). - **Windowed** — an aggregate, such as a maximum, computed over the window. Rate and windowed metrics are stamped with the window's timestamp (so scrapes within the same minute deduplicate); point-in-time metrics carry no explicit timestamp and use the scrape time. The windowed metrics (`workflow_max_queue_wait_seconds`, `workflow_max_total_latency_seconds`, and `step_max_duration_seconds`) report a **maximum per label group**. When you combine groups in a query, aggregate them with `max()` — a maximum of maximums is still a maximum — rather than `sum()` or `avg()`, which are not meaningful over these values. #### Workflow metrics These metrics are labeled by `workflow_name` and, where noted, `queue_name`. | Metric | Measurement | Description | | --- | --- | --- | | `workflow_started_rate` | Rate | Workflows created per second. Labeled by queue. | | `workflow_dequeued_rate` | Rate | Enqueued workflows dequeued per second. Workflows that were never enqueued are not counted. Labeled by queue. | | `workflow_success_rate` | Rate | Workflows that completed successfully per second. Labeled by queue. | | `workflow_failed_rate` | Rate | Workflows that terminated with an error (`ERROR` or `MAX_RECOVERY_ATTEMPTS_EXCEEDED`) per second. Labeled by queue. | | `workflow_cancelled_rate` | Rate | Workflows that were cancelled per second. Labeled by queue. | | `workflow_enqueued_count` | Point-in-time | Workflows currently in the `ENQUEUED` state. Labeled by queue. | | `workflow_pending_count` | Point-in-time | Workflows currently in the `PENDING` (executing) state. | | `workflow_oldest_enqueued_timestamp_seconds` | Point-in-time | Unix timestamp (seconds) of the oldest workflow currently `ENQUEUED`. Use `time() - ` to derive its age. No series is emitted when no workflows are enqueued. Labeled by queue. | | `workflow_oldest_pending_timestamp_seconds` | Point-in-time | Unix timestamp (seconds) of the oldest workflow currently `PENDING`. Use `time() - ` to derive its age. No series is emitted when no workflows are pending. | | `workflow_max_queue_wait_seconds` | Windowed | Maximum queue wait (created to first started), in seconds, across workflows that completed successfully in the window. Labeled by queue. | | `workflow_max_total_latency_seconds` | Windowed | Maximum end-to-end latency (created to completed), in seconds, across workflows that completed successfully in the window. Labeled by queue. | #### Step metrics These metrics are labeled by `step_name`. | Metric | Measurement | Description | | --- | --- | --- | | `step_success_rate` | Rate | Workflow steps that completed successfully per second. | | `step_failed_rate` | Rate | Workflow steps that terminated with an error per second. | | `step_max_duration_seconds` | Windowed | Maximum single-step duration, in seconds, across steps that completed successfully in the window. | #### Executor metrics | Metric | Measurement | Description | | --- | --- | --- | | `executor_count` | Point-in-time | Number of executors registered for the application, labeled by `status` and `application_version`. | ### Example Queries A few example [PromQL](https://prometheus.io/docs/prometheus/latest/querying/basics/) queries: ```promql # Number of workflows that completed successfully in the past hour, across all workflows and queues. # workflow_success_rate is a per-second gauge, so average it over the hour and multiply by 3600 seconds. sum(avg_over_time(dbos_conductor_v1_workflow_success_rate{application="my-app"}[1h])) * 3600 # Age, in seconds, of the oldest currently enqueued workflow time() - dbos_conductor_v1_workflow_oldest_enqueued_timestamp_seconds # Number of healthy executors per application sum by (application) (dbos_conductor_v1_executor_count{status="HEALTHY"}) ``` --- ## DBOS Conductor Overview When operating DBOS durable workflows in production, we strongly recommend connecting your application to Conductor. Conductor is the control plane for your durable workflows, providing: - [**High availability**](./distributed-recovery.md): In a distributed environment with many executors running durable workflows, Conductor automatically detects when a workflow is interrupted (for example, if its executor disconnects or crashes) and recovers the workflow to another healthy executor. - [**Workflow and queue observability**](./workflow-management.md): Conductor provides dashboards of all active and past workflows and all queued tasks as well as real-time workflow visualization. - [**Workflow and queue management**](./workflow-management.md): From the Conductor dashboard, you can pause any workflow execution, start any stopped or enqueued workflow, or restart any workflow from a specific step. This is useful for rapidly responding to incidents or debugging. - [**Managed Retention Policies**](./retention.md): From the Conductor dashboard, manage how much workflow history each of your applications should retain and for how long to retain it. - [**Autoscaling and version management**](./autoscaling.md): Conductor computes how many executors each version of your application needs from queue utilization, so autoscalers like KEDA can size a deployment per application version, drain old versions down to zero, and drive rollouts. - [**Observability Integrations**](./metrics.md): Conductor exposes metrics about your applications' workflows, steps, and executors from a Prometheus-compatible endpoint, so you can monitor your DBOS applications in Datadog, Grafana, or any other tool that understands the OpenMetrics format. - [**Programmatic access**](./reference/conductor-api.md): Conductor's workflow, queue, and schedule management is available over an OpenAPI-described HTTP API and from the [`dbosctl` command-line client](./reference/dbosctl.md), so you can script incident response and wire Conductor into your own tooling. Architecturally, Conductor is not part of your workflows orchestration path. If your connection to Conductor is interrupted, your applications will continue operating normally. Recovery, observability, and workflow management will automatically resume once connectivity is restored. ### Connecting To Conductor To connect your application to Conductor, first register your application on the [DBOS Console](https://console.dbos.dev). **The name you register must match the name you give your application in its configuration.** Next, generate an API key from the [key settings page](https://console.dbos.dev/settings/apikey). By default, API keys do not expire, though they may be revoked at any time. Finally, supply that API key to your DBOS application to connect it to Conductor. This initiates a websocket connection with Conductor: :::tip The application name in your DBOS configuration must match the name with which you registered your app in Conductor. The name also identifies the application's data in its system database; see [sharing a system database](../explanations/sharing-a-system-database.md) for more information. ::: **Python** ```python config: DBOSConfig = { "name": "my-app-name", "application_version": "0.1.0", "system_database_url": os.environ.get("DBOS_SYSTEM_DATABASE_URL"), "conductor_key": os.environ.get("DBOS_CONDUCTOR_KEY") } DBOS(config=config) ``` **TypeScript** ```typescript DBOS.setConfig({ "name": "my-app-name", "applicationVersion": "0.1.0", "systemDatabaseUrl": process.env.DBOS_SYSTEM_DATABASE_URL, }); const conductorKey = process.env.DBOS_CONDUCTOR_KEY await DBOS.launch({conductorKey}) ``` **Go** ```go conductorKey := os.Getenv("DBOS_CONDUCTOR_KEY") dbosContext, err := dbos.NewDBOSContext(context.Background(), dbos.Config{ AppName: "dbos-starter", ApplicationVersion: "0.1.0", DatabaseURL: os.Getenv("DBOS_SYSTEM_DATABASE_URL"), ConductorAPIKey: conductorKey, }) ``` **Java** ```java String conductorKey = System.getenv("DBOS_CONDUCTOR_KEY"); DBOSConfig config = DBOSConfig.defaults("dbos-java-starter") .withAppVersion("0.1.0") .withDatabaseUrl(System.getenv("DBOS_SYSTEM_JDBC_URL")) .withConductorKey(conductorKey) ``` ### Managing Conductor Applications You can view all applications registered with Conductor on the DBOS Console: On your application's page, you can see all executors (processes) running that application that are currently connected to Conductor. Executors are identified by a unique ID that they generate and print on startup. When you restart an executor, it generates a new ID. You can tag executors with custom metadata (such as region or instance type) using the `conductor_executor_metadata` configuration option (in TypeScript, the `conductorExecutorMetadata` launch option). This metadata is displayed on the dashboard to help you identify executors. Conductor uses a WebSocket-based protocol to exchange workflow metadata and commands with your application. An application is shown as _available_ in Conductor when at least one of its processes is connected. Conductor has no access to your application's database or other private data. As a result, workflow-related features are only available while your application is connected to Conductor over this metadata-only connection. :::tip For isolation, you should set up a separate Conductor app for each environment in which you run your DBOS application. For example, you may want to have separate dev, staging, and prod Conductor apps. To facilitate this, pass in your application name as an environment variable, for example: **Python** ```python config: DBOSConfig = { "name": os.environ.get("DBOS_APPLICATION_NAME"), "application_version": "0.1.0", "system_database_url": os.environ.get("DBOS_SYSTEM_DATABASE_URL"), "conductor_key": os.environ.get("DBOS_CONDUCTOR_KEY") } DBOS(config=config) ``` **TypeScript** ```typescript DBOS.setConfig({ "name": process.env.DBOS_APPLICATION_NAME!, "applicationVersion": "0.1.0", "systemDatabaseUrl": process.env.DBOS_SYSTEM_DATABASE_URL, }); const conductorKey = process.env.DBOS_CONDUCTOR_KEY await DBOS.launch({conductorKey}) ``` **Go** ```go conductorKey := os.Getenv("DBOS_CONDUCTOR_KEY") dbosContext, err := dbos.NewDBOSContext(context.Background(), dbos.Config{ AppName: os.Getenv("DBOS_APPLICATION_NAME"), ApplicationVersion: "0.1.0", DatabaseURL: os.Getenv("DBOS_SYSTEM_DATABASE_URL"), ConductorAPIKey: conductorKey, }) ``` **Java** ```java String appName = System.getenv("DBOS_APPLICATION_NAME") String conductorKey = System.getenv("DBOS_CONDUCTOR_KEY"); DBOSConfig config = DBOSConfig.defaults(appName) .withAppVersion("0.1.0") .withDatabaseUrl(System.getenv("DBOS_SYSTEM_JDBC_URL")) .withConductorKey(conductorKey) ``` ::: #### Metadata-Only Mode :::info Metadata-Only mode requires at least a [DBOS Teams](https://www.dbos.dev/dbos-pricing) plan. ::: If an application handles especially sensitive data, you may consider enabling metadata-only mode for it. In this mode, only workflow and step metadata (but not data, like workflow or step inputs or outputs) is sent to Conductor. As a result, workflow data will not be visible from the console. Note that Conductor does not store application data in any mode. You can also enable metadata-only mode from your application, so it is enforced by the application process regardless of the setting in the console, by setting `conductor_metadata_only_mode` in your [Python configuration](../python/reference/configuration.md#conductor-settings) or `conductorMetadataOnlyMode` in your [TypeScript launch options](../typescript/reference/dbos-class.md#dboslaunch). --- ## Permissions and API Keys DBOS Conductor controls access to your organization's applications, workflows, and settings using **role-based access control (RBAC)** for users and **scoped API keys** for applications and automation. This page describes the permission model, the built-in and custom roles, and how to create and manage API keys. You manage permissions and API keys from the [DBOS Console](https://console.dbos.dev). ### Permissions Every action in Conductor, like viewing a workflow, registering an application, or creating an API key, requires a specific permission. | Permission | Grants the ability to | | --- | --- | | `organization.read` | View organization details, members, and roles. | | `organization.write` | Manage the organization: rename it, add and remove members, create and assign roles, manage billing. | `application.read` | View applications and their workflows, queues, schedules, executors, and alerting rules. | | `application.write` | Register, update, and delete applications; manage workflows (cancel, resume, fork, delete, import); and manage schedules and alerting rules. | | `metric.read` | Read application metrics, including the [Prometheus-compatible metrics endpoint](./metrics.md). | | `token.read` | List API keys. | | `token.write` | Create and revoke API keys. | | `websocket.connect` | For an API key, connect a running application executor to Conductor over its websocket. | ### Roles A **role** is a named set of permissions. Each member of an organization is assigned exactly one role, which determines everything they can do in that organization. #### Built-in roles Every organization has two built-in roles: | Permission | Organization Member | Organization Admin | | --- | :---: | :---: | | `organization.read` | ✅ | ✅ | | `organization.write` | | ✅ | | `application.read` | ✅ | ✅ | | `application.write` | ✅ | ✅ | | `metric.read` | ✅ | ✅ | | `token.read` | ✅ | ✅ | | `token.write` | ✅ | ✅ | | `websocket.connect` | ✅ | ✅ | - **Organization Admin** holds every permission. Admins can manage the organization and its members and roles, in addition to managing applications, API keys, and metrics. - **Organization Member** holds every permission *except* `organization.write`. Members can manage applications, workflows, and API keys and view metrics, but cannot change organization settings, manage members, or manage roles. Built-in roles are shared by every organization. They cannot be deleted or renamed. #### Custom roles :::info Custom roles require at least a [DBOS Teams](https://www.dbos.dev/dbos-pricing) plan. ::: Organization admins can create **custom roles** with any combination of permissions. This is useful for granting narrower access than the built-in roles; for example, a read-only role that can view applications and metrics but not modify them. You can manage custom roles from the [organization settings](https://console.dbos.dev/settings/organization) page in the console. #### Assigning roles to members Organization admins manage members and their roles from the [organization settings](https://console.dbos.dev/settings/organization) page in the console: - **Change a member's role** to any role whose permissions the admin also holds. - **Remove a member** from the organization. Changing a member's role replaces their previous role. A member always has exactly one role at a time. ### API keys An **API key** authenticates a non-human caller. API keys are used by running DBOS applications connecting to Conductor, but can also be used by scripts and CI/CD automation that call the Conductor API. Like a role, every API key carries a set of permissions. They can also be scoped to specific applications. API keys do not expire, but can be revoked at any time. A key can be renamed after creation without changing its secret, from the console, with [`dbosctl api-key rename`](./reference/dbosctl.md#dbosctl-api-key-rename), or through the [Conductor API](./reference/conductor-api.md#roles-permissions-and-api-keys). #### Permissions and application scope An API key has two independent restrictions: - **Permissions** — the set of capabilities the key grants, drawn from the same [permission catalog](#permissions) as roles. - **Application scope** — either *all applications* in the organization (org-wide), or a specific list of applications. A key scoped to specific applications is rejected on any request targeting an application outside its list, even if it holds the required permission. For example, an API key with only `application.read` scoped to a single application can read that application's workflows and nothing else. #### Using an API key Supply the key to your DBOS application to connect it to Conductor, as described in [Connecting to Conductor](./overview.md#connecting-to-conductor). You can also use an API key to authenticate HTTP calls to the Conductor API (for example the [metrics endpoint](./metrics.md)), passing the key as a bearer token: ``` Authorization: Bearer dbos_... ``` --- ## Conductor API Conductor is the control plane for your durable workflows, and this HTTP API is how you drive it programmatically: register applications with Conductor and tune their settings, search workflows, cancel or fork them, inspect queues and schedules, drive schedules, read metrics and audit logs, and manage members, roles, and API keys. This is the Conductor half of the [DBOS Console](https://console.dbos.dev) — what the console shows for an application connected to Conductor, whether that application runs on your own infrastructure or on DBOS Cloud. DBOS Cloud's own operations, such as [deploying an application](./dbos-cloud/deploying-to-cloud.md) or [provisioning a database](./dbos-cloud/database-management.md), are not part of this API; they have their own [CLI](./dbos-cloud/cloud-cli.md). The API is described by an OpenAPI 3.1 specification generated directly from the running server, so it is never out of date with the deployment serving it. Both the console and the [`dbosctl` CLI](./dbosctl.md) drive Conductor through this API, using clients generated from that spec. ### Base URL | Deployment | Base URL | | --- | --- | | DBOS-hosted Conductor | `https://cloud.dbos.dev/conductor` | | [Self-hosted Conductor](../self-hosting/hosting-conductor.md) | `http://:8090` (port `8090` by default) | Every path is relative to that base, so the full URL of an operation is, for example: ``` https://cloud.dbos.dev/conductor/v2/orgs/my_org/apps/my-app/workflows ``` The paths themselves are identical in both deployments; only the base differs. ### The OpenAPI Specification There are three ways to obtain the spec. **From DBOS-hosted Conductor.** The spec is served publicly (no authentication required) and reflects the currently deployed version: ```shell curl -O https://cloud.dbos.dev/conductor/v2/openapi.json ``` If your toolchain does not yet support OpenAPI 3.1, request the 3.0 downgrade instead: ```shell curl -O https://cloud.dbos.dev/conductor/v2/openapi-3.0.json ``` The spec served here is Conductor's own, with only its `servers` entry repointed at `/conductor` so that generated clients resolve paths correctly through DBOS-hosted Conductor. **From a self-hosted Conductor.** The server mounts the spec and an interactive browser at its root, all unauthenticated: | Path | Serves | | --- | --- | | `/openapi.json` | OpenAPI 3.1 specification (JSON) | | `/openapi.yaml` | The same specification in YAML | | `/openapi-3.0.json` | OpenAPI 3.0 downgrade | | `/docs` | Interactive API browser | | `/schemas/*` | The JSON Schema documents referenced by the spec | For example, with the Docker Compose setup from the [Self-Hosting Guide](../self-hosting/hosting-conductor.md), open `http://localhost:8090/docs` to explore the API in your browser. **From the Conductor image.** Conductor's `openapi` subcommand prints the spec to stdout without connecting to a database or requiring any runtime configuration, which is convenient in CI and code generation pipelines. The image's entrypoint starts the server, so override it to reach the subcommand: ```shell docker run --rm --entrypoint ./dbos-conductor dbosdev/conductor:latest \ openapi > openapi.json docker run --rm --entrypoint ./dbos-conductor dbosdev/conductor:latest \ openapi -spec-version 3.0 > openapi-3.0.json ``` Pin a version tag rather than `latest` when the spec feeds code generation, so a regenerated client changes only when you choose to bump it. :::info This emits the *complete* route surface, including operations that a no-auth deployment does not register. See [Self-hosted differences](#self-hosted-differences) below. ::: ### Authentication All authenticated requests carry a bearer token: ``` Authorization: Bearer ``` Conductor accepts two kinds of token, distinguished by their prefix: | Credential | Description | | --- | --- | | **API key** | A key beginning with `dbos_`, created with `POST /v2/orgs/{orgName}/tokens/{tokenName}` (or from the console, or with [`dbosctl api-key create`](./dbosctl.md#dbosctl-api-key-create)). Keys authenticate machine-to-machine callers and can be scoped to specific applications and permissions. A key carries no user identity, so it cannot call `GET /v2/users/me`. | | **User JWT** | An OIDC-issued JSON Web Token identifying a human user. This is what the console and `dbosctl login` use. | Both are sent the same way; Conductor tells them apart by the `dbos_` prefix. Authorization is enforced per operation using the permission model described in [Permissions and API Keys](../permissions.md) — a caller needs `application.read` to list workflows, `application.write` to cancel one, `organization.write` to manage members, and so on. An unauthenticated request returns `401`; an authenticated request lacking the required permission returns `403`. ### Resource Naming Almost every operation is scoped to an organization, and most are additionally scoped to an application: ``` /v2/orgs/{orgName}/apps/{appName}/workflows/{workflowId} ``` | Path parameter | Constraints | | --- | --- | | `orgName` | 3–30 characters, matching `^[a-z0-9_]+$` | | `appName` | 3–256 characters, matching `^[a-z0-9-_]+$` | Only two operations sit outside an organization: `POST /v2/users` and `GET /v2/users/me`. ### How Operations Are Served Conductor answers some requests from its own database and delegates the rest to your running application. Which one applies is a property of the resource, not of whether the operation reads or writes: | Served from | Operations | | --- | --- | | Conductor's database | Users, organizations, roles, permissions, and API keys; application registration, settings, and executor listing; alerting rules, audit logs, and metrics | | Your application | Everything under workflows, queues, and schedules — reads as much as mutations — plus listing application versions and setting the latest one | Conductor dispatches the second group over the websocket each executor holds open, waits for the reply, and returns it. Those operations therefore need a healthy executor connected for the target application, and they read whatever that executor's system database holds — Conductor neither caches nor mirrors it. They fail with `503` when the application has no healthy executor connected, and `502` when every healthy executor fails to answer. A `503` from `GET .../workflows` means your application is not connected, not that it has no workflows. ### Errors Errors are returned as [RFC 9457 problem details](https://www.rfc-editor.org/rfc/rfc9457) with content type `application/problem+json`: ```json { "status": 404, "title": "Not Found", "detail": "workflow 8f4a1e0c-1b2d-4c9a-a3e5-77d2c9a1b6ef not found", "type": "about:blank" } ``` Validation failures add an `errors` array locating each individual problem: ```json { "status": 422, "title": "Unprocessable Entity", "detail": "validation failed", "errors": [ { "location": "body.limit", "message": "expected integer", "value": "ten" } ] } ``` Common statuses: | Status | Meaning | | --- | --- | | `400` / `422` | Malformed request or failed validation | | `401` | Missing, expired, or invalid credentials | | `403` | Authenticated, but lacking the required permission | | `404` | No such organization, application, or resource — or an operation not registered in this deployment mode | | `502` / `503` | The application serving this resource failed to answer, or has no healthy executor connected — see [How Operations Are Served](#how-operations-are-served) | ### Listing and Filtering List operations that can return large result sets accept `limit` and `offset` query parameters for paging, and `sortDesc=true` to return newest results first. Time windows are given as `startTime` and `endTime` in RFC 3339 format. Workflow listing comes in two flavors: - **`GET .../workflows`** takes a few common filters as query parameters — `status`, `workflowName`, `limit`, `offset`, `sortDesc`, `loadInput`, `loadOutput` — and is convenient for quick queries. - **`POST .../workflows/search`** takes a JSON body and supports the full filter set the console uses: arrays of `status`, `workflowName`, `workflowIds`, `workflowIdPrefix`, `queueName`, `scheduleName`, `user`, `executorId`, `appVersion`, `parentWorkflowId`, and `forkedFrom`, plus `startTime`/`endTime`, `completedAfter`/`completedBefore`, `dequeuedAfter`/`dequeuedBefore`, `hasParent`, `wasForkedFrom`, `queuesOnly`, `attributes`, and the same paging and sorting fields. Workflow inputs and outputs can be large, so they are omitted unless you ask for them with `loadInput` and `loadOutput`. ### Endpoint Reference The tables below are a map of the whole API. The generated spec is the authoritative reference for request and response schemas of each operation. #### Users and organizations | Operation | Endpoint | | --- | --- | | Register user | `POST /v2/users` | | Get current user | `GET /v2/users/me` | | Get organization | `GET /v2/orgs/{orgName}` | | Update organization | `PATCH /v2/orgs/{orgName}` | | Join organization | `POST /v2/orgs/{orgName}/join` | | Generate join secret | `POST /v2/orgs/{orgName}/secrets` | | List members | `GET /v2/orgs/{orgName}/members` | | Remove member | `DELETE /v2/orgs/{orgName}/members/{username}` | | List domain claims | `GET /v2/orgs/{orgName}/domain-claims` | | Claim a domain | `POST /v2/orgs/{orgName}/domain-claims` | | Release a domain claim | `DELETE /v2/orgs/{orgName}/domain-claims/{domain}` | A **domain claim** automatically adds users who register with an email at that domain to your organization. On DBOS-hosted Conductor a claim takes effect only after DBOS approves it; on a self-hosted deployment it takes effect immediately. Claims apply to new registrations only: approving one never moves users who already have accounts, and releasing one never removes them. #### Roles, permissions, and API keys | Operation | Endpoint | | --- | --- | | List grantable permissions | `GET /v2/orgs/{orgName}/permissions` | | List roles | `GET /v2/orgs/{orgName}/roles` | | Create role | `POST /v2/orgs/{orgName}/roles` | | Delete role | `DELETE /v2/orgs/{orgName}/roles/{roleName}` | | Grant role to a member | `PUT /v2/orgs/{orgName}/members/{username}/roles/{roleName}` | | List API keys | `GET /v2/orgs/{orgName}/tokens` | | Create API key | `POST /v2/orgs/{orgName}/tokens/{tokenName}` | | Rename API key | `PATCH /v2/orgs/{orgName}/tokens/{tokenName}` | | Delete API key | `DELETE /v2/orgs/{orgName}/tokens/{tokenName}` | Create an API key with an optional body scoping it to particular applications and permissions; omitting a field leaves that dimension unscoped: ```json { "appNames": ["my-app"], "permissions": ["application.read", "metric.read"] } ``` The response contains the key's secret. It is returned **once**, at creation, and cannot be retrieved afterwards. A key can be renamed afterwards with `PATCH /v2/orgs/{orgName}/tokens/{tokenName}` and a body of `{"newName": "..."}`; the secret itself never changes, so rotating it means deleting the key and creating a new one. See [Permissions and API Keys](../permissions.md) for the full list of permissions. #### Applications | Operation | Endpoint | | --- | --- | | List applications | `GET /v2/orgs/{orgName}/apps` | | Get application | `GET /v2/orgs/{orgName}/apps/{appName}` | | Register application | `PUT /v2/orgs/{orgName}/apps/{appName}` | | Update application | `PATCH /v2/orgs/{orgName}/apps/{appName}` | | Delete application | `DELETE /v2/orgs/{orgName}/apps/{appName}` | | List versions | `GET /v2/orgs/{orgName}/apps/{appName}/versions` | | Set latest version | `PATCH /v2/orgs/{orgName}/apps/{appName}/versions/latest` | | List executors | `GET /v2/orgs/{orgName}/apps/{appName}/executors` | | List metrics | `GET /v2/orgs/{orgName}/apps/{appName}/metrics` | `PATCH .../apps/{appName}` is where an application's tuning settings live: the executor timeout, the global workflow timeout, the [workflow retention thresholds](../retention.md), and private mode. :::info `GET .../metrics` returns metrics for one application over a time window. If you want to scrape Conductor from Prometheus, Datadog, or Grafana, use the OpenMetrics endpoint described in [Metrics](../metrics.md) instead. ::: #### Workflows | Operation | Endpoint | | --- | --- | | List workflows | `GET /v2/orgs/{orgName}/apps/{appName}/workflows` | | Search workflows | `POST /v2/orgs/{orgName}/apps/{appName}/workflows/search` | | Workflow aggregates | `POST /v2/orgs/{orgName}/apps/{appName}/workflows/aggregates` | | Step aggregates | `POST /v2/orgs/{orgName}/apps/{appName}/steps/aggregates` | | Get workflow | `GET /v2/orgs/{orgName}/apps/{appName}/workflows/{workflowId}` | | List steps | `GET .../workflows/{workflowId}/steps` | | List events | `GET .../workflows/{workflowId}/events` | | List notifications | `GET .../workflows/{workflowId}/notifications` | | List streams | `GET .../workflows/{workflowId}/streams` | | Cancel workflow | `POST .../workflows/{workflowId}/cancel` | | Resume workflow | `POST .../workflows/{workflowId}/resume` | | Fork workflow | `POST .../workflows/{workflowId}/fork` | | Delete workflow | `DELETE .../workflows/{workflowId}` | | Export workflow | `GET .../workflows/{workflowId}/export` | | Import workflow | `POST .../workflows/import` | | Bulk cancel | `POST .../workflows/bulk-cancel` | | Bulk resume | `POST .../workflows/bulk-resume` | | Bulk delete | `POST .../workflows/bulk-delete` | | Bulk fork from failure | `POST .../workflows/bulk-fork-from-failure` | The semantics of cancelling, resuming, and forking are described in [Workflow Management](../workflow-management.md). The bulk variants take an array of workflow IDs and apply the same operation to each, which is far cheaper than issuing the calls one at a time. **Export** and **import** move a workflow and its steps between deployments as a JSON document — useful for reproducing a production failure in a development environment. #### Queues | Operation | Endpoint | | --- | --- | | List queues | `GET /v2/orgs/{orgName}/apps/{appName}/queues` | | Get queue | `GET /v2/orgs/{orgName}/apps/{appName}/queues/{queueName}` | #### Autoscaling | Operation | Endpoint | | --- | --- | | Get autoscaling policy | `GET /v2/orgs/{orgName}/apps/{appName}/autoscaling-policy` | | Set autoscaling policy | `PUT /v2/orgs/{orgName}/apps/{appName}/autoscaling-policy` | | Delete autoscaling policy | `DELETE /v2/orgs/{orgName}/apps/{appName}/autoscaling-policy` | | Desired executors, all versions | `GET /v2/orgs/{orgName}/apps/{appName}/autoscale` | | Desired executors, one version | `GET /v2/orgs/{orgName}/apps/{appName}/autoscale/versions/{version}` | The policy names the queue whose backlog drives the executor count; the two `autoscale` operations return how many executors each application version needs right now. `{version}` is a registered version or `latest`. See [Autoscaling and Version Management](../autoscaling.md). #### Schedules | Operation | Endpoint | | --- | --- | | List schedules | `GET /v2/orgs/{orgName}/apps/{appName}/schedules` | | Get schedule | `GET /v2/orgs/{orgName}/apps/{appName}/schedules/{scheduleName}` | | Pause schedule | `POST .../schedules/{scheduleName}/pause` | | Resume schedule | `POST .../schedules/{scheduleName}/resume` | | Trigger schedule | `POST .../schedules/{scheduleName}/trigger` | | Backfill schedule | `POST .../schedules/{scheduleName}/backfill` | **Trigger** runs a scheduled workflow immediately, out of band, and returns the started workflow's ID. **Backfill** replays a schedule across a past time window, starting one workflow per occurrence the schedule would have fired, and returns all of their IDs. #### Alerting and audit logs | Operation | Endpoint | | --- | --- | | List alerting rules | `GET /v2/orgs/{orgName}/apps/{appName}/alerting-rules` | | Create alerting rule | `POST /v2/orgs/{orgName}/apps/{appName}/alerting-rules` | | Delete alerting rule | `DELETE /v2/orgs/{orgName}/apps/{appName}/alerting-rules/{ruleId}` | | List audit logs | `GET /v2/orgs/{orgName}/audit-logs` | Alerting rules are described in [Alerting](../alerting.md). Audit log listing accepts `startTime`, `endTime`, `operation`, `subject`, and `target` filters alongside `limit` and `offset`; see [Audit Logs](../audit-logs.md). ### Self-Hosted Differences A self-hosted Conductor can run with OIDC authentication enabled or with authentication disabled entirely (see the [Self-Hosting Guide](../self-hosting/hosting-conductor.md)). In no-auth mode there is no user identity and no multi-organization concept, so the operations that depend on them are **not registered at all** and respond `404`: - every organization operation: `getOrg`, `updateOrg`, `joinOrg`, `generateSecret`, `listMembers`, `removeMember`, `listDomainClaims`, `requestDomainClaim`, and `releaseDomainClaim`; - every role operation: `listRoles`, `createRole`, `deleteRole`, `grantRole`; - the user operations `registerUser` and `getCurrentUser`; - audit log listing, `listAuditLogs`. These operations are marked in the spec with the `x-dbos-requires-oauth` extension, so a generated client or a tool reading the spec can identify them without hardcoding a list: ```shell jq -r '.paths | to_entries[] | .key as $p | .value | to_entries[] | select(.value["x-dbos-requires-oauth"]) | "\(.key | ascii_upcase) \($p)"' openapi.json ``` Everything else — applications, workflows, queues, schedules, alerting, and metrics — behaves identically in all three modes. ### Generating a Client Because the spec is generated from the server rather than maintained by hand, generating your client from it is the recommended way to call the API. Any OpenAPI generator works. For example, with [`oapi-codegen`](https://github.com/oapi-codegen/oapi-codegen) for Go: ```shell curl -o openapi.json https://cloud.dbos.dev/conductor/v2/openapi.json go tool oapi-codegen -package conductor -generate client,types openapi.json > conductor.gen.go ``` Or with [`openapi-python-client`](https://github.com/openapi-generators/openapi-python-client): ```shell curl -o openapi-3.0.json https://cloud.dbos.dev/conductor/v2/openapi-3.0.json openapi-python-client generate --path openapi-3.0.json ``` Pin the generated client to a checked-in copy of the spec and regenerate deliberately, so that a change to the deployed API surfaces as a reviewable diff rather than as a silent change in your build. If you would rather not write a client at all, the [`dbosctl` CLI](./dbosctl.md) already covers the operational surface of this API from the command line. --- ## Account Management In this guide, you'll learn how to manage DBOS Cloud accounts. #### New User Registration You can sign up for an account on the [DBOS Cloud console](https://console.dbos.dev/login-redirect). Additionally, all `dbos-cloud` commands prompt you to register a new account if you don't already have one. #### Authenticating Programatically Sometimes, such as in a CI/CD pipeline, it is useful to authenticate programatically without providing credentials through a browser-based login portal. DBOS Cloud provides this capability with refresh tokens. To obtain a refresh token, run: ``` dbos-cloud login --get-refresh-token ``` This command has you authenticate through the browser, but obtains a refresh token and stores it in `.dbos/credentials`. Once you have your token, you can use it to authenticate programatically without going through the browser with the following command: ``` dbos-cloud login --with-refresh-token ``` Refresh tokens automatically expire after a year or after a month of inactivity. You can manually revoke them at any time: ``` dbos-cloud revoke ``` :::warning Until they expire or are revoked, refresh tokens can be used to log in to your account. Treat them as secrets and keep them safe! ::: #### Organization Management :::info This feature is currently only available to [DBOS Pro or Enterprise](https://www.dbos.dev/pricing) subscribers. ::: Organizations allow multiple users to collaboratively manage applications. When a user creates an account, they are automatically added to an organization containing only them, where the organization name is the same as their username. You can manage your organization from the [cloud console organizations page](https://console.dbos.dev/settings/organization): ![Organizations](./assets/cc-orgs.png) ##### Organization Admins The original creator of an organization is the organization admin. Only the organization admin can invite new users, delete existing users, or rename the organization. All users have full access to organization resources, including databases and applications. ##### Inviting New Users To invite a new user to your organization, click the "Generate Invite Link" button. This generates a **single-use** URL for joining your organization. When a user signs in to the cloud console using that URL, they are prompted to join your organization. If they do not have an account, they are prompted to create one. If they already have an account, they must delete all resources (applications and databases) before joining your organization. ##### Renaming Your Organization You can rename your organization by clicking the icon next to your organization name. Note that applications belonging to organizations are hosted at the URL `https://-.cloud.dbos.dev/`. **Therefore, renaming your organization changes your application URLs**. ##### Removing Users The organization admin can remove any other user from their organization. This immediately terminates their access to all organization resources. --- ## Application Management #### Deploying Applications To deploy your application to DBOS Cloud or update an existing application, run this command in its root directory: ```shell dbos-cloud app deploy ``` Each time you deploy an application, the following steps execute: 1. **Upload**: An archive of your application folder is created and uploaded to DBOS Cloud. This archive can be up to 500 MB in size. 2. **Configuration**: Your application's dependencies [are installed](#dependency-management). 3. **Migration**: If you specify database migrations in your `dbos-config.yaml`, these are run on your cloud database. 4. **Deployment**: Your application is deployed to a number of [Firecracker microVMs](https://firecracker-microvm.github.io/) also referred to as "executors." By default, these have 1 vCPU and 512MB of RAM. The amount of memory allocated to each microVM is [configurable](./cloud-cli.md#dbos-cloud-app-update). After an application is deployed, it is assigned a domain of the form `https://-.cloud.dbos.dev/`. If your account is part of an [organization](./account-management.md#organization-management), organization name is used instead of username. :::tip * Applications should serve requests from port 8000 (Python—the default port for FastAPI and Gunicorn) or 3000 (TypeScript—the default port for Express and Koa). * Multiple applications can connect to the same Postgres database server—they are deployed to isolated databases on that server. * To change the database server of a deployed application, use `dbos-cloud app change-database-instance`. ::: ##### Applications Configuration You need to provide a valid `dbos-config.yaml` file when deploying an application to DBOS Cloud. The required fields are: - **name**: Your application name. This is the name with which your application is registered. - **language**: `node` or `python` - **runtimeConfig.start**: the command used to start your application. For example, `fastapi run` Note that some fields from dbos-config.yaml will be **ignored** during cloud deployments: - **database_url** and **database** connection-related fields. DBOS Cloud automatically applies the connection information of your cloud database server. ##### Dependency Management **Python** For Python applications, DBOS Cloud installs all dependencies from your `requirements.txt` file. The maximum size of your application after all dependencies are installed is 2 GB. **TypeScript** For TypeScript applications, DBOS Cloud installs all dependencies from your `package-lock.json` file (or from `package.json` if no lockfile is provided). The maximum size of your application after all dependencies are installed is 2 GB. After all dependencies are installed, your application is compiled using `npm run build`. ##### Customizing MicroVM Setup DBOS Pro subscribers can provide a _setup script_ that runs before their application is configured. This script can customize the runtime environment for your application, for example installing system packages and libraries. A setup script must be specified in your `dbos-config.yaml` like so: ```yaml title="dbos-config.yaml" runtimeConfig: # Script DBOS Cloud runs to customize your application runtime. # Requires a DBOS Pro subscription. setup: - "./build.sh" # Command DBOS Cloud executes to start your application. start: ``` A setup script may install system packages or libraries or otherwise customize the microVM image. For example: ```shell title="build.sh" #!/bin/bash # Install the traceroute package for use in your application apt install traceroute ``` ##### Ignoring files with .dbosignore A `.dbosignore` file at the root of your project instructs the DBOS Cloud CLI to exclude resources from application deployment. The syntax for this file is similar to `.gitignore`: - Patterns are compatible with the [fast-glob library](https://www.npmjs.com/package/fast-glob) - Lines ending with `/` are transformed into a recursive ignore `/**` to exclude everything within a directory. - Lines starting with `#` are ignored. - Some patterns are automatically excluded: ```shell **/.dbos/** **/node_modules/** **/dist/** **/.git/** **/dbos-config.yaml **/venv/** **/.venv/** **/.python-version ``` #### Monitoring and Debugging Applications Here are some useful tools to monitor and debug applications: - The [cloud console](https://console.dbos.dev) provides a web UI for viewing your applications and their traces and logs. - To retrieve the last `N` seconds of your application's logs, run [`dbos-cloud app logs -l `](./cloud-cli.md#dbos-cloud-app-logs). Note that new log entries take a few seconds to appear. - To retrieve the status of a particular application, run [`dbos-cloud app status `](./cloud-cli.md#dbos-cloud-app-status). To list all applications, run [`dbos-cloud app list`](./cloud-cli.md#dbos-cloud-app-list). #### Managing Application Versions Each time you deploy an application, it creates a new version with a unique ID. You can view all previous versions of your application from the [cloud console](https://console.dbos.dev) or list them by running: ``` dbos-cloud app versions ``` You can redeploy a previous version of your application by passing `--previous-version ` to the [`app deploy`](./cloud-cli.md#dbos-cloud-app-deploy) command. ```shell dbos-cloud app deploy --previous-version ``` #### MicroVM Termination DBOS Cloud may, from time to time, stop your microVMs due to a variety of reasons, including app upgrade or scaling down (see below). When a microVM is stopped, DBOS Cloud performs the following steps: 1. Stops routing new HTTP traffic to the VM. 2. Sends SIGTERM to the `dbos` process (launched by your start command) and waits up to 10 seconds for it to exit. 3. If the process is still running, terminates it forcefully. Any `PENDING` workflows run by the microVM are then recovered on another VM. See below. #### Workflow Recovery When a microVM running in DBOS Cloud stops (either due to a process crash or when scaling down), all of its workflows are automatically recovered to another microVM running the same application version. If no other microVM of that application version exists, DBOS Cloud launches a new one and instructs it to recover the workflows. When you deploy a new version of your application, DBOS Cloud routes all requests and scheduled workflows to microVMs of the new application version. Then, DBOS Cloud attempts to decommission microVMs running the previous application version. If there are still `PENDING` or `ENQUEUED` workflows of that code version, DBOS Cloud leaves some number of microVMs alive to process those workflows until all are complete. Periodically, DBOS Cloud checks if there are any `PENDING` or `ENQUEUED` workflows not assigned to any microVM. If any are found, DBOS Cloud recovers them to a microVM of the appropriate application version (starting one if necessary). #### Automatic and Manual Scaling Accounts on the free 30-day trial are limited to 1 microVM per app. Apps for Pro accounts are auto-scaled. Auto-scaling occurs based on CPU or queue utilization. For the latter, the queue must have `worker_concurrency` (or `workerConcurrency`) set. DBOS Cloud computes the current parallel task capacity - the number of tasks that can be executed simultaneously - using the product of worker_concurrency and the current number of microVMs: `capacity = worker_concurrency * num_microvms`. Scaling up occurs when: - the average microVM CPU utilization exceeds 85% for several seconds, or - the number of enqueued tasks exceeds capacity for at least one of the queues. Inversely, the app scales down when the average CPU utilization drops below 40% and the number of enqueued tasks drops below capacity for all queues. To alter the auto-scaling behavior, you can manually set `min-executors` and/or `max-executors` using `app update` (see below). #### Updating Applications To change your executor RAM or the default number of microVMs, run: ```shell dbos-cloud app update [options] ``` See the [DBOS Cloud CLI reference](./cloud-cli.md#dbos-cloud-app-update) for a list of properties you can update. Note that `app update` does not trigger a redeploy of the code, which you can do with the [`app deploy`](./cloud-cli.md#dbos-cloud-app-deploy) command. #### Deleting Applications To delete an application, run: ```shell dbos-cloud app delete ``` You can also drop the application database with the `--dropdb` argument. As each application has its own isolated database, this does not affect your other applications. ```shell dbos-cloud app delete --dropdb ``` :::warning This is a destructive operation and cannot be undone. ::: --- ## Bringing Your Own Database In this guide, you'll learn how to bring your own Postgres database instance to DBOS Cloud and deploy your applications to it. #### Linking Your Database to DBOS Cloud To bring your own Postgres database instance to DBOS Cloud, you must first create a role DBOS Cloud can use to deploy and manage your apps. By default this role must be named `dbosadmin` and must have the `LOGIN` and `CREATEDB` privileges: ```sql CREATE ROLE dbosadmin WITH LOGIN CREATEDB PASSWORD ''; ``` If you cannot use the name `dbosadmin`, you can specify a different role name when linking your database with the `--dbos-admin-name` flag. Next, link your database instance to DBOS Cloud, entering the password for the admin role when prompted. You must choose a database instance name that is 3 to 16 characters long and contains only lowercase letters, numbers and underscores. ```shell dbos-cloud db link -H -p ``` You can now register and deploy applications with this database instance as normal! Check out our [applications management](./application-management.md) guide for details. :::tip DBOS Cloud is currently hosted in AWS us-east-1. For maximum performance, we recommend linking a database instance hosted there. ::: --- ## CI/CD Best Practices ### Staging and Production Environments To make it easy to test changes to your application without affecting your production users, we recommend using separate staging and production environments. You can do this by deploying your application with different names for staging and production. For example, when deploying `my-app` to staging, deploy using: ```shell dbos-cloud app deploy my-app-staging ``` When deploying to production, use: ```shell dbos-cloud app deploy my-app-prod ``` `my-app-staging` and `my-app-prod` are completely separate and isolated DBOS applications. There's nothing special about the `-staging` and `-prod` suffixes—you can use any names you like. :::info If you manually specify the application database name by setting `app_db_name` in `dbos-config.yaml`, you must ensure each environment uses a different value of `app_db_name`. ::: ### Authentication You should use [refresh tokens](./account-management#authenticating-programatically) to programmatically authenticate your CI/CD user with DBOS Cloud. :::info Upgrading to a DBOS Cloud paid plan will unlock [multi-user organizations](./account-management#organization-management) which you can use to setup dedicated users for CI/CD. ::: --- ## Cloud CLI Reference ### Installation To globally install the DBOS Cloud CLI, run the following command: ``` npm install -g @dbos-inc/dbos-cloud@latest ``` ### User Management Commands #### `dbos-cloud register` **Description:** This command creates and registers a new DBOS Cloud account. It provides a URL to a secure login portal you can use to create an account from your browser. **Arguments:** - `-u, --username `: Your DBOS Cloud username. Must be between 3 and 30 characters and contain only lowercase letters, numbers, and underscores (`_`). - `-s, --secret [string]`: (Optional) An [organization secret](./account-management.md#organization-management) given to you by an organization admin. If supplied, adds your newly registered account to the organization. :::info If you register with an email and password, you also need to verify your email through a link we email you. ::: --- #### `dbos-cloud login` **Description:** This command logs you in to your DBOS Cloud account. It provides a URL to a secure login portal you can use to authenticate from your browser. :::info When you log in to DBOS Cloud from an application, a token with your login information is stored in the `.dbos/` directory in your application package root. ::: --- #### `dbos-cloud logout` **Description:** This command logs you out of your DBOS Cloud account. --- ### Database Instance Management Commands #### `dbos-cloud db provision` **Description:** This command provisions a Postgres database instance to which your applications can connect. **Arguments:** - ``: The name of the database instance to provision. Must be between 3 and 30 characters and contain only lowercase letters, numbers, underscores, and dashes. - `-U, --username `: Your username for this database instance. Must be between 3 and 16 characters and contain only lowercase letters, numbers, and underscores. - `-W, --password [string]`: Your password for this database instance. If not provided, will be prompted on the command line. Passwords must contain 8 or more characters. --- #### `dbos-cloud db list` **Description:** This command lists all Postgres database instances provisioned by your account. **Arguments:** - `--json`: Emit JSON output **Output:** For each provisioned Postgres database instance, emit: - `PostgresInstanceName`: The name of this database instance. - `HostName`: The hostname of this database instance. - `Port`: The connection port for this database instance. - `Status`: The current status of this database instance (available or unavailable). - `AdminUsername`: The administrator username for this database instance. --- #### `dbos-cloud db status` **Description:** This command retrieves the status of a Postgres database instance **Arguments:** - ``: The name of the database instance whose status to retrieve. - `--json`: Emit JSON output **Output:** - `PostgresInstanceName`: The name of the database instance. - `HostName`: The hostname of the database instance. - `Port`: The connection port for the database instance. - `Status`: The current status of the database instance (available or unavailable). - `AdminUsername`: The administrator username for the database instance. --- #### `dbos-cloud db reset-password` **Description:** This command resets your password for a Postgres database instance. **Arguments:** - `[database-instance-name]`: The name of the database instance whose password to reset. - `-W, --password [string]`: Your new password for this database instance. If not provided, will be prompted on the command line. Passwords must contain 8 or more characters. --- #### `dbos-cloud db destroy` **Description:** This command destroys a previously-provisioned Postgres database instance. **Arguments:** - ``: The name of the database instance to destroy. --- #### `dbos-cloud db url` **Description:** This command retrieves your cloud database connection URL. **Arguments:** - `[database-instance-name]`: The name of the database instance to which to connect. - `-S, --show-password`: Whether to show your database password in the output. - `-W, --password [string]`: Your password for this database instance. If not provided, will be prompted on the command line. --- #### `dbos-cloud db link` **Description:** This command links your own Postgres database instance to DBOS Cloud. Before running this command, please first follow our [tutorial](./byod-management) to set up your Postgres database. :::info This feature is currently only available to [DBOS Pro or Enterprise](https://www.dbos.dev/pricing) subscribers. ::: **Arguments:** - ``: The name of the database instance to link. Must be between 3 and 30 characters and contain only lowercase letters, numbers, underscores, and dashes. - `-H, --hostname `: The hostname for your Postgres database instance (required). - `-p, --port [number]`: The connection port for your Postgres database instance (default: `5432`). - `-W, --password [string]`: The password for the `dbosadmin` role. If not provided, will be prompted on the command line. Passwords must contain 8 or more characters. - `--dbos-admin-name `: Specify a custom Postgres role name for DBOS Cloud to administer the database as (default: `dbosadmin`). --- #### `dbos-cloud db unlink` **Description:** This command unlinks a previously linked Postgres database instance. **Arguments:** - ``: The name of the database instance to unlink. --- ### Application Management Commands #### `dbos-cloud app deploy` **Description:** This command must be run from an application root directory. It executes the migration commands declared in `dbos-config.yaml`, deploys the application to DBOS Cloud (or updates its code if already deployed), and emits the URL at which the application is hosted, which is `https://-.cloud.dbos.dev/`. **Arguments:** - `[application-name]`: The name of the application to deploy. By default we obtain the application name from `dbos-config.yaml`. This argument overrides the package name. - `-d, --database `: The name of the Postgres database instance to which this application will connect. This may only be set the first time an application is deployed and cannot be changed afterwards. - `--verbose`: Logs debug information about the deployment process, including config file processing and files sent. - `--configFile`: DBOS config file path (default: `dbos-config.yaml`). - `-p, --previous-version `: The ID of a previous version of this application. If this is supplied, redeploy that version instead of deploying from the application directory. This will fail if the previous and current versions have different database schemas. You can list previous versions and their IDs with the [versions command](#dbos-cloud-app-versions). --- #### `dbos-cloud app update` **Description:** Update an application metadata in DBOS Cloud. Increasing RAM or adjusting autoscaling configuration requires a DBOS Pro subscription. **Arguments:** - `[application-name]`: The name of the application to update. - `--executors-memory-mib`: The amount of RAM, in MiB, to allocate to the application's executors. This value must be between 512 and 5120. Additional RAM is [billed](https://www.dbos.dev/dbos-pricing). - `--min-executors `: The minimum number of microVMs to be allocated to this application. Acts as a floor for autoscaling. - `--max-executors `: The maximum number of microVMs to be allocated to this application. The app won't auto-scale to a larger number. :::info This command does not trigger a redeployment of the application. To apply changes affecting the application's executors, you must redeploy the application with [`dbos-cloud app deploy`](#dbos-cloud-app-deploy). ::: --- #### `dbos-cloud app delete` **Arguments:** - `[application-name]`: The name of the application to delete. - `--dropdb`: Drop the application's database during deletion. **Description:** Delete an application from DBOS Cloud. If run in an application root directory with no application name provided, delete the local application. By default, this command does not drop your application's database. You can use the `--dropdb` parameter to drop your application's database (not the Postgres instance) and delete all application data. To destroy the previously-provisioned Postgres instance, please use [`dbos-cloud db destroy`](#dbos-cloud-db-destroy). --- #### `dbos-cloud app list` **Description:** List all applications you have registered with DBOS Cloud. **Arguments:** - `--json`: Emit JSON output **Output:** For each registered application, emit: - `Name`: The name of this application - `ID`: The unique ID DBOS Cloud assigns to this application. - `PostgresInstanceName`: The Postgres database instance to which this application is connected. - `ApplicationDatabaseName`: The database on this instance on which this application stores data. - `Status`: The current status of this application (available or unavailable). - `Version`: The currently deployed version of this application. - `AppURL`: The URL at which the application is hosted. --- #### `dbos-cloud app status` **Arguments:** - `[application-name]`: The name of the application to retrieve. **Description:** Retrieve an application's status. If run in an application root directory with no application name provided, retrieve the local application's status. **Arguments:** - `--json`: Emit JSON output **Output:** - `Name`: The name of this application - `ID`: The unique ID DBOS Cloud assigns to this application. - `PostgresInstanceName`: The Postgres database instance to which this application is connected. - `ApplicationDatabaseName`: The database on this instance on which this application stores data. - `Status`: The current status of this application (available or unavailable). - `Version`: The currently deployed version of this application. - `AppURL`: The URL at which the application is hosted. --- #### `dbos-cloud app versions` **Arguments:** - `[application-name]`: The name of the application to retrieve. **Description:** Retrieve a list of an application's past versions. A new version is created each time an application is deployed. If run in an application root directory with no application name provided, retrieve versions of the local application. **Arguments:** - `--json`: Emit JSON output **Output:** For each previous version of this application, emit: - `ApplicationName`: The name of this application. - `Version`: The ID of this version. - `CreationTime`: The timestamp (in UTC with [RFC3339](https://datatracker.ietf.org/doc/html/rfc3339) format) at which this version was created. --- #### `dbos-cloud app logs` **Description:** Retrieve an application's logs. **Arguments:** - `[application-name]`: The name of the application. - `-l, --last `: How far back to query, in seconds from current time. Default is 3600 (one hour). --- #### `dbos-cloud app cmd` **Description:** A debugging utility that lets you run a shell command on one of your app's executors. Prints the `stderr` and `stdout` output by the command. The command must finish in 10 seconds. Every command is also recorded, without its output, in the app logs at `WARN` level. Note that stopping the running `dbos` process destroys the executor and causes it to be replaced by a new one. **Arguments:** - `-e, --executor-id `: The ID of the executor to use (see app logs). - `-c, --command `: The shell command to run. --- #### `dbos-cloud app resource-usage` **Description:** Retrieve your applications' resource usage for a specific time interval. If no time range is provided, queries for a recent completed 1-minute interval of data. **Arguments:** - `-s, --since `: UTC time since which to start querying (formatted as 2006-01-02 15:04:05.000000). Defaults to the start of a 1-minute interval ~2 minutes ago. - `-u --upto `: UTC time up to which to start querying (formatted as 2006-01-02 15:04:05.000000). Defaults to the end of a 1-minute interval ~2 minutes ago. - `-g, --group-by `: Time interval for grouping data: 'minute', 'hour', or 'day', defaults to 'minute'. --- #### `dbos-cloud app change-database-instance` **Description:** This command must be run from an application root directory. It redeploys the application to a new database instance. **Arguments:** - `--verbose`: Logs debug information about the deployment process, including config file processing and files sent. - `-d, --database ` The name of the new database instance for this application. - `-p, --previous-version `: The ID of a previous version of this application. If this is supplied, redeploy that version instead of deploying from the application directory. --- #### `dbos-cloud app secrets create` **Description:** Create a new secret associated with an application, or update an existing secret. Secrets are made available to your application as environment variables. You must redeploy your application for a change in its secrets to take effect. **Arguments:** - `[application-name]`: The name of the application for which to create or update secrets. - `-s, --name ` The name of the secret to create or update. - `-v, --value`: The value of the secret. --- #### `dbos-cloud app secrets import` **Description:** Import all environment variables defined in a `.env` file as secrets, updating them if they already exist. Allowed syntax for the `.env` file is described [here](https://dotenvx.com/docs/env-file), note that interpolation is supported but command substitution and encryption are currently not. Secrets are made available to your application as environment variables. You must redeploy your application for a change in its secrets to take effect. **Arguments:** - `[application-name]`: The name of the application for which to import secrets. - `-d, --dotenv ` Path to the `.env` file to import. --- #### `dbos-cloud app secrets list` **Description:** List all secrets associated with an application (only their names, not their values). **Arguments:** - `[application-name]`: The name of the application for which to list secrets. --- #### `dbos-cloud app secrets delete` **Description:** Delete a secret associated with an application. :::warning This action is irreversible ::: **Arguments:** - `[application-name]`: The name of the application for which to list secrets. - `-s, --name ` The name of the secret to delete. --- ### Organization Management Commands #### `dbos-cloud org list` **Description:** List users in your organization **Arguments:** - `--json`: Emit JSON output --- #### `dbos-cloud org invite` **Description:** Generate an organization secret with which to invite another user into your organization. Organization secrets are single-use and expire after 24 hours. **Arguments:** - `--json`: Emit JSON output --- #### `dbos-cloud org join` **Description:** Join your account to an organization. This gives you full access to the organization's resources. **Arguments:** - ``: The name of the organization you intend to join. - ``: An organization secret given to you by an organization admin. --- #### `dbos-cloud org rename` **Description:** Rename your organization. Only the organization admin (the original creator of the organization) can run this command. After running this command, [log out](#dbos-cloud-logout) and [log back in](#dbos-cloud-login) to refresh your local context. **Arguments:** - ``: The current name of your organization. - ``: The new name for your organization. :::info Applications belonging to organizations are hosted at the URL `https://-.cloud.dbos.dev/`, so renaming your organization changes your application URLs. The old URLs are no longer accessible. ::: --- #### `dbos-cloud org remove` **Description:** Remove a user from an organization. Only the organization admin (the original creator of the organization) can run this command. **Arguments:** - ``: The user to remove from your organization. --- ### Workflow Management Commands #### `dbos-cloud workflow list` **Description:** List workflows run by your application in JSON format ordered by recency (most recently started workflows last). **Arguments:** - `[application-name]`: The name of your application - `-l, --limit ` Limit the results returned (default: "10") - `-o, --offset ` Skip workflows from the results returned. - `-u, --workflowUUIDs ` Retrieve specific UUIDs - `-U, --user ` Retrieve workflows run by this user - `-s, --start-time ` Retrieve workflows starting after this timestamp (ISO 8601 format) - `-e, --end-time ` Retrieve workflows starting before this timestamp (ISO 8601 format) - `-S, --status ` Retrieve workflows with this status (`PENDING`, `SUCCESS`, `ERROR`, `MAX_RECOVERY_ATTEMPTS_EXCEEDED`, `ENQUEUED`, `DELAYED`, or `CANCELLED`) - `-v, --application-version ` Retrieve workflows with this application version - `-n, --name ` Retrieve functions with this name --- #### `dbos-cloud workflow queue list` **Description:** Lists all currently enqueued functions in JSON format ordered by recency (most recently enqueued functions last). **Arguments:** - `[application-name]`: The name of your application - `-n, --name ` Retrieve functions with this name - `-s, --start-time ` Retrieve functions starting after this timestamp (ISO 8601 format) - `-e, --end-time ` Retrieve functions starting before this timestamp (ISO 8601 format) - `-S, --status ` Retrieve workflows with this status (`PENDING`, `SUCCESS`, `ERROR`, `MAX_RECOVERY_ATTEMPTS_EXCEEDED`, `ENQUEUED`, `DELAYED`, or `CANCELLED`) - `-l, --limit ` Limit the results returned - `-o, --offset ` Skip functions from the results returned (for pagination) - `-q, --queue ` Retrieve functions run on this queue #### `dbos-cloud workflow cancel` **Description:** Cancel a workflow, setting its status to `CANCELLED`. If the workflow is currently executing, cancelling it preempts its execution (interrupting it at the beginning of its next step). If the workflow is enqueued, cancelling removes it from the queue. **Arguments:** - `[application-name]`: The name of your application - `-w, --workflowid`: The ID of the workflow to cancel. #### `dbos-cloud workflow resume` **Description:** Resume a workflow from its last completed step. You can use this to resume workflows that are cancelled or that have exceeded their maximum recovery attempts. You can also use this to start an `ENQUEUED` workflow, bypassing its queue. **Arguments:** - `[application-name]`: The name of your application - `-w, --workflowid`: The ID of the workflow to resume. --- ## Database Management #### Provisioning Database Instances Before you can deploy an application to DBOS Cloud, you must provision a Postgres database instance (server) for it. You must choose a database instance name, username and password. :::info * Both the database instance name and username must be 3 to 16 characters long and contain only lowercase letters, numbers and underscores. * The username must start with a letter. * The usernames `dbosadmin`, `dbos`, `postgres` and `admin` are reserved and cannot be used. * The database password must contain between 8 and 128 characters, and cannot contain the characters `/`, `"`, `@`, `'`, or whitespaces. ::: Run this command and choose your database password when prompted: ```shell dbos-cloud db provision -U ``` :::info A Postgres database instance (server) can host many independent databases used by different applications. Each application is deployed to an isolated database by default; you can configure this through the `app_db_name` field in `dbos-config.yaml`. ::: :::info If you forget your database password, you can always [reset it](./cloud-cli.md#dbos-cloud-db-reset-password). ::: To see a list of all provisioned instances and their statuses, run: ```shell dbos-cloud db list ``` To retrieve the status of a particular instance, run: ```shell dbos-cloud db status ``` #### Database Schema Management Every time you deploy an application to DBOS Cloud, it runs all migrations defined in your `dbos-config.yaml`. Sometimes, it may be necessary to manually perform schema changes on a cloud database, for example to recover from a schema migration failure. To make this easier, you can retrieve your cloud database connection URL by running: ```shell dbos-cloud db url ``` You can then use it to run locally any migration command (for example, a down-migration command in your schema migration tool) and it will execute on your cloud database. :::warning While it is occasionally necessary, be careful when manually changing the schema on a production database. ::: :::warning Be careful making breaking schema changes such as deleting or renaming a column—they may break active workflows running on a previous application version. ::: #### Destroying Database Instances To destroy a database instance, run: ```shell dbos-cloud db destroy ``` :::warning Take care—this will irreversibly delete all data in the database instance. ::: --- ## Deploying to DBOS Cloud :::info To use DBOS Cloud, please [contact sales](https://dbos.dev/contact). ::: Any application built with DBOS can be deployed to DBOS Cloud. DBOS Cloud is a serverless platform for durably executed applications. It provides: - [**Application hosting and autoscaling**](./application-management.md): Managed hosting of your application in the cloud, automatically scaling to millions of users. Applications are charged only for the CPU time they actually consume. - [**Managed workflow recovery**](./application-management.md): If a cloud executor is interrupted, crashed, or restarted, each of its workflows is automatically recovered by another executor. - [**Workflow and queue management**](./workflow-management.md): Dashboards of all active and past workflows and all queued tasks, including their status, inputs, outputs, and steps. Cancel, resume, or fork any workflow execution and manage the tasks in your distributed queues. ### Deploying Your App to DBOS Cloud ##### 1. Install the DBOS Cloud CLI
The Cloud CLI requires Node.js 20 or later.
Instructions to install Node.js **macOS or Linux** Run the following commands in your terminal: ```bash curl -o- https://raw.githubusercontent.com/nvm-sh/nvm/v0.40.1/install.sh | bash export NVM_DIR="$HOME/.nvm" [ -s "$NVM_DIR/nvm.sh" ] && \. "$NVM_DIR/nvm.sh" # This loads nvm nvm install 22 nvm use 22 ``` **Windows** Download Node.js 20 or later from the [official Node.js download page](https://nodejs.org/en/download) and install it. After installing Node.js, create the following folder: `C:\Users\%user%\AppData\Roaming\npm` (`%user%` is the Windows user on which you are logged in).
Run this command to install it.
```shell npm i -g @dbos-inc/dbos-cloud@latest ```
##### 2. Create a requirements.txt File
Create a `requirements.txt` file listing your application's dependencies. As DBOS Cloud uses OpenTelemetry to export your application's logs and traces, you should install the DBOS OpenTelemetry dependencies through `pip install dbos[otel]` before generating `requirements.txt`.
```shell pip freeze > requirements.txt ```
##### 3. Define a Start Command
Set the `start` command in the `runtimeConfig` section of your [`dbos-config.yaml`](../../../python/reference/configuration.md) to your application's launch command. If your application includes an HTTP server, configure it to listen on port 8000. To test that it works, try launching your application with `dbos start`.
```yaml runtimeConfig: start: - "fastapi run" ```
##### 4. Deploy to DBOS Cloud
Run this single command to deploy your application to DBOS Cloud!
```shell dbos-cloud app deploy ```
##### 1. Install the DBOS Cloud CLI
Run this command to install the Cloud CLI globally.
```shell npm i -g @dbos-inc/dbos-cloud@latest ```
As DBOS Cloud uses OpenTelemetry to export your application's logs and traces, you must install the DBOS OpenTelemetry dependencies into your application:
```shell npm i @dbos-inc/otel@latest ```
##### 2. Define a Start Command
Set the `start` command in the `runtimeConfig` section of your [`dbos-config.yaml`](../../../typescript/reference/configuration.md) to your application's launch command. If your application includes an HTTP server, configure it to listen on port 3000. To test that it works, try launching your application with `npx dbos start`.
```yaml runtimeConfig: start: - "npm start" ```
##### 3. Deploy to DBOS Cloud
Run this single command to deploy your application to DBOS Cloud!
```shell dbos-cloud app deploy ```
##### 1. Install the DBOS Cloud CLI
The Cloud CLI requires Node.js 20 or later.
Instructions to install Node.js **macOS or Linux** Run the following commands in your terminal: ```bash curl -o- https://raw.githubusercontent.com/nvm-sh/nvm/v0.40.1/install.sh | bash export NVM_DIR="$HOME/.nvm" [ -s "$NVM_DIR/nvm.sh" ] && \. "$NVM_DIR/nvm.sh" # This loads nvm nvm install 22 nvm use 22 ``` **Windows** Download Node.js 20 or later from the [official Node.js download page](https://nodejs.org/en/download) and install it. After installing Node.js, create the following folder: `C:\Users\%user%\AppData\Roaming\npm` (`%user%` is the Windows user on which you are logged in).
Run this command to install it.
```shell npm i -g @dbos-inc/dbos-cloud@latest ```
##### 2. Configure DBOS
Your DBOSContext [Config](../../../golang/reference/dbos-context.md) must be set with: - `DatabaseURL` (or your custom `pgxpool`) must point to an environment variable named `DBOS_SYSTEM_DATABASE_URL`
```go dbos.Config{ AppName: "dbos-starter", ApplicationVersion: "0.1.0", DatabaseURL: os.Getenv("DBOS_SYSTEM_DATABASE_URL"), } ```
##### 3. Provide a configuration file
In a file named `dbos-config.yaml`, set these fields: - `name`: your application name. Must match `AppName` in your `dbos.Config` - `language`: must be `go`
```yaml name: your-app-name language: go ```
##### 4. Deploy to DBOS Cloud
Finally, build your application under the name `main`, against linux/amd64, then run this command to deploy your application to DBOS Cloud! :::info DBOS Cloud will serve HTTP traffic on port 8080. Make sure to use that port when configuring web servers. :::
```shell GOOS=linux GOARCH=amd64 go build -o main . dbos-cloud app deploy ```
### DBOS Cloud How-Tos #### HTTP Serving & Port Numbers DBOS Cloud provides your application with an HTTPS URL and routes traffic to it. It expects applications to listen for HTTP requests on port 3000 (TypeScript), port 8000 (Python), or port 8080 (Go and Java). #### Environment Management We recommend you configure your application with environment variables. You can use [DBOS Cloud secrets](./secrets.md) to pass environment variables to your cloud application. You can even [import secrets](./secrets.md#importing-secrets) into DBOS Cloud from a `.env` file. #### Database Setup ##### System Database DBOS Cloud automatically constructs a system database for your application. You can access it at `DBOS_SYSTEM_DATABASE_URL`. You should not attempt to override or modify this URL. ##### Cloud-Provided Application Databases You may connect an application database to your application through DBOS Cloud. If you do, connection information is provided via the `DBOS_DATABASE_URL` environment variable. You may optionally direct DBOS Cloud to run migrations to set up your application's database schema by specifying migration commands in your `dbos-config.yaml` file: ```yaml migrate: - npx knex migrate:latest ``` DBOS Cloud will provide a database URL to your migrations through the `DBOS_DATABASE_URL` environment variable. --- ## Monitoring Your Applications The [DBOS Cloud Console](https://console.dbos.dev) provides several tools to monitor your applications. #### Logs You can view your application's logs from your application's cloud console page. Logs are paginated and ordered chronologically. ![Logs](./assets/cc-logs.png) #### Traces You can view traces for all your applications from the cloud console [traces page](https://console.dbos.dev/traces). You can filter traces by application, time, operation, type, and status. Traces are sorted chronologically and displayed hierarchically. You can click on a trace or span to see detailed information about it. ![Logs](./assets/cc-traces.png) #### Grafana Dashboard You can launch a Grafana dashboard for DBOS Cloud from the cloud console [dashboard page](https://console.dbos.dev/dashboard). ##### Time Selection In the top-right corner, the Grafana dashboard provides a time selector, defaulting to the last hour. You can change this setting to navigate to a different window of time. All of the panels are filtered for the selected time interval. ![Time picker](./assets/time_picker.png) Under the time selector you can find time series for active CPU milliseconds used by your apps and the counts of logs and traces, summarized for every minute. Counts of warnings, errors and fatal errors are color coded as yellow, red and purple respectively. These panes have a matched time axis. You can click and drag across an interesting region in the series to "zoom in". ![Series](./assets/timeseries.png) Zooming in will update the time selector and, therefore, all other panels. You can then use the time selector to "zoom back out" or use the `<` and `>` buttons to move backwards and forwards in time. ##### Grafana Logs and Traces On the left side, the Grafana dashboard provides a log view with entries generated by your applications arranged chronologically. The pane displays up to 1,000 most recent log records in the selected time period. Special entries for application lifetime events are colored grey and labeled as `[APP REGISTER]`, `[APP DEPLOY]` and so on. These are generated by DBOS Cloud automatically and not shown in the summarized log counts. You can click on a log record to browse additional metadata. In the example below, we see logs for an example app first getting registered, then undergoing schema migration, then getting deployed: ![Logs](./assets/log.png) Under the logs pane, there is an expandable panel of traces. Each row corresponds to a handler, workflow, transaction or step span. Each span is timestamped and decorated with duration in milliseconds, the IDs of the trace and workflow it belongs to, its execution status, and other information. ##### Filtering In the top-right corner of the Grafana dashboard, there are filtering selectors: ![Filters](./assets/filters.png) 1. you can select a single `Application Name` to filter for. Refresh the browser to update the list of names for a new app. 2. you can paste a specific `Trace ID` to only view logs and spans for that Trace. To clear, erase the text and press "return." 3. similar to Trace ID you can copy-paste a specific `Workflow UUID` to filter by that. It is cleared the same way as Trace ID. 4. you can select `Min Severity` to filter logs and traces for a specific severity level or higher 5. you can select `Application Version` to filter all data for a specific version of your app 6. you can select an `Executor ID` to only show data for a specific Micro VM :::tip When turning on these filters, the time window filter also still applies. You may see more data for your selection if you "zoom out" in time. ::: When using `Workflow UUID` use `_` to match any one character and `%` to match any string (SQL 'like' notation). This is useful for selecting groups of scheduled workflows. For example you can use a string like `sched%T19%` to match any scheduled workflows that ran at 7PM on any of the days in the selected time interval. `Search` also supports this syntax. ##### RAM Time, Requests and CPU Milliseconds The dashboard tracks the total RAM 512MB-Hours, Requests and active CPU Milliseconds for all your apps. These totals are updated every time you refresh your dashboard. They are applied against your DBOS Pricing tier's [execution time limit](https://www.dbos.dev/pricing). Please allow up to 5 minutes of delay between an event happening and the dashboard refresh showing it. The number of total CPU milliseconds since the start of the month is in orange. The light orange "selection" number to the right of it changes with the selected app(s) and time window. You can select a particular app or workload and see how much it contributes to your total. ![Execution Seconds](./assets/execution-seconds.png) :::tip It is possible for one or two small API calls to not consume a measurable amount of CPU ms. It is also normal for an idle app to use a negligible amount of CPU ms for periodic health checks and background tasks. For best results, run an example workflow of at least 10 API calls (the more the better). Observe how much CPU ms your example uses and extrapolate to your monthly expected usage. ::: ##### Memory and CPU Metrics MicroVM Metrics is an expandable panel under the Traces panel. This shows the number of running executors, their RAM and CPU usage over time. These plots are filtered for time and the selected app. If you're running multiple VMs, the CPU % and RAM usage plots show a separate colored line for each selected VM. ##### Dashboards and Organizations If you are part of a multi-user organization, your Grafana dashboard will show data for all applications deployed by all users in the organization. The log entries for application lifetime events (labeled as `[APP REGISTER]`, `[APP DEPLOY]` and so on) are annotated with the email address of the user performing each action. --- ## Export Logs and Traces This tutorial shows how to configure your DBOS Cloud application to export OpenTelemetry logs and traces to a third party observability service. If your service accepts the OTEL format, you can skip steps 1 and 2. Simply pass environment variables like `OTEL_EXPORTER_OTLP_HEADERS` as [app secrets](./secrets.md) (see [step 3](#3-set-the-datadog-api-key-to-your-apps-environment)) and then configure logs and traces endpoints as shown in [step 4](#4-configure-your-app-to-export-logs-and-traces-to-otel-contrib). Other services may require additional software. Here we use Datadog as an example. We connect by installing the otel-contrib package in the App VM at deployment time and configuring it with the Datadog API key to export data. :::info These steps require a [DBOS Pro or Enterprise](https://www.dbos.dev/pricing) subscription. ::: ### 1. Create a Custom VM Setup Script In your app directory (next to `dbos-config.yaml`) create the following script called `build.sh`. Make sure to set its permissions to execute. ```bash #!/bin/bash # Download and install otel-contrib in the MicroVM curl -L -O https://github.com/open-telemetry/opentelemetry-collector-releases/releases/download/v0.121.0/otelcol-contrib_0.121.0_linux_amd64.deb dpkg -i otelcol-contrib_0.121.0_linux_amd64.deb rm otelcol-contrib_0.121.0_linux_amd64.deb # Configure and enable it cat < /etc/otelcol-contrib/config.yaml receivers: otlp: protocols: grpc: http: endpoint: "0.0.0.0:4318" processors: batch: exporters: datadog: api: site: datadoghq.com #this URL depends on your datadog region key: ${DATADOG_API_KEY} #this is passed in a secret or env (see below) service: pipelines: metrics: receivers: [otlp] processors: [batch] exporters: [datadog] traces: receivers: [otlp] processors: [batch] exporters: [datadog] logs: receivers: [otlp] processors: [batch] exporters: [datadog] EOF systemctl restart otelcol-contrib systemctl enable otelcol-contrib ``` ### 2. Configure Your App to Run the Script on Deploy Add the build.sh script as a custom setup to your runtimeConfig in your `dbos-config.yaml`. See [Customizing MicroVM Setup](./application-management#customizing-microvm-setup) for more info. ```yaml runtimeConfig: setup: - "./build.sh" start: - npm run start #or your custom start command ``` ### 3. Set the Datadog API Key to Your App's Environment After registering your app, set the API key like so: ```bash dbos-cloud app register -d dbos-cloud app secrets create -s DATADOG_API_KEY -v 678... #your key value ``` The script we created in step 1 will read this value and pass it to `otel-contrib`. ### 4. Configure your App to Export Logs and Traces to otel-contrib In the app code, when creating the `DBOS` object, pass in the Logs and Traces endpoints like so: **Python** ```python from dbos import DBOSConfig config: DBOSConfig = { "name": "your-app-name", "application_version": "0.1.0", "otlp_traces_endpoints": [ "http://0.0.0.0:4318/v1/traces" ], #match the config in step 1 above "otlp_logs_endpoints": [ "http://0.0.0.0:4318/v1/logs" ] } DBOS(config=config) ``` **Typescript** ```typescript DBOS.setConfig({ "name": "your-app-name", "applicationVersion": "0.1.0", "otlpTracesEndpoints": [ "http://0.0.0.0:4318/v1/traces" ], "otlpLogsEndpoints": [ "http://0.0.0.0:4318/v1/logs" ] }); await DBOS.launch(); ``` ### 5. Add RAM if Needed, and Deploy! Depending on your app’s other memory usage, you may need to increase your RAM limit to make room for the otel-contrib process. ```bash dbos-cloud app update --executors-memory-mib 1024 dbos-cloud app deploy ``` Within a few minutes of deploying you should see your logs appear in Datadog. --- ## Workflow Retention Policies You can configure workflow history retention policies for your application from the Retention Policy page of the DBOS Console. These settings let you configure how long workflow history is retained in your application's [system database](../../../explanations/system-tables.md). This is useful for managing the database disk usage of workflow history. Retention policies only delete the history of completed workflows (workflows with status `SUCCESS`, `ERROR`, `CANCELLED`, or `MAX_RECOVERY_ATTEMPTS_EXCEEDED`); workflows that are still running, enqueued, or delayed are never deleted. Deleting a workflow's history also deletes its steps, inputs, outputs, messages, events, and streams. #### Time Threshold If a time threshold is set, workflow history is only retained for X hours after a workflow completes. History of workflows that completed more than X hours ago is automatically deleted. Time-based retention is disabled by default. #### Rows Threshold If the rows threshold is set, history is only retained for the X most recently completed workflows. History of older completed workflows is automatically deleted. By default, the rows threshold is set to 1M rows. You can set both a rows threshold and a time threshold. #### Global Timeout If a global timeout is set, any workflow that has not completed X hours after it was created (started or enqueued) is automatically cancelled. By default, the global timeout is disabled. --- ## Secrets and Environment Variables We recommend using _secrets_ to securely manage your application's secrets and environment variables in DBOS Cloud. Secrets are key-value pairs that are securely stored in DBOS Cloud and made available to your application as environment variables. Redeploy your application for newly created or updated secrets to take effect. ### Managing and Using Secrets You can create or update a secret using the Cloud CLI: ``` dbos-cloud app env create -s -v ``` :::info A few secret names are reserved and cannot be used. These are `DBOS_DATABASE_URL` and `DBOS_APP_HOSTNAME`. ::: For example, to create a secret named `API_KEY` with value `abc123`, run: ``` dbos-cloud app env create -s API_KEY -v abc123 ``` When you next redeploy your application, its environment will be updated to contain the `API_KEY` environment variable with value `abc123`. You can access it like any other environment variable: **Python** ```python key = os.environ['API_KEY'] # Value is abc123 ``` **Typescript** ```typescript const key = process.env.API_KEY; // Value is abc123 ``` Additionally, you can manage your application's secrets from the secrets page of the [cloud console](https://console.dbos.dev). ### Importing Secrets You can import the contents of a `.env` file as secrets. Allowed syntax for the `.env` file is described [here](https://dotenvx.com/docs/env-file). Note that interpolation is supported but command substitution and encryption are currently not. Import a `.env` file with the following command: ```shell dbos-cloud app env import -d ``` For example: ```shell dbos-cloud app env import -d .env ``` ### Listing Secrets You can list the names of your application's secrets with: ```shell dbos-cloud app env list ``` ### Deleting a Secret You can delete an environment variable with: ```shell dbos-cloud app env delete -s ``` --- ## Workflow Management ### Viewing Workflows Navigate to the workflows tab of your application's page on the DBOS Console to see a list of its workflows: This includes **all** your application's workflows: those currently executing, those enqueued for execution, those that have completed successfully, and those that have failed. You can filter by time, workflow ID, workflow name, and workflow status (for example, you can search for all failed workflow executions in the past day). Click on a workflow to see details, including its input and output: Click "Show Workflow Steps" to view the workflow's execution as a trace timeline (showing the workflow, its steps, and its child workflows and their steps). For example, here is the trace of a workflow that processes multiple tasks concurrently by enqueueing child workflows: You can manage individual workflows directly from the DBOS Console. ##### Cancelling Workflows You can cancel any workflow that has not completed: `PENDING`, `ENQUEUED`, or `DELAYED`. Cancelling a workflow sets its status to `CANCELLED`. If the workflow is currently executing, cancelling it preempts its execution (interrupting it at the beginning of its next step). If the workflow is enqueued or delayed, cancelling removes it from the queue. ##### Resuming Workflows You can resume any `ENQUEUED`, `DELAYED`, `CANCELLED` or `MAX_RECOVERY_ATTEMPTS_EXCEEDED` workflow. Resuming a workflow resumes its execution from its last completed step. If the workflow is enqueued, this bypasses the queue to start it immediately. ##### Forking Workflows You can start a new execution of a workflow by **forking** it from a specific step. To do this, open the workflow steps view, select a particular step, and click "Fork". When you fork a workflow, DBOS generates a new workflow with a new workflow ID, copies to that workflow the original workflow's inputs and all its steps up to the selected step, then begins executing the new workflow from the selected step. Forking a workflow is useful for recovering from outages in downstream services (by forking from the step that failed after the outage is resolved) or for "patching" workflows that failed due to a bug in a previous application version (by forking from the bugged step to an application version on which the bug is fixed). --- ## dbosctl CLI Reference `dbosctl` is a command-line client for the [Conductor API](./conductor-api.md). It manages workflows, queues, schedules, applications, and API keys against DBOS-hosted Conductor or a [self-hosted Conductor](../self-hosting/hosting-conductor.md), with the target selected by a named **profile**. The [`dbosctl sysdb`](#system-database-commands) commands are the exception: they manage the Postgres [system database](../../explanations/system-tables.md) directly, so they take a database URL rather than a profile. ### Installation The install script detects your platform, verifies the download against the release checksums, and installs `dbosctl` to the first writable of `/usr/local/bin`, `~/.local/bin`, or the current directory: ```shell curl -sSfL https://raw.githubusercontent.com/dbos-inc/dbos-ctl/main/install.sh | sh ``` Set `VERSION` to pin a release, or `BIN_DIR` to choose where it lands: ```shell curl -sSfL https://raw.githubusercontent.com/dbos-inc/dbos-ctl/main/install.sh \ | VERSION=v0.1.0 BIN_DIR=~/.local/bin sh ``` You can also [download an archive directly](https://github.com/dbos-inc/dbos-ctl/releases). Builds are published for Linux, macOS, and Windows on both amd64 and arm64, alongside a `checksums.txt`. The binaries are statically linked, so they run on any Linux distribution, Alpine included. If you already have a Go toolchain (1.24 or later), you can install from source instead: ```shell go install github.com/dbos-inc/dbos-ctl/cmd/dbosctl@latest ``` However you install it, `dbosctl version` reports what you have — a downloaded release prints its tag, and a `go install` prints the module version it was built from. ### Quick Start Against DBOS-hosted Conductor: ```shell dbosctl config set managed --managed # create a profile pointing at cloud.dbos.dev dbosctl login # log in through the device-authorization flow dbosctl whoami # confirm who you are logged in as dbosctl app list ``` Against a self-hosted Conductor with OIDC authentication, name the issuer and client ID the deployment is configured with, then log in as you would against the managed service: ```shell dbosctl config set selfhosted --url https://conductor.example.com \ --issuer https://auth.example.com/realms/dbos --client-id dbos-cli dbosctl login --profile selfhosted dbosctl app list --profile selfhosted ``` Against a self-hosted Conductor running without authentication: ```shell dbosctl config set local --url http://localhost:8090 dbosctl app list --profile local ``` ### Profiles A profile is a named bundle of connection settings: which Conductor to talk to, how to authenticate, and the default organization and application. Profiles are stored in `config.yaml` under your OS configuration directory — `~/.config/dbos/config.yaml` on Linux, `~/Library/Application Support/dbos/config.yaml` on macOS. A profile must target either DBOS-hosted Conductor (`--managed`) or a self-hosted one (`--url`); the two are mutually exclusive. There are three common shapes: | Shape | How to create it | Authentication | Identity | | --- | --- | --- | --- | | DBOS-hosted | `dbosctl config set --managed` | User JWT or `dbos_` API key | Your real user, or none for an API key | | Self-hosted with OIDC | `dbosctl config set --url --issuer --client-id ` | User JWT or `dbos_` API key | Your real user, or none for an API key | | Self-hosted, no auth | `dbosctl config set --url ` | None | Always `local` | The two authenticated shapes accept the same credentials: Conductor tells a user JWT from an API key by the key's `dbos_` prefix, not by how it is deployed. What differs is where the OIDC settings come from — `--managed` derives them, self-hosted needs them spelled out. `--managed` points the profile at `cloud.dbos.dev` and derives everything else — the `/conductor` base URL, bearer authentication, and the OIDC tenant — automatically. Passing `--issuer`/`--client-id` implies bearer authentication, so `--auth` is only needed in the uncommon case of a self-hosted Conductor you reach with a `dbos_` API key but no OIDC login: pass `--auth bearer`. Because an API key carries no user identity, give that profile an `--org` as well. ### Authentication ```shell dbosctl login # OIDC device flow against the profile's issuer; stores a token dbosctl logout # discard the stored token for the current profile ``` `login` runs the [device-authorization flow](https://www.rfc-editor.org/rfc/rfc8628): it prints a URL and a code, you approve in a browser, and the resulting token is written to `credentials.json` (mode `0600`) next to `config.yaml`, keyed by profile. Tokens are refreshed automatically on expiry when the issuer returns a refresh token. There are two ways to bypass the login flow: - **`DBOS_TOKEN`** — a bearer token used as-is for a single invocation. - **API keys** — a `dbos_…` key from [`dbosctl api-key create`](#dbosctl-api-key-create) or the console, supplied through `DBOS_TOKEN`. Keys authenticate machine-to-machine calls such as `app list`, but carry no user identity, so `dbosctl whoami` still requires a user login. ### Configuration Precedence Each setting is resolved **flag → environment variable → profile**, so a flag always wins and the profile is the fallback: | Setting | Flag | Environment variable | | --- | --- | --- | | Profile | `--profile` | `DBOS_PROFILE` | | Conductor URL | `--url` | `DBOS_URL` | | Organization | `--org` | `DBOS_ORG` | | Application | `-a`, `--app` | `DBOS_APP` | | Bearer token | — | `DBOS_TOKEN` | | System database ([`sysdb`](#system-database-commands) only) | `-D`, `--db-url` | `DBOS_SYSTEM_DATABASE_URL` | | Output format | `-o`, `--output` | — | Flags are scoped to the command that uses them, so pass them **after** the command name (`dbosctl app list --org acme`). Each command's `--help` lists only the flags it honors. ### Common Flags These flags are accepted by most commands and are not repeated in the reference below: - `--profile `: Config profile to use. Overrides `$DBOS_PROFILE`. - `--url `: Conductor base URL. Overrides `$DBOS_URL` and the profile. - `--org `: Organization. Overrides `$DBOS_ORG` and the profile. - `-a, --app `: Application name, for application-scoped commands. Overrides `$DBOS_APP` and the profile. - `-o, --output `: Output format — `table` (default), `json`, or `ids`. The [`dbosctl sysdb`](#system-database-commands) commands accept none of them except `-o`. They do not talk to Conductor, so there is no profile, organization, or application to name. `sysdb reset` and `sysdb rename-application` print row counts, so they honor `-o` like every other command that prints data — `table` or `json`, but not `ids`. (`sysdb reset` has an `--app` flag of its own, read from the command line only.) ### Output Output is human-readable tables by default. `-o json` emits the raw API shape for scripting; it is never truncated or reprojected, so `dbosctl app list -o json` is exactly the JSON array the [Conductor API](./conductor-api.md) returned. ```shell dbosctl app list # aligned table dbosctl app list -o json # raw JSON array dbosctl whoami -o json # raw user profile ``` Detail views — `workflow get`, `queue get`, and `schedule get` — include an `applicationName` row naming the application that owns the object, when it has one. Commands with a natural identifier also accept `-o ids`, which prints one ID per line for piping. A literal `-` in place of arguments reads IDs from stdin: ```shell dbosctl workflow list -a myapp --status PENDING -o ids | dbosctl workflow cancel -a myapp - ``` ### Exit Codes | Code | Meaning | | --- | --- | | `0` | Success | | `1` | General error | | `2` | Usage error (bad flags or arguments) | | `3` | Authentication required (HTTP 401) — run `dbosctl login` | | `4` | Not found (HTTP 404) | | `130` | Interrupted (Ctrl-C) | ### Commands That Need a Running Application Conductor answers some commands itself and forwards the rest to your application over the websocket its executors hold open, as described in [How Operations Are Served](./conductor-api.md#how-operations-are-served). The split follows the resource, not the verb, so it cuts across the reference below: | | Commands | | --- | --- | | Forwarded to your application | All `workflow`, `queue`, and `schedule` commands — reads as much as mutations — plus `app versions` and `app set-version` | | Answered by Conductor | Everything else except `sysdb`: the rest of `app`, plus `api-key`, `permission`, `login`, `logout`, `whoami`, and `config` | | Answered by neither | The [`sysdb`](#system-database-commands) commands, which open the system database themselves | Commands in the first group fail if the application has no healthy executor connected, so a failure there usually means your application is not running rather than that nothing matched. ### Authentication Commands #### `dbosctl login` **Description:** Logs in to the current profile's Conductor using the OIDC device-authorization flow, storing the resulting token for later commands. --- #### `dbosctl logout` **Description:** Removes the stored login for the current profile. --- #### `dbosctl whoami` **Description:** Shows the logged-in identity. On a no-auth self-hosted target this always reports `local`. An API key carries no user identity, so this command requires a user login. --- ### Profile Commands #### `dbosctl config list` **Description:** Lists all profiles, marking the current one. --- #### `dbosctl config show` **Description:** Shows one profile's settings. **Arguments:** - `[profile]`: (Optional) The profile to show. Defaults to the current profile. --- #### `dbosctl config use` **Description:** Sets the current profile, used by any command that does not pass `--profile`. **Arguments:** - ``: The profile to make current. --- #### `dbosctl config set` **Description:** Creates or updates a profile. Only the flags you pass are changed; fields you do not name are left as they were. **Arguments:** - ``: The profile to create or update. - `--managed`: Make this a DBOS-hosted Conductor profile (production domain `cloud.dbos.dev`). Mutually exclusive with `--url`. - `--url `: Base URL of a self-hosted Conductor. Mutually exclusive with `--managed`. - `--issuer `: OIDC issuer URL. Implies bearer authentication. - `--client-id `: OIDC client ID. Implies bearer authentication. - `--audience `: OIDC audience, for bearer profiles that require it. - `--auth `: Force bearer authentication without OIDC, for a profile that authenticates with a `dbos_` API key only. The accepted value is `bearer`. - `--org `: Organization. Only needed when it cannot be derived from a login, as with an API-key-only profile. - `--app `: Default application for this profile. --- ### Application Commands #### `dbosctl app list` **Description:** Lists the applications in the organization. --- #### `dbosctl app get` **Description:** Shows one application's details. **Arguments:** - ``: The application's name. --- #### `dbosctl app register` **Description:** Registers an application with Conductor. The name must match the application name in your DBOS configuration — see [Connecting To Conductor](../overview.md#connecting-to-conductor). **Arguments:** - ``: The application's name. - `--private-mode`: Register the application in private mode, so it does not send workflow payload data — inputs, outputs, and events — to Conductor. Omit the flag to take Conductor's default; it can be changed later with [`dbosctl app update`](#dbosctl-app-update). --- #### `dbosctl app update` **Description:** Updates an application's tuning settings. Only the flags you pass are changed. **Arguments:** - ``: The application's name. - `--executor-timeout-secs `: Seconds before an idle executor is considered gone. - `--global-timeout-ms `: Global workflow timeout, in milliseconds. - `--gc-rows-threshold `: Number of most recently completed workflows whose history is kept; history of older completed workflows is garbage-collected. See [Workflow Retention Policies](../retention.md). - `--gc-time-threshold-ms `: Time, in milliseconds, after a workflow completes before its history is garbage-collected. - `--private-mode`: Whether the application is in private mode, in which it does not send workflow payload data — inputs, outputs, and events — to Conductor. Pass `--private-mode=false` to turn it back off. --- #### `dbosctl app delete` **Description:** Deletes an application. Prompts for confirmation when run interactively. **Arguments:** - ``: The application's name. - `--force`: Skip the confirmation prompt. Required when running non-interactively. --- #### `dbosctl app versions` **Description:** Lists an application's versions. **Arguments:** - ``: The application's name. --- #### `dbosctl app set-version` **Description:** Sets an application's latest version. **Arguments:** - ``: The application's name. - ``: The version to mark as latest. --- #### `dbosctl app executors` **Description:** Lists the executors currently connected to Conductor for an application. **Arguments:** - ``: The application's name. --- #### `dbosctl app metrics` **Description:** Lists an application's metrics over a time window. **Arguments:** - ``: The application's name. - `--since `: Report the window ending now and starting this long ago. Defaults to `24h`. --- ### Workflow Commands Workflow commands are application-scoped: pass `-a/--app`, or set `DBOS_APP`, or give the profile a default application. The `wf` alias is accepted in place of `workflow`. :::info `resume` is stricter than every other command here: it needs an executor running the application's **latest** version, not merely a healthy one. A deployment that is connected but still rolling out can therefore resume nothing, while everything else works. ::: #### `dbosctl workflow list` **Description:** Lists workflows matching the given filters. Without `--limit`, this returns *all* matching workflows, so that `dbosctl workflow list -o ids | dbosctl workflow cancel -` acts on the whole set. Pass `--limit`/`--offset` to bound or page the results. **Arguments:** - `-s, --status `: Filter by workflow status. Repeatable. - `-n, --name `: Filter by workflow name. Repeatable. - `--id `: Filter by workflow ID. Repeatable. - `-u, --user `: Filter by user. Repeatable. - `--queue `: Filter by queue name. Repeatable. - `--queued`: Return only workflows currently on a queue. - `--app-version `: Filter by application version. Repeatable. - `--since `: Window start, as RFC 3339 or a duration such as `1h`. - `--until `: Window end, as RFC 3339 or a duration such as `1h`. - `--desc`: Sort newest first. - `-l, --limit `: Maximum number of results. - `--offset `: Number of results to skip. --- #### `dbosctl workflow get` **Description:** Shows a workflow's details, including its input and output. **Arguments:** - ``: The workflow's ID. --- #### `dbosctl workflow steps` **Description:** Lists a workflow's steps. **Arguments:** - ``: The workflow's ID. --- #### `dbosctl workflow events` **Description:** Lists the events a workflow has set. **Arguments:** - ``: The workflow's ID. --- #### `dbosctl workflow cancel` **Description:** Cancels one or more workflows. Cancelling sets a workflow's status to `CANCELLED`: if it is executing, its execution is preempted at the start of its next step; if it is enqueued, it is removed from the queue. **Arguments:** - `...`: One or more workflow IDs. A literal `-` reads IDs from stdin, one per line. - `--children`: Also cancel the workflows' child workflows. --- #### `dbosctl workflow resume` **Description:** Resumes one or more workflows from their last completed step. If a workflow is enqueued, resuming it bypasses the queue and starts it immediately. **Arguments:** - `...`: One or more workflow IDs. A literal `-` reads IDs from stdin. - `--queue `: Resume onto this queue. --- #### `dbosctl workflow fork` **Description:** Forks a workflow into a new execution starting from a chosen step, copying the original workflow's inputs and its completed steps up to that point. Prints the new workflow's ID. **Arguments:** - ``: The workflow to fork. - `--start-step `: The step to fork from. - `--new-id `: ID for the forked workflow. Generated if not supplied. - `--app-version `: Application version to run the fork on. - `--queue `: Enqueue the fork onto this queue. --- #### `dbosctl workflow delete` **Description:** Deletes one or more workflows and their recorded history. **Arguments:** - `...`: One or more workflow IDs. A literal `-` reads IDs from stdin. - `--children`: Also delete the workflows' child workflows. --- ### Queue Commands #### `dbosctl queue list` **Description:** Lists an application's queue definitions. --- #### `dbosctl queue get` **Description:** Shows one queue's details. **Arguments:** - ``: The queue's name. --- ### Schedule Commands #### `dbosctl schedule list` **Description:** Lists an application's scheduled workflows. --- #### `dbosctl schedule get` **Description:** Shows one schedule's details. **Arguments:** - ``: The schedule's name. --- #### `dbosctl schedule pause` **Description:** Pauses a schedule, so that it stops starting new workflows. **Arguments:** - ``: The schedule's name. --- #### `dbosctl schedule resume` **Description:** Resumes a paused schedule. **Arguments:** - ``: The schedule's name. --- #### `dbosctl schedule trigger` **Description:** Runs a scheduled workflow immediately, out of band. Prints the started workflow's ID. **Arguments:** - ``: The schedule's name. --- #### `dbosctl schedule backfill` **Description:** Replays a schedule across a past time window, starting one workflow for each occurrence the schedule would have fired. Prints the started workflow IDs. **Arguments:** - ``: The schedule's name. - `--since `: Window start, as RFC 3339 or a duration such as `1h`. - `--until `: Window end, as RFC 3339 or a duration such as `1h`. --- ### API Key Commands The `token` and `apikey` aliases are accepted in place of `api-key`. #### `dbosctl api-key list` **Description:** Lists the organization's API keys. Secrets are not shown. --- #### `dbosctl api-key create` **Description:** Creates an API key and prints its secret. **The secret is shown once and cannot be retrieved afterwards.** By default the key is unscoped; narrow it with `--app` and `--permission`. See [Permissions and API Keys](../permissions.md). **Arguments:** - ``: A name for the key. - `--app `: Scope the key to these applications. Repeatable; defaults to all applications. - `--permission `: Grant these permissions, for example `application.read`. Repeatable. --- #### `dbosctl api-key rename` **Description:** Renames an API key. The secret is unchanged, so anything already using the key keeps working. **Arguments:** - ``: The key's current name. - ``: The key's new name. Fails if another key in the org already has it. --- #### `dbosctl api-key delete` **Description:** Deletes an API key, revoking it immediately. **Arguments:** - ``: The key's name. --- #### `dbosctl permission list` **Description:** Lists the permissions that can be granted to an API key or a role. --- ### System Database Commands `dbosctl sysdb` groups the commands that open a database instead of calling Conductor. They connect to a Postgres (or CockroachDB) [system database](../../explanations/system-tables.md) directly, so they take a database URL rather than a profile, and accept none of the [common flags](#common-flags) that resolve one. The system schema is shared by every DBOS SDK and the migrations are built into the `dbosctl` binary. **Shared arguments:** - `-D, --db-url `: The system database URL. Overrides `$DBOS_SYSTEM_DATABASE_URL`. - `--schema `: The schema holding the DBOS system tables. Defaults to `dbos`. Both are defined on `sysdb` itself, so every subcommand takes them: ```shell dbosctl sysdb migrate -D postgres://user:password@host:5432/dbos_sys DBOS_SYSTEM_DATABASE_URL=postgres://user:password@host:5432/dbos_sys dbosctl sysdb migrate ``` --- #### `dbosctl sysdb migrate` **Description:** Creates or upgrades the DBOS [system database](../../explanations/system-tables.md), applying every migration the schema is missing and creating the database and schema if they do not exist yet. By default, a DBOS application automatically creates these on startup. However, in production environments, a DBOS application may not run with sufficient privilege to create databases or tables. In that case, the `migrate` command can be run with a privileged user to create all DBOS database tables. After creating the DBOS database tables with this command, a DBOS application can run with minimum permissions, requiring only access to the DBOS schema in the application and system databases. Use the `-r/--app-role` flag to grant a role access to that schema. `migrate` is safe to re-run. Migrations already recorded are skipped, so a database that is up to date is left alone. **Arguments:** - `-r, --app-role `: The role your DBOS application runs as. It is granted the minimum permissions needed to use the DBOS schema. - `--no-listen-notify`: Leave out the triggers that fire `pg_notify`. See [LISTEN/NOTIFY](#listennotify) below. - `--print-migrations `: Instead of running the migrations, print their SQL to standard output — `all` for a fresh database, or a migration number to upgrade an existing one. - `--print-user-role`: Instead of running them, print the SQL statements granting `--app-role` access to the DBOS system tables. - `--cockroach`: Render the printed SQL for CockroachDB. Print mode only — see [CockroachDB](#cockroachdb) below. :::info SDK migration config settings After running `dbosctl sysdb migrate`, configure your application not to alter the system schema on startup where that option exists: **Python** ```python config: DBOSConfig = { "name": "my-app", "system_database_url": os.environ["DBOS_SYSTEM_DATABASE_URL"], "run_migrations": False, } ``` **TypeScript** ```typescript DBOS.setConfig({ name: "my-app", systemDatabaseUrl: process.env.DBOS_SYSTEM_DATABASE_URL, runMigrations: false, }); await DBOS.launch(); ``` **Java** ```java DBOSConfig config = DBOSConfig.defaults("my-app") .withDatabaseUrl(System.getenv("DBOS_SYSTEM_JDBC_URL")) .withMigrate(false); ``` ::: ##### Printing the SQL If your database is managed by a DBA, or its DDL goes through review, the print modes emit SQL for someone else to apply: ```shell dbosctl sysdb migrate --print-migrations all > migrations.sql dbosctl sysdb migrate --print-user-role -r my_app_role > grants.sql ``` ##### LISTEN/NOTIFY If your system database sits behind a connection pooler in transaction mode, pass `--no-listen-notify` to generate a system schema that doesn't use `pg_notify`: ```shell dbosctl sysdb migrate -D postgres://user:password@host:5432/dbos_sys --no-listen-notify ``` ##### CockroachDB Live migration can detect when the system database is running on CockroachDB. Since a printed script cannot detect the database engine, use `--cockroach` to specify Cockroach compatible system database migrations. ```shell dbosctl sysdb migrate --print-migrations all --cockroach > migrations.sql ``` --- #### `dbosctl sysdb reset` **Description:** Empties the DBOS system database, deleting all the rows in the DBOS system tables. The schema itself is left migrated and immediately usable, so the database does not have to be provisioned again. Prompts for confirmation when run interactively. **Arguments:** - `-a, --app `: Empty only the specified application's rows, for a [shared system database](../../explanations/sharing-a-system-database.md). Unlike the `--app` argument used by Conductor commands, this argument is only read from the command line — never from `$DBOS_APP` or a profile. - `--drop-database`: Drop the whole database instead of emptying the DBOS tables. Cannot be combined with `--app` or `--schema`. - `--force`: Skip the confirmation prompt. Required when running non-interactively. - `-o, --output `: Output format for the row counts — `table` (default) or `json`. When not using `--drop-database`, per-table row counts are written to standard output, with progress on standard error, so a scripted reset can read what it removed without parsing log lines. --- #### `dbosctl sysdb rename-application` **Description:** Transfers ownership of everything in the system database from an application's old name to its new one. The `rename-app` alias is accepted in place of `rename-application`. Prints the number of rows transferred, by table. Prompts for confirmation when run interactively. :::warning **Stop the application being renamed before running this.** Nothing here locks it out, and a running one keeps creating new records using its old name. ::: **Arguments:** - `-f, --from `: The application's previous name. Omit to only adopt unclaimed rows, which then requires `--adopt-unclaimed-rows`. - `-t, --to `: The application that ends up owning the rows. Required. Must not be only whitespace. - `--adopt-unclaimed-rows`: Also transfer rows no application owns (`application_name` is null). - `--batch-size `: Completed workflows and steps transferred per transaction. Defaults to 10000. - `--force`: Skip the confirmation prompt and the `--to` name checks. Required when running non-interactively. - `-o, --output `: Output format for the row counts — `table` (default) or `json`. By default, `rename-application` expects the `--to` name to be unused and to be DBOS Conductor compatible (between 3 and 256 characters, only numbers, lowercase ASCII letters, hyphens, and underscores). For interactive shells, `rename-application` will confirm with the user before renaming if either of these conditions are false. This check can be overridden with the `--force` argument. ##### Schema versions `reset` and `rename-application` both refuse a schema migrated past what your `dbosctl` version knows. Typically, in this case you just need to install a more recent version of `dbosctl`. `rename-application` also refuses a schema that predates the addition of application names. --- ### Other Commands #### `dbosctl version` **Description:** Prints the version, build metadata, and platform. Also available as `dbosctl --version`. --- #### `dbosctl completion` **Description:** Generates a shell autocompletion script. Run `dbosctl completion --help` for installation instructions for your shell. **Arguments:** - ``: The shell to generate a script for: `bash`, `zsh`, `fish`, or `powershell`. --- ## Workflow Retention Policies(Conductor) If you are using [Conductor](./overview.md), you can configure workflow history retention policies for your application from the Retention Policy page of the DBOS Console. These settings let you configure how long workflow history is retained in your application's [system database](../explanations/system-tables.md). This is useful for managing the database disk usage of workflow history. Retention policies only delete the history of completed workflows (workflows with status `SUCCESS`, `ERROR`, `CANCELLED`, or `MAX_RECOVERY_ATTEMPTS_EXCEEDED`); workflows that are still running, enqueued, or delayed are never deleted. Deleting a workflow's history also deletes its steps, inputs, outputs, messages, events, and streams. Retention runs in the background and deletes history in batches. If multiple applications [share a system database](../explanations/sharing-a-system-database.md), retention policies apply to the entire system database, including workflows owned by other applications. The most restrictive policy configured by any of these applications therefore applies to all of them, so we recommend configuring the same retention policies for every application that shares a system database. The global timeout applies only to workflows owned by the application for which it is configured. #### Time Threshold If a time threshold is set, workflow history is only retained for X hours after a workflow completes. History of workflows that completed more than X hours ago is automatically deleted. Time-based retention is disabled by default. #### Rows Threshold If the rows threshold is set, history is only retained for the X most recently completed workflows. History of older completed workflows is automatically deleted. Rows-based retention is disabled by default. You can set both a rows threshold and a time threshold. #### Global Timeout If a global timeout is set, any workflow that has not completed X hours after it was created (started or enqueued) is automatically cancelled. By default, the global timeout is disabled. --- ## Deploying Conductor on Kubernetes :::info Self-hosted Conductor is released under a [proprietary license](https://www.dbos.dev/conductor-license). Self-hosting Conductor for commercial or production use requires a [license key](./hosting-conductor.md#licensing). ::: ### Overview This guide covers deploying DBOS Conductor and the DBOS Console on Kubernetes. It maps the components and production requirements from the [Self-Hosting Guide](./hosting-conductor.md) onto Kubernetes resources, then walks through a full deployment on AWS EKS. The Kubernetes manifests are portable to any conformant cluster. --- ### Deployments **Database** — Conductor's [Postgres database](./hosting-conductor.md#components) runs outside the cluster (this guide uses RDS). **Conductor** — A single-container [Deployment](https://kubernetes.io/docs/concepts/workloads/controllers/deployment/), not a [StatefulSet](https://kubernetes.io/docs/concepts/workloads/controllers/statefulset/), because all state lives in Postgres. It listens on port 8090 and reads its [required environment variables](./hosting-conductor.md#conductor) from Secrets. To run multiple replicas for [high availability](./hosting-conductor.md#high-availability), each pod must advertise its own address; `conductor.yaml` below sets `DBOS__ADVERTISE_ADDRESS` from the pod IP. **Console** — A single-container Deployment listening on port 8080, behind a Service that publishes port 80. Its `DBOS_CONDUCTOR_URL` is the in-cluster Conductor Service address, `conductor.dbos.svc.cluster.local:8090`. :::info Updating Conductor Conductor is architecturally **out-of-band** — it is not on the critical path of your application. To upgrade, update the container image tag in `conductor.yaml` and `console.yaml`, (`latest` by default) then `kubectl rollout restart`. Prefer updating both Conductor and the console together. Applications seamlessly reconnect to the new Conductor version with no impact on their availability. ::: :::info Register applications After deploying Conductor and Console, [register your application, and generate an API key](../overview.md#connecting-to-conductor). The application connects to Conductor via WebSocket using this API key and the Conductor URL. With the [Ingress](#ingress) below, that URL is your Ingress hostname plus the `/conductor-api` prefix: ```bash DBOS_CONDUCTOR_KEY= DBOS_CONDUCTOR_URL=wss:///conductor-api ``` Because this is a `wss://` connection, your application verifies the Ingress TLS certificate. ::: ### Authentication Conductor supports OAuth 2.0 with any OIDC-compliant provider. See [Security](./hosting-conductor.md#security) for the provider setup and environment variables. :::warning Conductor performs **no authentication** unless OAuth is enabled, so anyone who can reach the Ingress has full admin access. Configure OAuth before exposing this deployment to any untrusted network. ::: When configuring your OAuth provider, the callback URL and allowed web origin are your Ingress hostname (`https:///oauth/callback` and `https://`). The OAuth settings are not secrets, so they can be set directly in the Deployment manifests — Conductor and the Console each need their own set. ### Ingress All external traffic enters through an Ingress that meets the [reverse proxy requirements](./hosting-conductor.md#reverse-proxy-and-tls) and routes by path: `/conductor-api/...` to Conductor, everything else to the Console. This guide uses [ingress-nginx](https://kubernetes.github.io/ingress-nginx/), but any ingress controller meeting those requirements will work. The `ingress.yaml` below defines the routing it must implement. Set idle timeouts to 3600 seconds on both the ingress controller and the cloud load balancer in front of it (for example, AWS ELB). ### Security Best Practices **Secret management** — Store the [Conductor secrets](./hosting-conductor.md#secrets) as Kubernetes Secrets and inject them via `secretKeyRef`. For Git-safe storage, encrypt with [Sealed Secrets](https://github.com/bitnami-labs/sealed-secrets), [SOPS](https://github.com/getsops/sops), or a cloud-native secrets manager (AWS Secrets Manager, [Vault](https://developer.hashicorp.com/vault/docs/platform/k8s/vso), etc.). **Network policies** — Apply a default-deny ingress policy to the namespace, then add explicit allow rules for each pod. If Conductor and Console are co-located, allow traffic from the Console to Conductor on port 8090. Keep [outbound HTTPS](./hosting-conductor.md#network-access) open from the Conductor pod for license validation, which means its nodes need a route to the internet, such as a NAT gateway for private subnets. **RBAC** — Restrict which ServiceAccounts can read Secrets in the namespace. Conductor credentials (database URLs, license key, API key) should only be accessible to the pods that need them. --- ### Walkthrough (AWS EKS) **EKS (AWS)** In addition to DBOS Conductor and the DBOS Console, the infrastructure includes the following components: | Component | Role | |-----------|------| | **RDS** | Database for Conductor operating state | | **Reverse Proxy (Nginx Ingress)** | TLS termination, path-based routing, WebSocket support | | **Sealed Secrets** | Encrypts secrets at rest; decrypts them in-cluster |
Set environment variables Set these variables before proceeding — replace the placeholder values with your own: ```bash # Your AWS account ID (12-digit number) AWS_ACCOUNT_ID=123456789012 # AWS region for all resources AWS_REGION=us-west-2 # PostgreSQL admin password (used for the RDS master user) POSTGRES_PASSWORD='choose-a-secure-password' # Password for the Conductor database role CONDUCTOR_ROLE_PASSWORD='choose-another-secure-password' # Conductor license key (from DBOS Console or sales) CONDUCTOR_LICENSE_KEY='your-license-key' ```
#### Infrastructure
CLI tools required on your workstation | Tool | Purpose | Install | |------|---------|---------| | **AWS CLI** | AWS account access | [Install guide](https://docs.aws.amazon.com/cli/latest/userguide/getting-started-install.html) | | **eksctl** | Create and manage EKS clusters | [Install guide](https://eksctl.io/installation/) | | **kubectl** | Interact with Kubernetes | Included with eksctl, or [install separately](https://kubernetes.io/docs/tasks/tools/) | | **Helm** | Install cluster add-ons (Ingress, Sealed Secrets) | `brew install helm` or [Install guide](https://helm.sh/docs/intro/install/) | | **kubeseal** | Encrypt Kubernetes secrets | `brew install kubeseal` or [Install guide](https://github.com/bitnami-labs/sealed-secrets#kubeseal) | | **openssl** | Generate self-signed TLS certificate | Pre-installed on macOS/Linux | Verify your AWS credentials are configured: ```bash aws sts get-caller-identity ```
**DBOS Conductor License Key** Obtain a development license key from the [DBOS Console](https://console.dbos.dev/settings/license-key) or [contact DBOS sales](https://www.dbos.dev/contact) for a pro license key. You can follow this guide with a development license key for evaluation, but you will be limited to one executor per application. **Create an EKS Cluster** Create a managed EKS cluster with two nodes. This takes approximately 15 minutes.
Create EKS cluster ```bash eksctl create cluster \ --name dbos-conductor \ --region $AWS_REGION \ --version 1.32 \ --nodegroup-name default \ --node-type t3.medium \ --nodes 2 \ --managed ``` `eksctl` automatically: - Creates a VPC with public and private subnets - Configures the [Amazon VPC CNI](https://docs.aws.amazon.com/eks/latest/userguide/managing-vpc-cni.html), which supports NetworkPolicy enforcement - Sets up your `~/.kube/config` to point at the new cluster Once complete, verify the cluster is ready: ```bash kubectl get nodes ``` You should see two nodes in `Ready` status: ``` NAME STATUS ROLES AGE VERSION ip-192-168-xx-xx.us-west-2.compute.internal Ready 2m v1.32.x ip-192-168-xx-xx.us-west-2.compute.internal Ready 2m v1.32.x ```
**Create a Namespace** All resources in this guide are deployed to a dedicated `dbos` namespace: ```bash kubectl create namespace dbos ``` **Provision an RDS PostgreSQL Instance**
RDS provisioning commands Find the VPC and private subnets that `eksctl` created: ```bash # Get the VPC ID VPC_ID=$(aws ec2 describe-vpcs \ --filters "Name=tag:alpha.eksctl.io/cluster-name,Values=dbos-conductor" \ --query "Vpcs[0].VpcId" --output text --region $AWS_REGION) echo "VPC: $VPC_ID" # Get the private subnets PRIVATE_SUBNETS=($(aws ec2 describe-subnets \ --filters "Name=vpc-id,Values=$VPC_ID" \ "Name=tag:aws:cloudformation:logical-id,Values=SubnetPrivate*" \ --query "Subnets[*].SubnetId" --output text --region $AWS_REGION)) echo "Private subnets: ${PRIVATE_SUBNETS[@]}" ``` Create a DB subnet group from the private subnets: ```bash aws rds create-db-subnet-group \ --db-subnet-group-name dbos-conductor-db \ --db-subnet-group-description "DBOS Conductor RDS subnets" \ --subnet-ids "${PRIVATE_SUBNETS[@]}" \ --region $AWS_REGION ``` Create a security group that allows PostgreSQL access from the EKS nodes: ```bash # Get the EKS cluster security group EKS_SG=$(aws ec2 describe-security-groups \ --filters "Name=vpc-id,Values=$VPC_ID" \ "Name=tag:aws:eks:cluster-name,Values=dbos-conductor" \ --query "SecurityGroups[0].GroupId" \ --output text --region $AWS_REGION) echo "EKS SG: $EKS_SG" # Create a security group for RDS RDS_SG=$(aws ec2 create-security-group \ --group-name dbos-conductor-rds \ --description "Allow PostgreSQL from EKS nodes" \ --vpc-id $VPC_ID \ --query "GroupId" --output text --region $AWS_REGION) echo "RDS SG: $RDS_SG" # Allow inbound PostgreSQL from EKS nodes aws ec2 authorize-security-group-ingress \ --group-id $RDS_SG \ --protocol tcp --port 5432 \ --source-group $EKS_SG \ --region $AWS_REGION ``` Create the RDS instance: ```bash aws rds create-db-instance \ --db-instance-identifier dbos-conductor-pg \ --db-instance-class db.t4g.micro \ --engine postgres \ --engine-version 16 \ --master-username postgres \ --master-user-password "$POSTGRES_PASSWORD" \ --allocated-storage 20 \ --db-subnet-group-name dbos-conductor-db \ --vpc-security-group-ids $RDS_SG \ --no-publicly-accessible \ --region $AWS_REGION ``` Wait for the instance to become available (this takes a few minutes): ```bash aws rds wait db-instance-available \ --db-instance-identifier dbos-conductor-pg \ --region $AWS_REGION ``` Get the RDS endpoint: ```bash RDS_ENDPOINT=$(aws rds describe-db-instances \ --db-instance-identifier dbos-conductor-pg \ --query "DBInstances[0].Endpoint.Address" \ --output text --region $AWS_REGION) echo "RDS endpoint: $RDS_ENDPOINT" ```
Create the databases and roles from a pod inside the cluster (since the RDS instance is not publicly accessible):
Create databases and roles ```bash kubectl run pg-setup --restart=Never \ --namespace dbos \ --image=postgres:16 \ --env="PGPASSWORD=$POSTGRES_PASSWORD" \ --command -- bash -c " psql -h $RDS_ENDPOINT -U postgres -c 'CREATE DATABASE dbos_conductor;' psql -h $RDS_ENDPOINT -U postgres -c \"CREATE ROLE dbos_conductor_role WITH LOGIN PASSWORD '$CONDUCTOR_ROLE_PASSWORD';\" psql -h $RDS_ENDPOINT -U postgres -c 'GRANT ALL PRIVILEGES ON DATABASE dbos_conductor TO dbos_conductor_role;' psql -h $RDS_ENDPOINT -U postgres -d dbos_conductor -c 'GRANT ALL ON SCHEMA public TO dbos_conductor_role;' " # Wait for the pod to finish, then clean up sleep 15 && kubectl logs pg-setup -n dbos && kubectl delete pod pg-setup -n dbos ``` This creates: - `dbos_conductor` — Conductor's internal database (application registry, metadata) - `dbos_conductor_role` — a dedicated role for Conductor's database access
**Install Cluster Add-ons** We install two Helm charts that the later sections depend on.
Helm installs (Nginx Ingress, Sealed Secrets) **Nginx Ingress Controller** — reverse proxy and TLS termination: ```bash helm repo add ingress-nginx https://kubernetes.github.io/ingress-nginx helm repo update helm install ingress-nginx ingress-nginx/ingress-nginx \ --namespace ingress-nginx --create-namespace \ --set controller.service.type=LoadBalancer ``` **Sealed Secrets** — encrypt secrets for safe Git storage: ```bash helm repo add sealed-secrets https://bitnami-labs.github.io/sealed-secrets helm install sealed-secrets sealed-secrets/sealed-secrets \ --namespace kube-system ``` Verify all add-ons are running: ```bash # Ingress controller kubectl get pods -n ingress-nginx # Sealed Secrets controller kubectl get pods -n kube-system -l app.kubernetes.io/name=sealed-secrets ```
#### Secrets Several components need sensitive credentials. We use [Bitnami Sealed Secrets](https://github.com/bitnami-labs/sealed-secrets): create a regular Secret, encrypt it with `kubeseal`, and apply the encrypted `SealedSecret` to the cluster. The controller decrypts it in-cluster into a standard Kubernetes Secret that pods can reference. The encrypted form is safe to commit to Git. **Secrets Inventory** | Secret | Keys | Used by | |--------|------|---------| | `conductor-db` | `database-url` | Conductor — connection to `dbos_conductor` database | | `conductor-license` | `license-key` | Conductor — production license | **Create and Seal Secrets**
kubeseal commands Create each secret, pipe it through `kubeseal`, and save the encrypted form: ```bash # 1. Conductor database credentials (dedicated role) kubectl create secret generic conductor-db \ --namespace dbos \ --from-literal=database-url="postgresql://dbos_conductor_role:${CONDUCTOR_ROLE_PASSWORD}@${RDS_ENDPOINT}:5432/dbos_conductor?sslmode=require" \ --dry-run=client -o yaml | \ kubeseal --controller-name=sealed-secrets --controller-namespace=kube-system --format yaml \ > sealed-conductor-db.yaml # 2. Conductor license key kubectl create secret generic conductor-license \ --namespace dbos \ --from-literal=license-key="$CONDUCTOR_LICENSE_KEY" \ --dry-run=client -o yaml | \ kubeseal --controller-name=sealed-secrets --controller-namespace=kube-system --format yaml \ > sealed-conductor-license.yaml ```
**Apply and Verify** ```bash kubectl apply -f sealed-conductor-db.yaml kubectl apply -f sealed-conductor-license.yaml ``` Verify the controller has decrypted them into regular Kubernetes Secrets: ```bash kubectl get secrets -n dbos ``` ``` NAME TYPE DATA AGE conductor-db Opaque 1 10s conductor-license Opaque 1 10s ``` #### Ingress With the Nginx Ingress Controller installed, you have a load balancer in front of the cluster. This section creates a TLS certificate and an Ingress resource so that all services are reachable over HTTPS. This walkthrough uses a self-signed certificate on the load balancer's hostname. For production, use [cert-manager](https://cert-manager.io/) with a real domain. Get the load balancer hostname: ```bash ELB_HOSTNAME=$(kubectl get svc -n ingress-nginx ingress-nginx-controller \ -o jsonpath='{.status.loadBalancer.ingress[0].hostname}') echo $ELB_HOSTNAME ``` Save this value — you'll need it throughout the rest of the guide. It looks like `xxxxxxxx.us-west-2.elb.amazonaws.com`.
Create a self-signed TLS certificate ```bash openssl req -x509 -nodes -days 365 -newkey rsa:2048 \ -keyout tls.key -out tls.crt \ -subj "/CN=dbos-conductor" \ -addext "subjectAltName=DNS:${ELB_HOSTNAME}" kubectl create secret tls dbos-tls \ --cert=tls.crt --key=tls.key \ --namespace dbos ``` :::note The CN is kept short because OpenSSL's CN field has a 64-character limit — the actual hostname is covered by the SAN extension. Your browser will show a certificate warning for the self-signed cert — accept it to proceed. ::: :::warning Applications must trust this certificate Applications connect to Conductor over `wss://`, so they verify the Ingress certificate and will fail the TLS handshake against one they do not trust. **For production**, issue a certificate for a domain you control, for example with [cert-manager](https://cert-manager.io/). Note that no public CA will issue a certificate for an `*.elb.amazonaws.com` hostname, so this requires your own domain pointed at the load balancer. **To evaluate with the self-signed certificate**, your applications must be configured to trust it. Distribute it using a ConfigMap and configure `SSL_CERT_FILE` (Go, Python) / `NODE_EXTRA_CA_CERTS` (TypeScript) or the JDK `javax.net.ssl.trustStore` (Java) accordingly. :::
ingress.yaml The Ingress routes `/conductor-api/...` to the Conductor service and everything else to the Console. A regex rewrite strips the `/conductor-api` prefix so Conductor sees requests at `/`. Replace `` with the `$ELB_HOSTNAME` value you retrieved above. ```yaml apiVersion: networking.k8s.io/v1 kind: Ingress metadata: name: dbos-ingress namespace: dbos annotations: nginx.ingress.kubernetes.io/use-regex: "true" nginx.ingress.kubernetes.io/rewrite-target: /$2 nginx.ingress.kubernetes.io/proxy-read-timeout: "3600" nginx.ingress.kubernetes.io/proxy-send-timeout: "3600" spec: ingressClassName: nginx tls: - hosts: - secretName: dbos-tls rules: - host: http: paths: # Both paths are regexes, so ordering matters: ingress-nginx sorts # locations longest-path-first, which puts /conductor-api ahead of # the Console catch-all. - path: /conductor-api(/|$)(.*) pathType: ImplementationSpecific backend: service: name: conductor port: number: 8090 - path: /()(.*) pathType: ImplementationSpecific backend: service: name: console port: number: 80 ``` The `host` in both `tls` and `rules` must match — without it, Nginx serves its default fake certificate instead of `dbos-tls`. | Request path | Backend | |---|---| | `/conductor-api/websocket//` | conductor:8090 → `/websocket//` | | `/conductor-api/healthz` | conductor:8090 → `/healthz` | | `/conductor-api/v1/metrics` | conductor:8090 → `/v1/metrics` | | `/` | console:80 | | `/conductor/applications` | console:80 (UI page) | - **`rewrite-target: /$2`** — strips the `/conductor-api` prefix using the second capture group. The Console catch-all uses `/()(.*)` so `$2` passes the full path through unchanged. - **`proxy-read-timeout` / `proxy-send-timeout`** — set to 3600s to keep Conductor's long-lived WebSocket connections alive.
**Apply the Ingress** ```bash kubectl apply -f ingress.yaml ``` **WebSocket Configuration** The application connects to Conductor via a long-lived WebSocket. Three layers must be configured to prevent idle connections from being dropped: | Layer | Setting | Default | Suggested | Why | |-------|---------|---------|----------|-----| | **Nginx Ingress** | `proxy-read-timeout` | 60s | 3600s | Prevents Nginx from closing an idle WebSocket | | **Nginx Ingress** | `proxy-send-timeout` | 60s | 3600s | Same, for the send direction | | **AWS ELB** | idle timeout | 60s | 3600s | Prevents the load balancer from closing an idle TCP connection | The Nginx timeouts are already set via the Ingress annotations. Nginx handles the `Connection: Upgrade` and `Upgrade: websocket` headers automatically — no additional annotation is needed for the protocol upgrade itself. The AWS load balancer idle timeout is configured separately on the `ingress-nginx-controller` Service: ```bash kubectl patch svc ingress-nginx-controller -n ingress-nginx -p \ '{"metadata":{"annotations":{"service.beta.kubernetes.io/aws-load-balancer-connection-idle-timeout":"3600"}}}' ``` :::note The DBOS SDK sends periodic ping frames that keep the connection active under normal conditions. Albeit the SDK will reconnect automatically, increasing the ELB idle timeout will prevent network hiccups from dropping the connection. ::: #### Deployments Conductor is the core service that manages workflow recovery and the application registry. It connects to the `dbos_conductor` database using the `dbos_conductor_role` credentials.
conductor.yaml ```yaml apiVersion: apps/v1 kind: Deployment metadata: name: conductor namespace: dbos spec: replicas: 1 selector: matchLabels: app: conductor template: metadata: labels: app: conductor spec: containers: - name: conductor # Untagged resolves to :latest. For production, pin an explicit # version so rollouts are reproducible: dbosdev/conductor: image: dbosdev/conductor env: - name: DBOS__CONDUCTOR_DB_URL valueFrom: secretKeyRef: name: conductor-db key: database-url - name: DBOS_CONDUCTOR_LICENSE_KEY valueFrom: secretKeyRef: name: conductor-license key: license-key # Peers forward tasks to each other at this address. It defaults to # 127.0.0.1, which only works for a single replica — set it to the # pod IP before scaling up. - name: DBOS__ADVERTISE_ADDRESS valueFrom: fieldRef: fieldPath: status.podIP ports: - containerPort: 8090 readinessProbe: httpGet: path: /healthz port: 8090 initialDelaySeconds: 5 periodSeconds: 10 livenessProbe: httpGet: path: /healthz port: 8090 initialDelaySeconds: 15 periodSeconds: 30 resources: requests: cpu: 250m memory: 256Mi limits: cpu: "1" memory: 512Mi --- apiVersion: v1 kind: Service metadata: name: conductor namespace: dbos spec: selector: app: conductor ports: - port: 8090 targetPort: 8090 ``` Both sensitive values (`DBOS__CONDUCTOR_DB_URL` and `DBOS_CONDUCTOR_LICENSE_KEY`) are pulled from the Sealed Secrets created in the [Secrets](#secrets) section.
The Console is the web UI for managing applications, monitoring workflows, and generating API keys. In this example, it connects to Conductor via internal cluster DNS.
console.yaml ```yaml apiVersion: apps/v1 kind: Deployment metadata: name: console namespace: dbos spec: replicas: 1 selector: matchLabels: app: console template: metadata: labels: app: console spec: containers: - name: console # As with Conductor, pin an explicit version in production and keep # the two in step: dbosdev/console: image: dbosdev/console env: - name: DBOS_CONDUCTOR_URL value: "conductor.dbos.svc.cluster.local:8090" ports: - containerPort: 8080 readinessProbe: httpGet: path: /health port: 8080 initialDelaySeconds: 5 periodSeconds: 10 livenessProbe: httpGet: path: /health port: 8080 initialDelaySeconds: 10 periodSeconds: 30 resources: requests: cpu: 100m memory: 128Mi limits: cpu: 500m memory: 256Mi --- apiVersion: v1 kind: Service metadata: name: console namespace: dbos spec: selector: app: console ports: - port: 80 targetPort: 8080 ```
Deploy both with: ```bash kubectl apply -f conductor.yaml kubectl apply -f console.yaml ``` Verify both pods are running: ```bash kubectl get pods -n dbos ``` ``` NAME READY STATUS RESTARTS AGE conductor-xxxxxxxxx-xxxxx 1/1 Running 0 2m console-xxxxxxxxx-xxxxx 1/1 Running 0 30s ``` **Access the Console and Generate an API Key** At this point, your self-hosted Conductor deployment is fully operational! Open `https:///` in your browser (accept the self-signed cert warning), then follow the [Conductor setup instructions](../overview.md#connecting-to-conductor) to: 1. Register your application 2. Generate an API key #### Cleanup To tear down all AWS resources when done, delete them in this order. ```bash # 1. Delete the RDS instance and wait for it to be gone aws rds delete-db-instance --db-instance-identifier dbos-conductor-pg \ --skip-final-snapshot --region $AWS_REGION aws rds wait db-instance-deleted \ --db-instance-identifier dbos-conductor-pg \ --region $AWS_REGION # 2. Delete the DB subnet group (must be empty) aws rds delete-db-subnet-group --db-subnet-group-name dbos-conductor-db --region $AWS_REGION # 3. Delete the RDS security group (no longer attached to any instance) RDS_SG=$(aws ec2 describe-security-groups \ --filters "Name=group-name,Values=dbos-conductor-rds" \ --query "SecurityGroups[0].GroupId" --output text --region $AWS_REGION) aws ec2 delete-security-group --group-id $RDS_SG --region $AWS_REGION # 4. Delete the EKS cluster (includes VPC, security groups, and node group) eksctl delete cluster --name dbos-conductor --region $AWS_REGION ``` --- ## Self-Hosting Guide :::info Self-hosted Conductor is released under a [proprietary license](https://www.dbos.dev/conductor-license) and requires a [license key](#licensing). ::: You can self-host Conductor and the DBOS Console on any infrastructure that runs containers. This guide covers what a self-hosted deployment consists of and what it needs in production, independent of where you run it. See our [Kubernetes guide](./hosting-conductor-with-kubernetes.md) for a specific walkthrough. ### Components A self-hosted deployment has three parts: | Component | Image | Port | Role | |---|---|---|---| | **Conductor** | [`dbosdev/conductor`](https://hub.docker.com/r/dbosdev/conductor) | 8090 | The control plane your applications connect to over WebSocket. | | **DBOS Console** | [`dbosdev/console`](https://hub.docker.com/r/dbosdev/console) | 8080 | Conductor's web UI. | | **Postgres** | Any Postgres | 5432 | Conductor's own database, holding its registry of applications, users, and settings. | Conductor's database is separate from the system databases your DBOS applications use. Conductor never connects to your applications' databases; it exchanges workflow metadata and commands with your applications over their WebSocket connections. In addition to the Console, you can use [Conductor's API](../reference/conductor-api.md) and the [dbosctl CLI](../reference/dbosctl.md) to manage your applications and their workflows. ### Trying It Locally with Docker Compose For development and trial purposes, you can self-host Conductor and the DBOS Console on your development machine using Docker Compose. To do this, you need a development license key, which can be obtained from the DBOS Console [here](https://console.dbos.dev/settings/license-key). See [licensing](#licensing) for more information. You should export this license key as an environment variable: ```shell export DBOS_CONDUCTOR_LICENSE_KEY= ``` You can trial self-hosted Conductor with this `docker-compose.yml`:
docker-compose.yml ```yml title="docker-compose.yml" # Docker Compose configuration for self-hosting DBOS Conductor and the DBOS Console. # This configuration is for development purposes only. # Commercial or production use of DBOS Conductor or the DBOS Console requires a paid license key. services: # ============================================ # Postgres # ============================================ postgres: image: postgres:16 container_name: dbos-postgres environment: POSTGRES_USER: postgres POSTGRES_PASSWORD: ${PGPASSWORD:-dbos} POSTGRES_DB: dbos_conductor volumes: - postgres_data:/var/lib/postgresql/data networks: - dbos-network healthcheck: test: ["CMD-SHELL", "pg_isready -U postgres"] interval: 10s timeout: 5s retries: 5 # ============================================ # Conductor # ============================================ conductor: image: dbosdev/conductor container_name: dbos-conductor environment: DBOS__CONDUCTOR_DB_URL: postgresql://postgres:${PGPASSWORD:-dbos}@postgres:5432/dbos_conductor?sslmode=disable # License Key (required) DBOS_CONDUCTOR_LICENSE_KEY: ${DBOS_CONDUCTOR_LICENSE_KEY} # OAuth configuration # DBOS_OAUTH_ENABLED: "true" # DBOS_OAUTH_ISSUER: "https://your-oauth-provider.com/" # DBOS_OAUTH_AUDIENCE: "your-api-audience" ports: - "8090:8090" depends_on: postgres: condition: service_healthy networks: - dbos-network healthcheck: test: ['CMD', 'curl', '-f', 'http://localhost:8090/healthz'] interval: 30s timeout: 3s retries: 3 start_period: 5s # ============================================ # DBOS Console # ============================================ console: image: dbosdev/console container_name: dbos-console environment: # Conductor URL (defaults to conductor:8090 for same Docker network) # Override to connect to remote Conductor # DBOS_CONDUCTOR_URL=conductor.example.com:8090 (remote) DBOS_CONDUCTOR_URL: '${DBOS_CONDUCTOR_URL:-conductor:8090}' # OAuth configuration (uncomment and configure to enable authentication) # DBOS_OAUTH_ENABLED: 'true' # DBOS_OAUTH_AUTHORIZATION_URL: 'https://your-oauth-provider.com/[...]/authorize' # DBOS_OAUTH_TOKEN_URL: 'https://your-oauth-provider.com/[...]/token' # DBOS_OAUTH_CLIENT_ID: 'your-client-id' # DBOS_OAUTH_SCOPE: 'openid profile email' # DBOS_OAUTH_USERINFO_URL: 'https://your-oauth-provider.com/[...]/userinfo' # DBOS_OAUTH_LOGOUT_URL: 'https://your-oauth-provider.com/[...]/logout' # DBOS_OAUTH_AUDIENCE: 'your-api-identifier' ports: # Expose console on port 80 (or override with DBOS_CONSOLE_PORT env var) - '${DBOS_CONSOLE_PORT:-80}:8080' depends_on: conductor: condition: service_healthy networks: - dbos-network healthcheck: test: ['CMD', 'curl', '-f', 'http://localhost:8080/health'] interval: 30s timeout: 3s retries: 3 start_period: 5s # ============================================ # Networks # ============================================ networks: dbos-network: driver: bridge name: dbos-network # ============================================ # Volumes # ============================================ volumes: postgres_data: ```
Start Conductor and the DBOS Console with `docker compose up`. After all containers have launched, navigate to http://localhost to view the self-hosted console. ### Connecting Applications To connect your application to self-hosted Conductor, first [follow these steps](../overview.md#connecting-to-conductor) in your self-hosted DBOS Console to register an application, generate an API key, and set it in your application. :::tip When self-hosting Conductor, make sure you register your application and generate your key in your self-hosted console, not at https://console.dbos.dev. ::: Then, provide your application with a websockets URL to your self-hosted Conductor server. For example, for the Docker Compose setup above, this URL is `ws://localhost:8090/`. In production, use a `wss://` URL that goes through your [reverse proxy](#reverse-proxy-and-tls). **Python** ```python config: DBOSConfig = { "name": "my-app-name", "application_version": "0.1.0", "system_database_url": os.environ.get("DBOS_SYSTEM_DATABASE_URL"), "conductor_key": os.environ.get("DBOS_CONDUCTOR_KEY"), "conductor_url": os.environ.get("DBOS_CONDUCTOR_URL"), } DBOS(config=config) ``` **TypeScript** ```typescript DBOS.setConfig({ "name": "my-app-name", "applicationVersion": "0.1.0", "systemDatabaseUrl": process.env.DBOS_SYSTEM_DATABASE_URL, }); const conductorKey = process.env.DBOS_CONDUCTOR_KEY; const conductorURL = process.env.DBOS_CONDUCTOR_URL; await DBOS.launch({conductorKey, conductorURL}); ``` **Go** ```go conductorKey := os.Getenv("DBOS_CONDUCTOR_KEY") conductorURL := os.Getenv("DBOS_CONDUCTOR_URL") dbosContext, err := dbos.NewDBOSContext(context.Background(), dbos.Config{ AppName: "dbos-starter", ApplicationVersion: "0.1.0", DatabaseURL: os.Getenv("DBOS_SYSTEM_DATABASE_URL"), ConductorURL: conductorURL, ConductorAPIKey: conductorKey, }) ``` **Java** ```java String conductorKey = System.getenv("DBOS_CONDUCTOR_KEY"); String conductorDomain = System.getenv("DBOS_CONDUCTOR_URL"); DBOSConfig config = DBOSConfig.defaults("dbos-java-starter") .withAppVersion("0.1.0") .withDatabaseUrl(System.getenv("DBOS_SYSTEM_JDBC_URL")) .withConductorKey(conductorKey) .withConductorDomain(conductorDomain); ``` ### Licensing For development, testing, or evaluation purposes, you can obtain a trial Conductor key from the [DBOS Console](https://console.dbos.dev/settings/license-key). A license agreement is required for production use. To obtain a production license, please [contact sales](https://www.dbos.dev/contact). You can provide your key to Conductor using the `DBOS_CONDUCTOR_LICENSE_KEY` environment variable. ### Deploying to Production The Docker Compose setup above is for development only. A production deployment runs the same containers with a managed Postgres database, a reverse proxy, secret storage, and [authentication](#security). :::tip For a complete production deployment, including infrastructure, secrets, ingress, and authentication, follow the [Kubernetes guide](./hosting-conductor-with-kubernetes.md). The requirements below apply to any platform. ::: #### Conductor Run Conductor as a stateless container service with an orchestrator like Kubernetes, ECS, Cloud Run, Nomad, or plain VMs. Because all state lives in Postgres, instances are interchangeable, and you can run several for [high availability](#high-availability). Conductor requires these environment variables: | Environment variable | Description | |---|---| | `DBOS__CONDUCTOR_DB_URL` | Connection string for Conductor's Postgres database. We recommend a dedicated database role. | | `DBOS_CONDUCTOR_LICENSE_KEY` | Your [license key](#licensing). | #### DBOS Console Run the Console as a stateless container service listening on port 8080. Set `DBOS_CONDUCTOR_URL` in the Console container to the bare `host:port` of your Conductor service (for example, `conductor.internal:8090`). This differs from the `DBOS_CONDUCTOR_URL` your applications use, which is a full WebSocket URL. Without [OAuth authentication](#security), the Console has no user or organization management. #### Reverse Proxy and TLS Place Conductor and the Console behind a reverse proxy or load balancer (such as Nginx, an ingress controller, or a cloud load balancer) that **supports WebSockets** and does **TLS termination** (Conductor and the Console serve plain HTTP). Route traffic to Conductor on port 8090 and to the Console on port 8080. We recommend setting long idle timeouts on the proxy and on any load balancer in front of it to handle network hiccups (for example, 3600 seconds). The DBOS SDK sends periodic pings and reconnects automatically after a disconnect. #### Network Access - **Outbound HTTPS from Conductor.** Conductor validates its license key against `https://cloud.dbos.dev` at startup and exits if it cannot reach it. Hosts in private networks need a route to the internet, such as a NAT gateway. For air-gapped deployments, [contact sales](https://www.dbos.dev/contact). - **Console to Conductor.** The Console must reach Conductor on port 8090. - **Conductor to Conductor.** In a [highly available](#high-availability) deployment, Conductor instances must reach each other directly. Conductor never needs access to your applications' databases, and your applications need only outbound access to the reverse proxy. #### Secrets Conductor's database URL and license key are secrets. Store them in your platform's secret store (such as Kubernetes Secrets, AWS Secrets Manager, or Vault) and inject them as environment variables. The Conductor API keys your applications use to connect are also secrets, and belong in each application's secret store. The OAuth settings below are not secrets and can be set directly in your deployment configuration. ### High Availability For production deployments that require fault tolerance, you can run multiple Conductor instances in a highly available configuration. All Conductor instances connect to the same Postgres database, which holds all Conductor state, so you can run multiple Conductor instances in multiple availability zones (or other failure domains) behind a load balancer. In a highly available configuration, you should additionally use a highly available Postgres database, such as AWS RDS or Aurora in a multi-AZ replicated configuration, or equivalent offerings from other Postgres providers. #### How It Works Each of your DBOS application's executors maintains a long-lived WebSocket connection to Conductor. When you run multiple Conductor instances behind a load balancer, the load balancer distributes these connections across instances, so each executor connects to (and is owned by) exactly one Conductor instance at a time. When a request (for example, from the DBOS Console) needs to reach a particular executor, it may land on any Conductor instance. If that instance does not own the target executor's connection, it looks up the owning instance in Postgres and forwards the request to it directly. The owning instance then relays the request to the executor over its WebSocket. This means **Conductor instances must be able to reach each other directly over the network**, in addition to being reachable through the load balancer. This peer-to-peer traffic flows directly between instances. To enable peer forwarding, each Conductor instance must advertise an address that its peers can use to reach it. Set the `DBOS__ADVERTISE_ADDRESS` environment variable to a routable address (a hostname or IP, without a port); peers connect to this address on the Conductor port (`8090` by default, configurable with `DBOS__CONDUCTOR_PORT`). #### Configuration summary | Environment variable | Default | Description | |---|---|---| | `DBOS__ADVERTISE_ADDRESS` | `127.0.0.1` | Routable address or URL peers use to forward requests to this instance. **Must be set** for multi-instance deployments. | | `DBOS__CONDUCTOR_PORT` | `8090` | Port Conductor listens on and advertises to peers. | ### Security To securely self-host Conductor in production, you should set up authentication and authorization for all API calls made to it. :::warning Conductor performs **no authentication** unless OAuth is enabled. Without it, all API requests run as a built-in `local` organization admin, and Conductor does not verify API keys on incoming WebSocket connections. Anyone who can reach Conductor can register applications, cancel, resume, fork, or delete workflows, and create API keys. Configure OAuth before exposing Conductor to any untrusted network. ::: You can integrate Conductor with any OAuth-compatible single-sign on (SSO) experience. To do this, first register the DBOS Console as an application and Conductor as an API (audience) with your OAuth provider. Configure the following with your provider: - `https://your-domain/oauth/callback` as a callback URL - `https://your-domain` as an allowed web origin - Authorization code with PKCE as an allowed grant type - `openid profile email` as valid scopes. Then, set these environment variables in your Conductor container: ```yml DBOS_OAUTH_ENABLED: "true" DBOS_OAUTH_ISSUER: "https://your-oauth-provider.com/" DBOS_OAUTH_AUDIENCE: "your-api-audience" ``` And set these environment variables in your DBOS Console container: ```yml DBOS_OAUTH_ENABLED: 'true' DBOS_OAUTH_AUTHORIZATION_URL: 'https://your-oauth-provider.com/[...]/authorize' DBOS_OAUTH_TOKEN_URL: 'https://your-oauth-provider.com/[...]/token' DBOS_OAUTH_CLIENT_ID: 'your-client-id' DBOS_OAUTH_SCOPE: 'openid profile email' DBOS_OAUTH_USERINFO_URL: 'https://your-oauth-provider.com/[...]/userinfo' DBOS_OAUTH_LOGOUT_URL: 'https://your-oauth-provider.com/[...]/logout' DBOS_OAUTH_AUDIENCE: 'your-api-audience' ``` These values correspond to the client credentials and endpoints provided by your OAuth identity provider (such as Google, Auth0, or Okta). None of these values are secrets. When properly configured, the DBOS Console will redirect users to your SSO login page and enforce authentication on access. This will also enable user and organization management features. ### Upgrading You can upgrade Conductor and the DBOS Console by simply upgrading the container versions and restarting the service. Because Conductor is entirely out-of-band, this will have no impact on your DBOS applications' availability; your apps will seamlessly reconnect to your new Conductor version. We recommend regularly upgrading Conductor and the DBOS Console to the latest versions to take advantage of new features. We always guarantee it is safe to upgrade directly from any past version to any future version. For the best experience, we recommend upgrading Conductor and the DBOS Console together and not using a version of the DBOS Console more recent than your version of Conductor. ### Scaling Architecturally, Conductor is entirely off your workflows orchestration path. As such, it requires minimal resources to serve large application deployments. A single server hosting the Conductor service can serve tens of thousands of application servers processing millions of workflows per second. --- ## Workflow Management(Conductor) :::info Workflow observability and management features are only available for applications connected to [Conductor](./overview.md). ::: ### Viewing Workflows Navigate to the workflows tab of your application's page on the DBOS Console to see a list of its workflows: This includes **all** your application's workflows: those currently executing, those enqueued for execution, those that have completed successfully, and those that have failed. You can filter by time, workflow ID, workflow name, and workflow status (for example, you can search for all failed workflow executions in the past day). Click on a workflow to see details, including its input and output: Click "Show Workflow Steps" to view the workflow's execution as a trace timeline (showing the workflow, its steps, and its child workflows and their steps). For example, here is the trace of a workflow that processes multiple tasks concurrently by enqueueing child workflows: You can manage individual workflows directly from the DBOS Console. ##### Cancelling Workflows You can cancel any workflow that has not completed: `PENDING`, `ENQUEUED`, or `DELAYED`. Cancelling a workflow sets its status to `CANCELLED`. If the workflow is currently executing, cancelling it preempts its execution (interrupting it at the beginning of its next step). If the workflow is enqueued or delayed, cancelling removes it from the queue. ##### Resuming Workflows You can resume any `ENQUEUED`, `DELAYED`, `CANCELLED` or `MAX_RECOVERY_ATTEMPTS_EXCEEDED` workflow. Resuming a workflow resumes its execution from its last completed step. If the workflow is enqueued, this bypasses the queue to start it immediately. ##### Forking Workflows You can start a new execution of a workflow by **forking** it from a specific step. To do this, open the workflow steps view, select a particular step, and click "Fork". When you fork a workflow, DBOS generates a new workflow with a new workflow ID, copies to that workflow the original workflow's inputs and all its steps up to the selected step, then begins executing the new workflow from the selected step. Forking a workflow is useful for recovering from outages in downstream services (by forking from the step that failed after the outage is resolved) or for "patching" workflows that failed due to a bug in a previous application version (by forking from the bugged step to an application version on which the bug is fixed). ### Exporting Workflows You can export a workflow from one application to another by clicking the "Export" button in the workflow details panel. This copies all information on that workflow (and optionally its children) to the other application's system database. This is most useful for copying workflows from a production to development environment, for example to examine and (using fork) reproduce a bug that originally occurred in production. --- ## DBOS Examples ## Featured Examples import { FaHackerNews, FaSlack, FaForwardFast, FaPerson } from "react-icons/fa6"; import { HiMiniQueueList } from "react-icons/hi2"; import { BiAddToQueue } from "react-icons/bi"; import { MdOutlineShoppingCart } from "react-icons/md"; import { SiApachekafka } from "react-icons/si"; import { IoEarth } from "react-icons/io5"; import { RiCalendarScheduleLine } from "react-icons/ri"; import { IoIosChatboxes } from "react-icons/io"; import { PiFileMagnifyingGlassBold } from "react-icons/pi"; import { RiCustomerService2Line } from "react-icons/ri"; import { TbClock2 } from "react-icons/tb"; import { VscGraphLine } from "react-icons/vsc"; import { FiInbox } from "react-icons/fi";
--- ## Comparing DBOS and Temporal DBOS and Temporal both provide durable workflows. The main difference is that Temporal implements durable workflows in a heavyweight orchestration service, whereas DBOS implements them in a Postgres-backed library. In our opinion, the DBOS architecture is simpler to adopt and operate. :::info To learn how to migrate an application from Temporal to DBOS, see the [migration guide](./migrating-from-temporal.md). ::: ### Simpler Architecture Temporal is designed around a central workflow server that orchestrates workflow execution on a cluster of workers. The central server runs workflow code, dispatching steps to workers. Workers execute steps, then return their output to the orchestrator, which durably checkpoints it then dispatches the next step. Because of this design, adding Temporal to an application requires rearchitecting it. First, you must move all your workflow and activity (step) code to run on a cluster of Temporal workers. You must also rewrite all interactions between your application and its workflows to go through the orchestration server and its client APIs. Then, if you self-host Temporal, you must also operate a highly available Temporal cluster and its supporting datastores (typically a durable database such as Cassandra and a visibility store such as Elasticsearch), effectively adding another distributed system alongside your application. Alternatively, you can use their managed cloud service, but that places critical workflow state in a third-party service. In either case, the Temporal server and its data stores are on the critical path for workflow execution and are single points of failure for your system; if they have downtime your application becomes unavailable. By contrast, DBOS is an open-source Postgres-backed library. To add DBOS to an application, you install the open-source library and annotate workflows and steps. The library uses Postgres to checkpoint workflow progress and recover workflows from failure. Because DBOS uses Postgres for orchestration, you don't need to change how your application is architected or deployed—it can run on any infrastructure connected to any Postgres-compatible database. You scale DBOS by scaling Postgres, and Postgres [scales well](https://www.dbos.dev/blog/benchmarking-workflow-execution-scalability-on-postgres). ### Advantages of DBOS #### Improved Operational Reliability The only point of failure in DBOS is Postgres. If your organization already uses Postgres, DBOS does not add any new infrastructural dependencies or points of failure to your application's architecture. By contrast, the Temporal architecture adds multiple new points of failure: the Temporal orchestration server and the datastores it relies on (most commonly Cassandra for durability and Elasticsearch for observability). Your team is responsible for operating them, and if they have downtime, your application becomes unavailable. #### >10x Better Latency In DBOS, the only overhead required to call a step is checkpointing its output. This requires a single Postgres write, which typically takes 1-2ms. In Temporal, a step requires an async dispatch from the central server, which takes [tens to hundreds of ms](https://temporal.io/blog/reduce-latency-and-speed-up-your-temporal-workflows). Thus, DBOS is preferred for interactive or otherwise latency-sensitive workflows. #### Privacy-Preserving Architecture Because DBOS stores workflow data in your Postgres database, it is intrinsically privacy-preserving: you own your data, you store it in your Postgres, and it is never stored or sent anywhere else. By contrast, to use Temporal, you must send potentially sensitive data (including workflow and step checkpoints) to the Temporal server for storage. #### Rich Workflow Introspection and Management Because DBOS is built on Postgres, it provides rich SQL-backed workflow introspection and management. You can search workflows by name, time, queue, version, or custom properties, introspect individual steps, and pause, cancel, or resume workflows. All these capabilities are available both programmatically and through a web UI. One particularly powerful and unique feature is **fork**: you can restart a workflow from a specific step, either programmatically or from the UI. This is useful for recovering from an unexpected failure in a step, such as a failure due to a bug or an outage. For example, if a large number of billing workflows fail overnight due to an outage in a payment API, you can use fork to restart them all from the payment step after the outage is resolved. #### Durable Workflow Queues DBOS provides durable workflow queues with managed flow control. Using queues, you can manage how many workflows can execute concurrently (globally, per-worker, and per-tenant) as well as which workers can execute which workflows. Temporal does not have comparable queueing or flow control abstractions, making it harder to control when and where workflows execute. Learn more about DBOS queues in the queues tutorial ([Python](../python/tutorials/queue-tutorial.md), [TypeScript](../typescript/tutorials/queue-tutorial.md), [Go](../golang/tutorials/queue-tutorial.md), [Java](../java/tutorials/queue-tutorial.md)). --- ## Concurrent Executions DBOS guarantees that every workflow runs to completion: if an executor crashes or becomes unreachable, another executor recovers its `PENDING` workflows and re-executes them from their last completed step. The component responsible for recovery, e.g., [DBOS Conductor](../conductor/overview.md), detects unhealthy executors and triggers recovery of its workflows. Sometimes, for example during the rollout of a new application image, that observation can be wrong, and a "zombie" executor could still be running your workflow. This means the same workflow instance could be running on two executors. (DBOS detects and prevents concurrent executions of the same workflow on the same executor.) DBOS is designed so that step and workflow invariants are preserved during these situations: steps get at-least-once guarantees and workflow outcomes are persisted exactly-once. ### Workflow Ownership DBOS detects concurrent executions by tracking which execution **owns** each workflow. When an execution starts running a workflow (because the workflow was started, dequeued, recovered, or resumed), it generates a unique ownership token, records it in the workflow's row in the [`workflow_status`](./system-tables.md#dbosworkflow_status) table, and keeps it in memory. Control-plane operations, such as cancelling, resuming, rewinding, or recovering a workflow, clear the recorded token. Every time an execution writes a checkpoint for the workflow, it first checks, in the same database transaction, that the recorded token still matches its own. If the token no longer matches, the execution has lost ownership: another execution has taken over the workflow, or it was cancelled. The checkpoint is not written, and the execution stops running the workflow and **parks**, _i.e._, it waits for the workflow's recorded outcome to become visible in the database, then delivers that recorded outcome through its own handle. For example, if a "zombie" executor keeps running a workflow after it has been recovered elsewhere, its next checkpoint fails the ownership check, so it stops, and its handle returns the result recorded by the execution that owns the workflow. Similarly, when you cancel a running workflow, its execution stops at its next checkpoint. When an execution loses ownership, DBOS throws an exception inside the workflow (`DBOSWorkflowConflictIDError` in Python, `DBOSWorkflowConflictError` in TypeScript, `DBOSWorkflowExecutionConflictException` in Java) or returns an error ([`ErrConflictingWorkflowID`](../golang/reference/workflows-steps.md#errors) in Go). Do not catch and ignore that error: no subsequent work in the workflow will be made durable. Separately, if a single execution records a result for a step and then tries to record a different result for the same step, DBOS throws a step nondeterminism error (`DBOSStepNondeterminismError` in Python and TypeScript). This indicates that the workflow is not deterministic (see determinism requirements in [Python](../python/tutorials/workflow-tutorial.md#determinism) and [TypeScript](../typescript/tutorials/workflow-tutorial.md#determinism)). --- ## DBOSify: Drop-in Temporal Replacement You can run your existing Temporal code on DBOS with [**DBOSify**](https://github.com/dbos-inc/dbosify-py). DBOSify is a drop-in replacement for the [Temporal Python SDK](https://github.com/temporalio/sdk-python) that uses Postgres (through [DBOS Transact](https://github.com/dbos-inc/dbos-transact-py)) instead of a Temporal server. It runs your workflows, activities, signals, updates, queries, retries, and recovery with no infrastructure except Postgres. To migrate, you import `dbosify` instead of `temporalio` and connect your clients and workers to a Postgres database instead of a Temporal server. :::info DBOSify only supports Python for now. For architectural details and detailed feature compatibility, see the [DBOSify architecture page](https://github.com/dbos-inc/dbosify-py/blob/main/docs/ARCHITECTURE.md). ::: :::tip To rewrite a Temporal application to use DBOS natively (in Python, TypeScript, Go, or Java), see [Migrating From Temporal](./migrating-from-temporal.md). ::: ### Using DBOSify Install DBOSify: ```shell pip install dbosify ``` Then import `dbosify` instead of `temporalio` and connect to Postgres instead of a Temporal server:
DBOSify Example ```python import asyncio import os from datetime import timedelta from dbosify import activity, workflow from dbosify.client import Client from dbosify.worker import Worker # A connection string to your Postgres database, instead of a Temporal server address DB_URL = os.environ.get("DBOS_SYSTEM_DATABASE_URL") @activity.defn async def compose_greeting(name: str) -> str: return f"Hello, {name}!" @workflow.defn class GreetingWorkflow: @workflow.run async def run(self, name: str) -> str: return await workflow.execute_activity( compose_greeting, name, start_to_close_timeout=timedelta(seconds=10) ) async def main() -> None: worker = Worker( DB_URL, task_queue="greetings", workflows=[GreetingWorkflow], activities=[compose_greeting], ) async with worker: async with await Client.connect(DB_URL) as client: result = await client.execute_workflow( GreetingWorkflow.run, "World", id="greeting-1", task_queue="greetings" ) print(result) # Hello, World! if __name__ == "__main__": asyncio.run(main()) ```
### Connection API Where a Temporal application connects to a Temporal server, a DBOSify application connects to Postgres. Both the client and the worker take a Postgres connection string. For more control, you can also construct a client from a `dbos.DBOSClient` and a worker from a `dbos.DBOSConfig`. **Client:** Connect a client with `Client.connect`: ```python client = await Client.connect( system_database_url, # Postgres connection string namespace="default", # optional; each namespace maps to its own Postgres schema ) ``` For full control, instead build a [`dbos.DBOSClient`](../python/reference/client.md) yourself and pass it to the `Client(...)` constructor. **Worker:** To configure a `Worker`, pass it a Postgres connection string or a [`dbos.DBOSConfig`](../python/reference/configuration.md): ```python worker = Worker( config, # Postgres connection string or a dbos.DBOSConfig task_queue="greetings", # required namespace="default", # optional workflows=[GreetingWorkflow], activities=[compose_greeting], ) await worker.run() # or use `async with worker:` ``` --- ## Migrating From Temporal This guide explains how to migrate a Temporal application to DBOS, with a focus on how each major Temporal feature translates to DBOS. :::info For a high-level comparison of DBOS and Temporal's architectures, see [Comparing DBOS and Temporal](./comparing-temporal.md). ::: :::tip Also check out [DBOSify](./dbosify.md), a drop-in replacement for the Temporal Python SDK backed by Postgres. ::: ### Workflows The core feature of both DBOS and Temporal is durably executed workflows. Both DBOS and Temporal automatically recover workflows from the last completed step (activity) after any failure. Both DBOS and Temporal support extremely long-running workflows, including workflows that run for weeks or months. **Temporal:** ```python @workflow.defn class OrderWorkflow: @workflow.run async def run(self, order: Order) -> str: result = await workflow.execute_activity( validate_order, order, start_to_close_timeout=timedelta(seconds=30), ) confirmation = await workflow.execute_activity( process_payment, result, start_to_close_timeout=timedelta(seconds=60), ) return confirmation ``` **DBOS:** **Python** ```python @DBOS.workflow() def order_workflow(order: Order) -> str: result = validate_order(order) confirmation = process_payment(result) return confirmation ``` Learn more in the [workflows tutorial](../python/tutorials/workflow-tutorial.md). **TypeScript** ```typescript async function orderWorkflow(order: Order): Promise { const result = await validateOrder(order); const confirmation = await processPayment(result); return confirmation; } const orderWorkflowFn = DBOS.registerWorkflow(orderWorkflow); ``` Learn more in the [workflows tutorial](../typescript/tutorials/workflow-tutorial.md). **Go** ```go func OrderWorkflow(ctx dbos.DBOSContext, order Order) (string, error) { result, err := dbos.RunAsStep(ctx, func(stepCtx context.Context) (string, error) { return validateOrder(stepCtx, order) }, dbos.WithStepName("validateOrder")) if err != nil { return "", err } confirmation, err := dbos.RunAsStep(ctx, func(stepCtx context.Context) (string, error) { return processPayment(stepCtx, result) }, dbos.WithStepName("processPayment")) return confirmation, err } ``` Learn more in the [workflows tutorial](../golang/tutorials/workflow-tutorial.md). **Java** ```java @Workflow(name = "orderWorkflow") public String orderWorkflow(Order order) { String result = dbos.runStep(() -> validateOrder(order), "validateOrder"); String confirmation = dbos.runStep(() -> processPayment(result), "processPayment"); return confirmation; } ``` Learn more in the [workflows tutorial](../java/tutorials/workflow-tutorial.md). #### Starting Workflows In Temporal, workflows are started through a client connected to the Temporal server. The workflow task is then picked up by a worker, which executes the workflow logic. In DBOS, workflows can be started directly within your application process. Alternatively, you can enqueue workflows from a separate process using the DBOS Client, which connects directly to the DBOS system database. **Temporal:** ```python client = await Client.connect("localhost:7233") handle = await client.start_workflow( OrderWorkflow.run, order, id="order-123", task_queue="orders", ) result = await handle.result() ``` **DBOS:** **Python** ```python # Starting a workflow from in your application with SetWorkflowID("order-123"): handle = DBOS.start_workflow(order_workflow, order) result = handle.get_result() ``` ```python # Starting a workflow from another application using the DBOS Client client = DBOSClient(system_database_url=os.environ["DBOS_SYSTEM_DATABASE_URL"]) handle = client.enqueue({"workflow_name": "order_workflow", "queue_name": "orders"}, order) result = handle.get_result() ``` Learn more in the [workflows tutorial](../python/tutorials/workflow-tutorial.md). **TypeScript** ```typescript // Starting a workflow from in your application const handle = await DBOS.startWorkflow(orderWorkflowFn, {workflowID: "order-123"})(order); const result = await handle.getResult(); ``` ```typescript // Starting a workflow from another application using the DBOS Client const client = await DBOSClient.create({systemDatabaseUrl: process.env.DBOS_SYSTEM_DATABASE_URL!}); await client.enqueue( { workflowName: "orderWorkflow", queueName: "orders" }, order, ); ``` Learn more in the [workflows tutorial](../typescript/tutorials/workflow-tutorial.md). **Go** ```go // Starting a workflow from in your application handle, err := dbos.RunWorkflow(dbosContext, OrderWorkflow, order, dbos.WithWorkflowID("order-123")) result, err := handle.GetResult() ``` ```go // Starting a workflow from another application using the DBOS Client client, err := dbos.NewClient(context.Background(), dbos.ClientConfig{ DatabaseURL: os.Getenv("DBOS_SYSTEM_DATABASE_URL"), }) handle, err := dbos.Enqueue[Order, string](client, "orders", "OrderWorkflow", order) result, err := handle.GetResult() ``` Learn more in the [workflows tutorial](../golang/tutorials/workflow-tutorial.md). **Java** ```java // Starting a workflow from in your application WorkflowHandle handle = dbos.startWorkflow( () -> proxy.orderWorkflow(order), new StartWorkflowOptions().withWorkflowId("order-123") ); String result = handle.getResult(); ``` ```java // Starting a workflow from another application using the DBOS Client var client = new DBOSClient(dbUrl, dbUser, dbPassword); var options = new EnqueueOptions("orderWorkflow", "com.example.OrderImpl", QueueName.of("orders")); var handle = client.enqueueWorkflow(options, new Object[]{order}); Object result = handle.getResult(); ``` Learn more in the [workflows tutorial](../java/tutorials/workflow-tutorial.md). #### Workflow IDs and Idempotency Both systems support workflow IDs to provide idempotent execution. In Temporal, the workflow ID is passed when starting a workflow. In DBOS, you set the workflow ID before invoking the workflow. If a workflow with the same ID has already executed, DBOS returns the previously recorded result instead of running the workflow again. One important difference is how workflow executions are identified. Temporal uniquely identifies an execution using a combination of workflow ID and run ID, so a workflow may have multiple run instances over time. DBOS, by contrast, treats each execution as uniquely identified by its workflow ID, so a workflow ID corresponds to exactly one execution. **Python** ```python with SetWorkflowID("payment-idempotency-key"): order_workflow(order) ``` Learn more in the [workflows tutorial](../python/tutorials/workflow-tutorial.md#workflow-ids-and-idempotency). **TypeScript** ```typescript const handle = await DBOS.startWorkflow(orderWorkflowFn, {workflowID: "payment-idempotency-key"})(order); ``` Learn more in the [workflows tutorial](../typescript/tutorials/workflow-tutorial.md#workflow-ids-and-idempotency). **Go** ```go handle, err := dbos.RunWorkflow(dbosContext, OrderWorkflow, order, dbos.WithWorkflowID("payment-idempotency-key")) ``` Learn more in the [workflows tutorial](../golang/tutorials/workflow-tutorial.md#workflow-ids-and-idempotency). **Java** ```java dbos.startWorkflow( () -> proxy.orderWorkflow(order), new StartWorkflowOptions().withWorkflowId("payment-idempotency-key") ); ``` Learn more in the [workflows tutorial](../java/tutorials/workflow-tutorial.md#workflow-ids-and-idempotency). #### Determinism Both DBOS and Temporal require workflows to be deterministic. Non-deterministic operations (API calls, random numbers, current time) must happen inside activities/steps, not directly in the workflow function. #### Durable Timers Temporal's `workflow.sleep()` maps directly to `DBOS.sleep()`. Both are durable and persist across restarts. **Temporal:** ```python await workflow.sleep(timedelta(hours=24)) ``` **DBOS:** **Python** ```python DBOS.sleep(86400) # seconds ``` Learn more in the [workflows tutorial](../python/tutorials/workflow-tutorial.md#durable-sleep). **TypeScript** ```typescript await DBOS.sleep(86400000); // milliseconds ``` Learn more in the [workflows tutorial](../typescript/tutorials/workflow-tutorial.md#durable-sleep). **Go** ```go dbos.Sleep(ctx, 24 * time.Hour) ``` Learn more in the [workflows tutorial](../golang/tutorials/workflow-tutorial.md#durable-sleep). **Java** ```java dbos.sleep(Duration.ofHours(24)); ``` Learn more in the [workflows tutorial](../java/tutorials/workflow-tutorial.md#durable-sleep). #### Continue-as-New A common pattern in Temporal is to use an extremely long-running workflow as a durable object. Applications interact with it via signals and queries and periodically refresh its state with `continue_as_new` to avoid Temporal's workflow size limits. In DBOS, there are no workflow history limits beyond the underlying database column and storage limits. However, instead of maintaining extremely long-lived workflows, which can slow down replay during recovery, we generally recommend storing long-lived objects directly in your database and interacting with them through shorter-lived workflows. To coordinate those interactions, you can use DBOS durable queues, especially partitioned queues, to control concurrency and ordering. This approach provides similar guarantees while avoiding the complexity of managing extremely long-running workflows. If you have workflows with many steps, another useful pattern is to use an outer control workflow that orchestrates smaller sub-workflows (child workflows). This improves observability (because you can easily isolate each sub-workflow) and can speed up recovery and replay. ### Activities → Steps Temporal activities map to DBOS steps. Both are where side effects and non-deterministic operations happen. The key architectural difference is how they are executed. In Temporal, activities are dispatched to workers, often running in separate processes, through the Temporal server. This introduces a network round trip between the workflow and the worker executing the activity. Temporal also supports local activities that run in the same process as the workflow, but they come with several limitations. In DBOS, steps run in the same process as the workflow and are invoked like regular function calls. DBOS automatically checkpoints the step's result to your database, guaranteeing durability without requiring a separate worker process. Because execution happens in place, steps typically have lower latency and less overhead compared to remotely dispatched activities. **Temporal:** ```python @activity.defn async def send_email(to: str, body: str) -> bool: response = requests.post(EMAIL_API, json={"to": to, "body": body}) return response.ok ``` **DBOS:** **Python** ```python @DBOS.step() def send_email(to: str, body: str) -> bool: response = requests.post(EMAIL_API, json={"to": to, "body": body}) return response.ok ``` Learn more in the [steps tutorial](../python/tutorials/step-tutorial.md). **TypeScript** ```typescript const sendEmail = DBOS.registerStep(async (to: string, body: string): Promise => { const response = await fetch(EMAIL_API, { method: "POST", body: JSON.stringify({ to, body }), }); return response.ok; }); ``` Learn more in the [steps tutorial](../typescript/tutorials/step-tutorial.md). **Go** ```go // Steps are called inline using RunAsStep result, err := dbos.RunAsStep(ctx, func(stepCtx context.Context) (bool, error) { return sendEmail(stepCtx, to, body) }, dbos.WithStepName("sendEmail")) ``` Learn more in the [steps tutorial](../golang/tutorials/step-tutorial.md). **Java** ```java // Steps are called inline using dbos.runStep boolean result = dbos.runStep(() -> sendEmail(to, body), "sendEmail"); ``` Learn more in the [steps tutorial](../java/tutorials/step-tutorial.md). #### Retries Both systems support configurable retries with exponential backoff. **Temporal**: ```python result = await workflow.execute_activity( send_email, args=[to, body], start_to_close_timeout=timedelta(seconds=30), retry_policy=RetryPolicy( initial_interval=timedelta(seconds=1), backoff_coefficient=2.0, maximum_attempts=5, ), ) ``` **DBOS:** **Python** ```python @DBOS.step(retries_allowed=True, max_attempts=5, interval_seconds=1.0, backoff_rate=2.0) def send_email(to: str, body: str) -> bool: response = requests.post(EMAIL_API, json={"to": to, "body": body}) return response.ok ``` Learn more in the [steps tutorial](../python/tutorials/step-tutorial.md#configurable-retries). **TypeScript** ```typescript const sendEmail = DBOS.registerStep( async (to: string, body: string): Promise => { const response = await fetch(EMAIL_API, { method: "POST", body: JSON.stringify({ to, body }), }); return response.ok; }, { retriesAllowed: true, maxAttempts: 5, intervalSeconds: 1.0, backoffRate: 2.0 } ); ``` Learn more in the [steps tutorial](../typescript/tutorials/step-tutorial.md#configurable-retries). **Go** ```go result, err := dbos.RunAsStep(ctx, func(stepCtx context.Context) (bool, error) { return sendEmail(stepCtx, to, body) }, dbos.WithStepName("sendEmail"), dbos.WithStepMaxRetries(5), dbos.WithBaseInterval(1 * time.Second), dbos.WithBackoffFactor(2.0), ) ``` Learn more in the [steps tutorial](../golang/tutorials/step-tutorial.md#configurable-retries). **Java** ```java boolean result = dbos.runStep( () -> sendEmail(to, body), new StepOptions("sendEmail") .withMaxAttempts(5) .withRetryInterval(Duration.ofSeconds(1)) .withBackoffRate(2.0) ); ``` Learn more in the [steps tutorial](../java/tutorials/step-tutorial.md#configurable-retries). #### Heartbeats Temporal activities support heartbeats for long-running operations so the server knows the activity is still alive. DBOS does not require heartbeats because there is no central orchestrator monitoring activity execution; instead, steps run directly in your application process. #### Database Operations DBOS provides a special type of step called a transaction that executes database operations in a single database transaction, co-committed with the DBOS checkpoint. This provides exactly-once semantics for database writes, which is stronger than the at-least-once semantics offered by Temporal. **Python** Learn more in the [transactions tutorial](../python/tutorials/transaction-tutorial.md). ```python import os from dbos import SQLAlchemyDatasource from sqlalchemy import text ds = SQLAlchemyDatasource.create(os.environ["APP_DATABASE_URL"]) @ds.transaction() def update_order_status(order_id: str, status: str) -> None: ds.sql_session().execute( text("UPDATE orders SET status = :status WHERE id = :id"), {"status": status, "id": order_id} ) ``` **TypeScript** Learn more in the [transactions tutorial](../typescript/tutorials/transaction-tutorial.md). ```typescript const dataSource = new KnexDataSource('app-db', { client: 'pg', connection: process.env.DBOS_DATABASE_URL }); async function updateOrderStatus(orderId: string, status: string) { await dataSource.client('orders') .where({ id: orderId }) .update({ status }); } const updateOrderStatusTx = dataSource.registerTransaction(updateOrderStatus); ``` ### Signals → Messages Temporal signals allow external processes to send data to a running workflow. In DBOS, the equivalent mechanism is **messages (notifications)**, which external processes send using `send()` and workflows read using `recv()`. **Temporal:** ```python # In the workflow @workflow.defn class OrderWorkflow: def __init__(self): self.payment_status = None @workflow.signal async def payment_received(self, status: str): self.payment_status = status @workflow.run async def run(self, order: Order): # ... start order processing ... await workflow.wait_condition(lambda: self.payment_status is not None) if self.payment_status == "paid": # handle success else: # handle failure # Sending the signal handle = client.get_workflow_handle("order-123") await handle.signal(OrderWorkflow.payment_received, "paid") ``` **DBOS:** **Python** ```python # In the workflow @DBOS.workflow() def order_workflow(order: Order): # ... start order processing ... payment_status = DBOS.recv("payment_status", timeout_seconds=3600) if payment_status is not None and payment_status == "paid": # handle success else: # handle failure # Sending the message DBOS.send("order-123", "paid", topic="payment_status") ``` Learn more in the [workflow communication tutorial](../python/tutorials/workflow-communication.md#workflow-messaging-and-notifications). **TypeScript** ```typescript // In the workflow async function orderWorkflow(order: Order) { // ... start order processing ... const paymentStatus = await DBOS.recv("payment_status", 3600); if (paymentStatus !== null && paymentStatus === "paid") { // handle success } else { // handle failure } } const orderWorkflowFn = DBOS.registerWorkflow(orderWorkflow); // Sending the message await DBOS.send("order-123", "paid", "payment_status"); ``` Learn more in the [workflow communication tutorial](../typescript/tutorials/workflow-communication.md#workflow-messaging-and-notifications). **Go** ```go // In the workflow func OrderWorkflow(ctx dbos.DBOSContext, order Order) (string, error) { // ... start order processing ... paymentStatus, err := dbos.Recv[string](ctx, "payment_status", 1*time.Hour) if err != nil { return "", err } if paymentStatus == "paid" { // handle success } else { // handle failure } // ... } // Sending the message err := dbos.Send(dbosContext, "order-123", "paid", "payment_status") ``` Learn more in the [workflow communication tutorial](../golang/tutorials/workflow-communication.md#workflow-messaging-and-notifications). **Java** ```java // In the workflow @Workflow(name = "orderWorkflow") public void orderWorkflow(Order order) { // ... start order processing ... Optional paymentStatus = dbos.recv("payment_status", Duration.ofHours(1)); if (paymentStatus.map("paid"::equals).orElse(false)) { // handle success } else { // handle failure } } // Sending the message dbos.send("order-123", "paid", "payment_status"); ``` Learn more in the [workflow communication tutorial](../java/tutorials/workflow-communication.md#workflow-messaging-and-notifications). Messages are persisted to the database, so they remain available even after the workflow completes. ### Queries → Events Temporal queries allow external code to synchronously read the state of a workflow. In DBOS, the equivalent mechanism is **events**, which workflows publish using `set_event()` and external processes read using `get_event()`. **Temporal:** ```python @workflow.defn class OrderWorkflow: def __init__(self): self.progress = 0 @workflow.query def get_progress(self) -> int: return self.progress @workflow.run async def run(self, order: Order): self.progress = 25 await workflow.execute_activity(validate_order, order, ...) self.progress = 50 # ... # Querying workflow state handle = client.get_workflow_handle("order-123") progress = await handle.query(OrderWorkflow.get_progress) ``` **DBOS:** **Python** ```python @DBOS.workflow() def order_workflow(order: Order): DBOS.set_event("progress", 25) validate_order(order) DBOS.set_event("progress", 50) # ... # Reading workflow state progress = DBOS.get_event("order-123", "progress") ``` Learn more in the [workflow communication tutorial](../python/tutorials/workflow-communication.md#workflow-events). **TypeScript** ```typescript async function orderWorkflow(order: Order) { await DBOS.setEvent("progress", 25); await validateOrder(order); await DBOS.setEvent("progress", 50); // ... } const orderWorkflowFn = DBOS.registerWorkflow(orderWorkflow); // Reading workflow state const progress = await DBOS.getEvent("order-123", "progress"); ``` Learn more in the [workflow communication tutorial](../typescript/tutorials/workflow-communication.md#workflow-events). **Go** ```go func OrderWorkflow(ctx dbos.DBOSContext, order Order) (string, error) { dbos.SetEvent(ctx, "progress", 25) // ... validate order ... dbos.SetEvent(ctx, "progress", 50) // ... } // Reading workflow state progress, err := dbos.GetEvent[int](dbosContext, "order-123", "progress", 30*time.Second) ``` Learn more in the [workflow communication tutorial](../golang/tutorials/workflow-communication.md#workflow-events). **Java** ```java @Workflow(name = "orderWorkflow") public void orderWorkflow(Order order) { dbos.setEvent("progress", 25); // ... validate order ... dbos.setEvent("progress", 50); // ... } // Reading workflow state Optional progress = dbos.getEvent("order-123", "progress", Duration.ofSeconds(30)); ``` Learn more in the [workflow communication tutorial](../java/tutorials/workflow-communication.md#workflow-events). Events are persisted to the database, so they remain available even after the workflow completes. ### Task Queues → Queues Temporal task queues control which workers execute which workflows. DBOS queues serve a similar purpose but also provide built-in advanced concurrency control and rate limiting. **Temporal:** ```python # Worker listens to a task queue worker = Worker( client, task_queue="order-processing", workflows=[OrderWorkflow], activities=[validate_order, process_payment], ) await worker.run() # Start workflow on a specific queue handle = await client.start_workflow( OrderWorkflow.run, order, id="order-123", task_queue="order-processing" ) ``` **DBOS:** **Python** ```python # Register a queue with concurrency limits DBOS.register_queue("order-processing", global_concurrency=10) # Enqueue a workflow handle = DBOS.enqueue_workflow("order-processing", order_workflow, order) result = handle.get_result() ``` Learn more in the [queues tutorial](../python/tutorials/queue-tutorial.md). **TypeScript** ```typescript // Register a queue with concurrency limits await DBOS.registerQueue("order-processing", { globalConcurrency: 10 }); // Enqueue a workflow const handle = await DBOS.startWorkflow(orderWorkflowFn, { queueName: "order-processing" })(order); const result = await handle.getResult(); ``` Learn more in the [queues tutorial](../typescript/tutorials/queue-tutorial.md). **Go** ```go // Define a queue with concurrency limits queue := dbos.NewWorkflowQueue(dbosContext, "order-processing", dbos.WithGlobalConcurrency(10)) // Enqueue a workflow handle, err := dbos.RunWorkflow(dbosContext, OrderWorkflow, order, dbos.WithQueue(queue.Name)) result, err := handle.GetResult() ``` Learn more in the [queues tutorial](../golang/tutorials/queue-tutorial.md). **Java** ```java // Register a queue with concurrency limits (after dbos.launch()) dbos.registerQueue("order-processing", QueueOptions.setConcurrency(10)); // Enqueue a workflow WorkflowHandle handle = dbos.startWorkflow( () -> proxy.orderWorkflow(order), new StartWorkflowOptions().withQueue("order-processing") ); String result = handle.getResult(); ``` Learn more in the [queues tutorial](../java/tutorials/queue-tutorial.md). DBOS queues provide features that Temporal task queues don't have out of the box: - **Global concurrency limits**: Limit total concurrent executions across all workers. - **Per-worker concurrency**: Limit concurrent executions per process. - **Global rate limiting**: Limit executions per time period across all workers. - **Partitioned queues**: Create per-tenant sub-queues with independent concurrency limits. - **Priority**: Process higher-priority workflows first. - **Deduplication**: Prevent duplicate workflows in the queue. - **Debouncing**: Delay a workflow's execution until some time has passed since it was last called. ### Scheduled Workflows Both DBOS and Temporal let you run workflows on a cron schedule: **Temporal:** ```python await client.create_schedule( "daily-report", Schedule( action=ScheduleActionStartWorkflow( DailyReportWorkflow.run, id="daily-report", task_queue="reports", ), spec=ScheduleSpec(cron_expressions=["0 9 * * *"]), ), ) ``` **DBOS:** **Python** ```python DBOS.create_schedule( schedule_name="daily-report", workflow_fn=daily_report_workflow, schedule="0 9 * * *", ) ``` DBOS schedules also support pausing, resuming, backfilling missed runs, and triggering immediate execution. Learn more in the [scheduling tutorial](../python/tutorials/scheduled-workflows.md). **TypeScript** ```typescript await DBOS.createSchedule({ scheduleName: "daily-report", workflowFn: dailyReportWorkflow, schedule: "0 9 * * *", }); ``` DBOS schedules also support pausing, resuming, backfilling missed runs, and triggering immediate execution. Learn more in the [scheduling tutorial](../typescript/tutorials/scheduled-workflows.md). ### Child Workflows Both Temporal and DBOS support calling a child workflow from within another workflow. **Temporal:** ```python @workflow.defn class ParentWorkflow: @workflow.run async def run(self): result = await workflow.execute_child_workflow( ChildWorkflow.run, args=[data] ) ``` **DBOS:** **Python** ```python @DBOS.workflow() def parent_workflow(): # Call directly (runs inline) result = child_workflow(data) # Or start in background handle = DBOS.start_workflow(child_workflow, data) result = handle.get_result() ``` Learn more in the [workflows tutorial](../python/tutorials/workflow-tutorial.md#starting-workflows-in-the-background). **TypeScript** ```typescript async function parentWorkflow() { // Call directly (runs inline) const result = await childWorkflowFn(data); // Or start in background const handle = await DBOS.startWorkflow(childWorkflowFn)(data); const result2 = await handle.getResult(); } const parentWorkflowFn = DBOS.registerWorkflow(parentWorkflow); ``` Learn more in the [workflows tutorial](../typescript/tutorials/workflow-tutorial.md#starting-workflows-in-the-background). **Go** ```go func ParentWorkflow(ctx dbos.DBOSContext, input string) (string, error) { // Start child workflow in background handle, err := dbos.RunWorkflow(ctx, ChildWorkflow, data) if err != nil { return "", err } result, err := handle.GetResult() return result, err } ``` Learn more in the [workflows tutorial](../golang/tutorials/workflow-tutorial.md). **Java** ```java @Workflow(name = "parentWorkflow") public String parentWorkflow() { // Call directly (runs inline) String result = proxy.childWorkflow(data); // Or start in background WorkflowHandle handle = dbos.startWorkflow( () -> proxy.childWorkflow(data), new StartWorkflowOptions() ); return handle.getResult(); } ``` Learn more in the [workflows tutorial](../java/tutorials/workflow-tutorial.md#starting-workflows-in-the-background). ### Codecs and Encryption In Temporal, you can define a codec to encrypt workflow information before it is stored on a Temporal server to limit Temporal's access to sensitive data. In DBOS, this is rarely necessary because data is stored **only** in your own database. However, if it is necessary to store sensitive data encrypted, you can use a custom serializer ([Python](../python/reference/contexts.md#custom-serialization), [TypeScript](../typescript/reference/configuration.md#custom-serialization)) to encrypt your data before storing it and decrypt it before retrieving it. ### What's Different in DBOS #### No Orchestration Server DBOS has no central server to manage, operate, or scale. Your workflows run in your application process and checkpoint directly to your database. This eliminates a major source of operational complexity and latency. #### Fork DBOS can [fork a workflow](../python/tutorials/workflow-management.md) from a specific step, re-executing it from that point. This is powerful for recovering from failures, for example, restarting thousands of failed workflows from a specific step after an outage is resolved. #### Workflow Streaming DBOS provides [streaming](../python/tutorials/workflow-communication.md#workflow-streaming), an append-only stream that workflows can write to and clients can read from in real time. This is useful for streaming LLM outputs, progress updates, or real-time data from long-running workflows. #### Database Integration & SQL-Based Introspection DBOS integrates deeply with your database. For example, you can [enqueue workflows directly from Postgres PL/pgSQL function](./portable-workflows.md#per-workflow-enqueue). You can also use [transactional steps](#database-operations) to perform database operations in workflows with exactly-once semantics. Moreover, because all workflow state is stored in your database, you can query it with SQL. DBOS also provides programmatic APIs to [list, search, and manage workflows](../python/tutorials/workflow-management.md) by status, name, time, queue, or custom properties. #### Queue Flow Control Using DBOS queues, you can manage how many workflows can execute concurrently (globally, per-worker, and per-tenant) as well as which workers can execute which workflows. Temporal does not have comparable queueing or flow control abstractions, making it harder to control when and where workflows execute. ### Automating Temporal -> DBOS Migration With coding agents, you can largely automate a migration from Temporal to DBOS. To do this, we recommend using DBOS skills and prompts to give your coding agent access to the latest information on DBOS: - [AI-assisted development in Python](../python/prompting.md) - [AI-assisted development in TypeScript](../typescript/prompting.md) - [AI-assisted development in Go](../golang/prompting.md) - [AI-assisted development in Java](../java/prompting.md) --- ## Cross-Language Interaction DBOS supports multiple languages—Python, TypeScript, Go, and Java—each with its own SDK. A client in one language can connect to the [system database](./system-tables.md) of an application written in another language to exchange data through workflows, messages, events, and streams. Applications in different languages can also [share a single system database](./sharing-a-system-database.md). However, each language has a native serialization format that the other languages can't read. The **portable JSON** serialization format solves this by providing a common data representation that all SDKs can read and write, and can even be read and written from the database without any DBOS code at all. ### Default Serialization Is Language-Specific By default, each DBOS SDK serializes data using its language's default format. These default formats are chosen for their fidelity to the wide range of data structures and objects available in each language: | Language | Default Format | Format Name | |------------|---------------------|-----------------| | Python | pickle | `py_pickle` | | TypeScript | SuperJSON | `js_superjson` | | Go | encoding/json | `DBOS_JSON` | | Java | Jackson | `java_jackson` | As the set of data structures and classes varies from language to language, data written in one language's default format cannot be read by the other languages. For example, a Python workflow that writes an event using pickle produces a binary blob that TypeScript and Java can't deserialize. ### Portable JSON Format The `portable_json` format is straightforward use of JSON that all SDKs can read and write. While a smaller subset of language constructs can be serialized, any DBOS application in any language can read or write it. **Supported types:** - JSON primitives: `null`, booleans, numbers, and strings - JSON arrays (ordered lists of JSON values) - JSON objects (maps with strings as keys and JSON values) **Type conversions:** Some language built-in and library types are mapped to equivalent JSON constructs. When these values are decoded, the recipient must restore them to the appropriate language equivalent. - Date/time values are converted to [RFC 3339](https://www.rfc-editor.org/rfc/rfc3339) UTC strings (e.g., `"2025-06-15T14:30:00.000Z"`) | Language | Type | Portable Representation | |------------|------------------------|-------------------------| | Python | `datetime` | RFC 3339 UTC string | | Python | `date` | ISO 8601 string | | Python | `Decimal` | Numeric string | | Python | `set`, `tuple` | JSON array | | Go | `time.Time` | RFC 3339 UTC string | | Java | `Instant` | RFC 3339 UTC string | | Java | `BigDecimal` | Numeric string | | TypeScript | `Date` | RFC 3339 UTC string | | TypeScript | `BigInt` | Numeric string | | TypeScript | `Map` (string keys) | JSON object | | TypeScript | `Set` | JSON array | ### Using Portable Serialization You can opt in to portable serialization at the workflow or operation level. Workflows started with portable serialization return their results or exceptions in portable format. Workflows started with portable serialization also write their events and streams in portable JSON by default, but this can be overridden for each operation. #### Per-Workflow (Enqueue) When enqueuing or starting a workflow from a `DBOSClient`, or when enqueueing a workflow to another application [sharing the same system database](./sharing-a-system-database.md), set the serialization format in the enqueue options. This ensures the workflow's arguments are serialized in portable format that can be read by the target language. If multiple applications [share the system database](./sharing-a-system-database.md), also name the application that owns the workflow, so that application runs it. You can also enqueue a workflow using the PL/pgSQL function [`dbos.enqueue_workflow`](system-tables.md#dbosenqueue_workflow). Only portable serialization is allowed when enqueuing using PL/pgSQL. **Python** ```python from dbos import DBOSClient, WorkflowSerializationFormat client = DBOSClient( system_database_url=db_url, # The name of the application that implements process_order application_name="order-service", ) handle = client.enqueue( { "workflow_name": "process_order", "queue_name": "orders", "serialization_type": WorkflowSerializationFormat.PORTABLE, }, "order-123", ) ``` **TypeScript** ```typescript import { DBOSClient } from "@dbos-inc/dbos-sdk"; const client = await DBOSClient.create({ systemDatabaseUrl: process.env.DBOS_SYSTEM_DATABASE_URL!, // The name of the application that implements process_order applicationName: "order-service", }); const handle = await client.enqueue( { workflowName: "process_order", queueName: "orders", serializationType: "portable", }, "order-123", ); ``` **Java** ```java import dev.dbos.transact.DBOSClient; import dev.dbos.transact.EnqueueOptions; import dev.dbos.transact.workflow.QueueName; import dev.dbos.transact.workflow.SerializationStrategy; var client = new DBOSClient(dbUrl, dbUser, dbPassword); var options = new EnqueueOptions("process_order", QueueName.of("orders")) // The name of the application that implements process_order .withApplicationName("order-service") .withSerialization(SerializationStrategy.PORTABLE); var handle = client.enqueueWorkflow(options, new Object[] {"order-123"}); ``` **Go** ```go import "github.com/dbos-inc/dbos-transact-golang/dbos" client, _ := dbos.NewClient(context.Background(), dbos.ClientConfig{ DatabaseURL: os.Getenv("DBOS_SYSTEM_DATABASE_URL"), // The name of the application that implements process_order AppName: "order-service", }) // In Go, use dbos.PortableWorkflowArgs to request a portable enqueue args := dbos.PortableWorkflowArgs{ PositionalArgs: []any{"order-123"}, } handle, err := dbos.Enqueue[any]( client, "orders", "process_order", args, ) ``` **PL/pgSQL** ```sql DECLARE workflow_id text; workflow_id := dbos.enqueue_workflow( workflow_name => 'processOrder', class_name => 'com.example.OrderProcessor', queue_name => 'orders', positional_args => ARRAY['"order-123"'::json] ); ``` #### Per-Workflow (via Annotation or Decorator) You can set the serialization strategy directly on the workflow annotation or decorator so that the workflow uses portable serialization by default when started: **Python** ```python from dbos import DBOS, WorkflowSerializationFormat @DBOS.workflow(serialization_type=WorkflowSerializationFormat.PORTABLE) def process_order(order_id: str): # All inputs, outputs, events, and streams for this workflow # use portable JSON serialization by default return f"processed: {order_id}" ``` **TypeScript** Using a decorator: ```typescript import { DBOS } from "@dbos-inc/dbos-sdk"; export class Orders { @DBOS.workflow({ serialization: "portable" }) static async processOrder(orderId: string): Promise { // All inputs, outputs, events, and streams for this workflow // use portable JSON serialization by default return `processed: ${orderId}`; } } ``` Or using `registerWorkflow`: ```typescript async function processOrder(orderId: string): Promise { return `processed: ${orderId}`; } const processOrderWorkflow = DBOS.registerWorkflow(processOrder, { name: "processOrder", serialization: "portable", }); ``` **Go** In Go, portable serialization is set per-invocation using the `WithPortableWorkflow` option on `RunWorkflow`: ```go handle, err := dbos.RunWorkflow(dbosContext, processOrder, "order-123", dbos.WithPortableWorkflow(), ) ``` **Java** ```java import dev.dbos.transact.workflow.SerializationStrategy; import dev.dbos.transact.workflow.Workflow; @Workflow(serializationStrategy = SerializationStrategy.PORTABLE) public String processOrder(String orderId) { // All inputs, outputs, events, and streams for this workflow // use portable JSON serialization by default return "processed: " + orderId; } ``` :::note The default serialization strategy only affects invocations that are aware of the annotation / decorator. This makes the default useful for unit testing, but the actual serialization strategy used will depend on how the workflow is enqueued by the client. ::: #### For Workflow Communication Setting the serialization format at the workflow level affects the default for `setEvent` and `writeStream`. However, individual operations can override this—for example, a workflow running with native serialization may want to publish a specific event in portable format for cross-language consumption, or a portable workflow may need to record an event with the greater flexibility afforded by the native serializer. Each language's `setEvent` and `writeStream` methods accept a serialization parameter for this purpose. `send` is a special case, because messages target a different workflow and the sender does not know what serialization that workflow expects. In every language, a `send` from inside a workflow defaults to that workflow's serialization format. You should therefore always set the serialization format explicitly on `send` when communicating cross-language. You can also send a message to a workflow using the PL/pgSQL function [`dbos.send_message`](system-tables.md#dbossend_message). Only portable serialization is allowed when sending a message using PL/pgSQL. Note, there is no PL/pgSQL version of `setEvent` or `writeStream`. :::info Step outputs always use the native serializer regardless of the workflow's serialization strategy. Steps are internal to a workflow and are not read by other languages, so the native serializer's greater flexibility is preferred. ::: **Python** ```python from dbos import DBOS, WorkflowSerializationFormat # Send a message readable by any language DBOS.send( destination_id="workflow-123", message={"status": "complete", "count": 42}, topic="updates", serialization_type=WorkflowSerializationFormat.PORTABLE, ) # Set an event readable by any language DBOS.set_event( "progress", {"percent": 75}, serialization_type=WorkflowSerializationFormat.PORTABLE, ) # Write to a stream readable by any language DBOS.write_stream( "results", {"item": "processed"}, serialization_type=WorkflowSerializationFormat.PORTABLE, ) ``` **TypeScript** ```typescript import { DBOS } from "@dbos-inc/dbos-sdk"; // Send a message readable by any language await DBOS.send( "workflow-123", { status: "complete", count: 42 }, "updates", undefined, // idempotencyKey { serializationType: "portable" } ); // Set an event readable by any language await DBOS.setEvent( "progress", { percent: 75 }, { serializationType: "portable" } ); // Write to a stream readable by any language await DBOS.writeStream( "results", { item: "processed" }, { serializationType: "portable" } ); ``` **Go** ```go import "github.com/dbos-inc/dbos-transact-golang/dbos" // Send a message readable by any language dbos.Send(ctx, "workflow-123", map[string]any{"status": "complete", "count": 42}, "updates", dbos.WithPortableSend(), ) // Set an event readable by any language dbos.SetEvent(ctx, "progress", map[string]any{"percent": 75}, dbos.WithPortableSetEvent(), ) // Write to a stream readable by any language dbos.WriteStream(ctx, "results", map[string]any{"item": "processed"}, dbos.WithPortableWriteStream(), ) ``` **Java** ```java import dev.dbos.transact.DBOS; import dev.dbos.transact.workflow.SerializationStrategy; // Send a message readable by any language dbos.send( "workflow-123", Map.of("status", "complete", "count", 42), "updates", null, // idempotencyKey SerializationStrategy.PORTABLE ); // Set an event readable by any language dbos.setEvent( "progress", Map.of("percent", 75), SerializationStrategy.PORTABLE ); ``` **PL/pgSQL** ```sql PERFORM dbos.send_message( destination_id => 'workflow-123', message => '{"status": "complete", "count": 42}'::json, topic => 'updates' ); ``` ### Portable Errors When a workflow using portable serialization fails, its error is serialized in a standard JSON structure that all languages can inspect: ```json { "name": "ValueError", "message": "Order not found", "code": 404, "data": {"orderId": "order-123"} } ``` | Field | Type | Description | |-----------|----------------------|----------------------------------------| | `name` | string | The error type/class name | | `message` | string | Human-readable error message | | `code` | number, string, null | Optional application-specific error code | | `data` | any JSON value, null | Optional structured error details | #### Raising Portable Errors You can explicitly raise a portable error from a workflow: **Python** ```python from dbos import PortableWorkflowError raise PortableWorkflowError( message="Order not found", name="NotFoundError", code=404, data={"orderId": "order-123"}, ) ``` **TypeScript** ```typescript import { PortableWorkflowError } from "@dbos-inc/dbos-sdk"; throw new PortableWorkflowError( "Order not found", "NotFoundError", 404, { orderId: "order-123" }, ); ``` **Go** ```go import "github.com/dbos-inc/dbos-transact-golang/dbos" return nil, &dbos.PortableWorkflowError{ Name: "NotFoundError", Message: "Order not found", Code: 404, Data: map[string]any{"orderId": "order-123"}, } ``` **Java** ```java import dev.dbos.transact.json.PortableWorkflowException; throw new PortableWorkflowException( "Order not found", "NotFoundError", 404, Map.of("orderId", "order-123") ); ``` #### Reading Portable Errors When a workflow that used portable serialization fails, other languages receive the error as a `PortableWorkflowError` (Python/TS/Go) or `PortableWorkflowException` (Java) with the `name`, `message`, `code`, and `data` fields populated. If a workflow fails with a non-portable exception while using portable serialization, DBOS automatically converts it to the portable error format on a best-effort basis, extracting the error type name, message, and any common error code attributes. ### Input Validation and Coercion When a workflow is started via portable JSON—whether from another language, a `DBOSClient`, or a direct database insert—the arguments arrive as plain JSON values. JSON has a limited type system: numbers are untyped (no distinction between `int`, `long`, `double`, or other language-specific offerings), there is no native date type (dates arrive as strings), and collection types may not match the target language's expectations (e.g., a JSON array becomes a generic `ArrayList` in Java, not a typed list or object array). Each SDK provides a way to validate these arguments so that the workflow function receives the types it expects. Note that while workflow argument validation is possible, return values, messages, and events are not automatically coerced, as the expected types are not known at runtime. These must be validated and coerced manually. Each SDK's approach is documented in its language-specific reference: - **[Java — Automatic Coercion](../java/reference/workflows-steps.md#input-validation-and-coercion)**: Java automatically coerces portable JSON arguments to match the workflow method's parameter types (e.g., `Integer` → `long`, ISO-8601 strings → `Instant`). No opt-in required. - **[TypeScript — Input Schema (Zod)](../typescript/reference/workflows-steps.md#input-validation-and-coercion)**: TypeScript workflows can specify an `inputSchema` (compatible with [Zod](https://zod.dev/)) that validates and optionally transforms arguments before the workflow runs. - **Go — Automatic Coercion**: Go automatically coerces portable JSON arguments to match the workflow function's parameter types using type assertion. No opt-in required. - **[Python — Argument Validator (Pydantic)](../python/reference/decorators.md#input-validation-and-coercion)**: Python workflows can specify `validate_args=pydantic_args_validator` to validate arguments against the function's type hints using [Pydantic](https://docs.pydantic.dev/). ### Further Reading - **Serialization strategy reference:** - [Python Serialization Strategy](../python/reference/contexts.md#serialization-strategy) - [TypeScript Serialization Strategy](../typescript/reference/methods.md#serialization-strategy) - [Go Portable Options](../golang/reference/methods.md#portable-serialization-options-and-types) - [Java Serialization Strategy](../java/reference/methods.md#serialization-strategy) - **Custom serialization configuration:** - [Python Custom Serialization](../python/reference/contexts.md#custom-serialization) - [TypeScript Custom Serialization](../typescript/reference/configuration.md#custom-serialization) - [Java Custom Serialization](../java/reference/lifecycle.md#custom-serialization) - **System tables:** The [`serialization` column](./system-tables.md) in system tables records which format was used for each piece of serialized data. --- ## Sharing a System Database Multiple DBOS applications, potentially in different languages, can share a single system database. Each application is identified by its configured name and owns everything it creates: workflows, steps, queues, schedules, and application versions. Applications sharing a system database are isolated from one another by default, but can freely interoperate by naming each other. For example, one application can enqueue another's workflows and wait for their results. ### Application Names and Ownership Every application is identified by the `name` in its configuration, so each application sharing a system database must have a distinct name. Ownership determines which application runs what: - A workflow is dequeued, run, and recovered only by the application that owns it. - A queue is polled only by the application that registered it, even if another application enqueues workflows on it. - A schedule is fired only by the application that created it, and its workflows are owned by that application. - Application versions are tracked per application, so one application's deployments do not affect which version its peers consider latest. [Retention policies](../conductor/retention.md) are an exception: their time and rows thresholds apply to the entire system database, including workflows owned by other applications. The global timeout remains scoped to the application that configures it. Queue, schedule, and version names remain globally unique across all applications sharing a system database; registering a name that a different application already owns raises an error. Workflow IDs are also unique across the entire system database, so ID-addressed operations (retrieving a workflow's handle, status, or result by ID, and sending messages or reading events and streams) work across applications regardless of ownership. Observability queries (`list_workflows`, `list_queues`, `list_schedules`) are scoped to the calling application by default. ### Calling Another Application's Workflows To run another application's workflow, enqueue it by name, naming the application that implements it. The enqueued workflow is owned by the target application, which dequeues and runs it on its latest application version. Because workflow IDs are global, you can then wait for the result from the returned handle. **Python** ```python from dbos import DBOS, EnqueueOptions options: EnqueueOptions = { "workflow_name": "process_order", "queue_name": "orders", # The name of the application that implements process_order "application_name": "order-service", } handle = DBOS.enqueue_workflow_with_options(options, "order-123") result = handle.get_result() ``` **TypeScript** ```typescript const handle = await DBOS.enqueueWorkflowWithOptions({ workflowName: "process_order", queueName: "orders", // The name of the application that implements process_order applicationName: "order-service", }, "order-123"); const result = await handle.getResult(); ``` **Go** ```go handle, err := dbos.Enqueue[any](ctx, "orders", "process_order", "order-123", // The name of the application that implements process_order dbos.WithEnqueueApplicationName("order-service"), ) if err != nil { return err } result, err := handle.GetResult() ``` **Java** ```java var options = new EnqueueOptions("process_order", QueueName.of("orders")) // The name of the application that implements process_order .withApplicationName("order-service"); WorkflowHandle handle = dbos.enqueueWorkflow(options, new Object[] {"order-123"}); Object result = handle.getResult(); ``` If the applications are written in different languages, also set the serialization type to portable so the target application can read the arguments. See [Cross-Language Interaction](./portable-workflows.md) for details. You can do the same from a DBOS Client ([Python](../python/reference/client.md), [Java](../java/reference/client.md)), which additionally supports registering queues, creating schedules, and debouncing workflows on behalf of a named application. Always set the client's application name if multiple applications share a system database. ### Unowned Rows Workflows, queues, schedules, and versions created before upgrading to a DBOS version supporting application names, or created by a client with no application name, are owned by no application. Every application treats unowned rows as its own: any application may dequeue an unowned workflow (claiming ownership of it when it does), and every application fires unowned schedules and polls unowned queues. :::warning To ensure no rows are unowned, always pass an application name to your clients when interacting with a system database shared by multiple applications ::: Before adding a second application to a system database, explicitly transfer ownership of any unowned rows to the first application: ```shell dbosctl sysdb rename-application --to my-app --adopt-unclaimed-rows ``` ### Renaming an Application Because ownership is recorded under the application's name, renaming an application requires transferring ownership of its rows. To rename an application, first stop it, then run [`dbosctl sysdb rename-application`](../conductor/reference/dbosctl.md#dbosctl-sysdb-rename-application), then restart it under its new name: ```shell dbosctl sysdb rename-application --from old-name --to new-name --db-url ``` --- ## DBOS System Database DBOS records application execution history in several system tables. These tables are located in your system database, whose location you configure when you launch your application. A PostgreSQL system database also includes PL/pgSQL functions that you can call from stored procedures or triggers. ### System Database Functions :::info[Reminder] PL/pgSQL functions can only be called from code running in the same database. ::: :::info Releases of DBOS Transact periodically update the system database functions, typically to add new parameters. All changes are backwards-compatible. Updates are applied by atomically dropping and recreating the functions. Therefore, it is not recommended to create database objects whose dependency on these functions [is tracked by Postgres](https://www.postgresql.org/docs/current/ddl-depend.html), such as views or SQL-standard-body (`BEGIN ATOMIC`) functions. ::: #### dbos.enqueue_workflow ```sql CREATE FUNCTION dbos.enqueue_workflow( workflow_name TEXT, queue_name TEXT, positional_args JSON[] DEFAULT ARRAY[]::JSON[], named_args JSON DEFAULT '{}'::JSON, class_name TEXT DEFAULT NULL, config_name TEXT DEFAULT NULL, workflow_id TEXT DEFAULT NULL, app_version TEXT DEFAULT NULL, timeout_ms BIGINT DEFAULT NULL, deadline_epoch_ms BIGINT DEFAULT NULL, deduplication_id TEXT DEFAULT NULL, priority INTEGER DEFAULT NULL, queue_partition_key TEXT DEFAULT NULL, authenticated_user TEXT DEFAULT NULL, authenticated_roles TEXT DEFAULT NULL, delay_until_epoch_ms BIGINT DEFAULT NULL, application_name TEXT DEFAULT NULL ) RETURNS TEXT ``` PL/pgSQL function for enqueuing a workflow on a [durable queue](../architecture.md#durable-queues). **Parameters:** - `workflow_name`: The name of workflow to enqueue. - `queue_name`: The durable queue on which to enqueue this workflow. - `positional_args`: An array of positional parameters for the enqueued workflow. Must use [Portable JSON Format](portable-workflows.md#portable-json-format). Defaults to an empty array - `named_args`: The named parameters (for languages that support them, like Python). Must use [Portable JSON Format](portable-workflows.md#portable-json-format) and be a JSON object. Defaults to an empty object (`{}`). - `class_name`: The class name of workflow to enqueue. Defaults to null. - `config_name`: The config name of workflow to enqueue. For languages that support it, this is usually exposed as workflow class instance name. Defaults to null. - `workflow_id`: Specify the idempotency ID to assign to the enqueued workflow. If left undefined, a random UUID is generated. - `app_version`: The version of your application that should process this workflow. If left undefined, it will be updated to the current version when the workflow is first dequeued by a process running the latest application version. - `timeout_ms`: Set a timeout for the enqueued workflow. When the timeout expires, the workflow and all its children are cancelled. The timeout does not begin until the workflow is dequeued and starts execution. - `deadline_epoch_ms`: Set a deadline for the enqueued workflow. If the workflow is executing when the deadline arrives, the workflow and all its children are cancelled. - `deduplication_id`: At any given time, only one workflow with a specific deduplication ID can be enqueued in the specified queue. If a workflow with a deduplication ID is currently enqueued, delayed, or actively executing (status ENQUEUED, DELAYED, or PENDING), subsequent workflow enqueue attempt with the same deduplication ID in the same queue will raise an exception. - `priority`: The priority of the enqueued workflow in the specified queue. Workflows with the same priority are dequeued in FIFO (first in, first out) order. Priority values can range from 0 to 2,147,483,647, where a low number indicates a higher priority. If a priority is not supplied, it defaults to 0, the highest priority. - `queue_partition_key`: Set a queue partition key for the workflow. Use if and only if the queue is partitioned. Partitioned queues apply their per-partition flow control limits (concurrency and rate limits) to each partition separately. - `authenticated_user`: The authenticated user to associate with the enqueued workflow. Defaults to null. - `authenticated_roles`: The authenticated roles to associate with the enqueued workflow, as a JSON-encoded array of strings (e.g. `'["admin", "reader"]'`). Defaults to null. - `delay_until_epoch_ms`: A Unix epoch timestamp (in milliseconds) before which the workflow should not be dequeued. The workflow is enqueued with `DELAYED` status until this time arrives, after which it becomes `ENQUEUED` and eligible to run. Must be `>= 0`. Defaults to null (the workflow is enqueued immediately). - `application_name`: The application that owns and runs the enqueued workflow. #### dbos.send_message ```sql CREATE FUNCTION dbos.send_message( destination_id TEXT, message JSON, topic TEXT DEFAULT NULL, message_id TEXT DEFAULT NULL ) RETURNS VOID ``` PL/pgSQL for sending a message to a workflow, similar to `DBOS.send`. Messages can optionally be associated with a topic. **Parameters:** - `destination_id`: The workflow to which to send the message. - `message`: The message to send. Must use [Portable JSON Format](portable-workflows.md#portable-json-format). - `topic`: A topic with which to associate the message. Messages are enqueued per-topic on the receiver. - `message_id`: A unique ID for the message. If a message ID is set, the message will only be sent once no matter how many times `send_message` is called with this ID. ### System Database Tables #### dbos.workflow_status This table stores workflow execution information. Each row represents a different workflow execution. **Columns:** - **workflow_uuid**: The unique identifier of the workflow execution. - **status**: The status of the workflow execution. One of `PENDING`, `SUCCESS`, `ERROR`, `MAX_RECOVERY_ATTEMPTS_EXCEEDED`, `ENQUEUED`, `DELAYED`, or `CANCELLED`. - **name**: The name of the workflow function. - **authenticated_user**: The user who ran the workflow. Empty string if not set. - **assumed_role**: The role used to run this workflow. Empty string if authorization is not required. - **authenticated_roles**: All roles the authenticated user has, if any. - **created_at**: The epoch timestamp of when this workflow was created (enqueued or started). - **updated_at**: The latest epoch timestamp when this workflow status was updated. - **application_version**: The application version of this workflow code. - **class_name**: The class name of the workflow function. - **config_name**: The name of the configured instance of this workflow, if any. - **recovery_attempts**: The number of attempts (so far) to recover this workflow. - **queue_name**: If this workflow is or was enqueued, the name of the queue. - **executor_id**: The ID of the executor that ran this workflow. - **workflow_timeout_ms**: The timeout of the workflow, if specified. - **workflow_deadline_epoch_ms**: The deadline at which the workflow times out, if the workflow has a timeout. Derived when the workflow starts by adding the timeout to the workflow start time (which may be different than the creation time for enqueued workflows). - **started_at_epoch_ms**: If this workflow was enqueued, the time at which it was dequeued and began execution. - **deduplication_id**: The deduplication key for this workflow, if any. - **priority**: The priority of this workflow on its queue, if enqueued. Defaults to 0 if not specified. Lower priorities execute first. - **queue_partition_key**: The key associated with the workflow, if on a partitioned queue. - **forked_from**: The ID of the workflow that this was forked from, if applicable. - **was_forked_from**: Whether this workflow has ever been forked from by another workflow. - **parent_workflow_id**: The ID of the parent workflow, if this workflow was started as a child of another workflow. - **delay_until_epoch_ms**: For workflows in the `DELAYED` state, the epoch timestamp at which the workflow should transition to `ENQUEUED` and become eligible for execution. - **owner_xid**: The ownership token of the execution currently running this workflow, used to [detect concurrent executions](./concurrent-executions.md#workflow-ownership). Null when no execution owns the workflow. - **creator_xid**: Internal transaction ID used to prevent duplicate workflow starts. - **application_id**: Internal field used only in DBOS Cloud. - **serialization**: The name of the serialization format used for this workflow's inputs, output, and error (e.g. `java_jackson`, `py_pickle`, `portable_json`). Null if the default serializer was used. - **rate_limited**: Whether this workflow was dequeued from a rate-limited queue. - **completed_at**: The epoch timestamp (in milliseconds) at which this workflow reached a terminal state (`SUCCESS`, `ERROR`, `CANCELLED`, or `MAX_RECOVERY_ATTEMPTS_EXCEEDED`). Null while the workflow is still active. - **attributes**: Custom key-value attributes attached to this workflow, if any. Stored in Postgres as GIN-indexed JSONB, so workflows can be efficiently searched by attribute. - **schedule_name**: If this workflow was started by a [scheduled workflow](#dbosworkflow_schedules), the name of its schedule. - **debounce_deadline_epoch_ms**: If this workflow is debounced with a debounce timeout, the epoch timestamp past which its execution can no longer be delayed. - **is_debounced**: Whether this workflow was created by a debouncer. - **application_name**: The application that owns this workflow. #### dbos.workflow_input This table stores workflow inputs. Each row represents a different workflow execution and is written when the workflow is created. **Columns:** - **workflow_uuid**: The unique identifier of the workflow execution. - **inputs**: The serialized inputs of the workflow execution. - **retention_timestamp**: The epoch timestamp (in milliseconds) when this row was written. Used when applying [retention policies](../conductor/retention.md). #### dbos.workflow_output This table stores workflow outputs. Each row represents a different workflow execution and is written when the workflow completes. **Columns:** - **workflow_uuid**: The unique identifier of the workflow execution. - **output**: The serialized workflow output, if any. - **error**: The serialized error thrown by the workflow, if any. - **retention_timestamp**: The epoch timestamp (in milliseconds) when this row was written. Used when applying [retention policies](../conductor/retention.md). #### dbos.operation_outputs This table stores the outputs of workflow steps. Each row represents a different workflow step execution. Executions of DBOS methods like `DBOS.sleep` and `DBOS.send` are also recorded here as steps, as is enqueueing or starting a child workflow. **Columns:** - **workflow_uuid**: The unique identifier of the workflow execution this function belongs to. - **function_id**: The monotonically increasing ID of the step (starts from 0, or from 1 in Python) within the workflow, based on the order in which steps execute. - **function_name**: The name of the step. - **output**: The serialized step output, if any. - **error**: The serialized error thrown by the step, if any. - **child_workflow_id**: If the step starts a new child workflow, its ID. - **started_at_epoch_ms**: The epoch timestamp of when this step started execution. - **completed_at_epoch_ms**: The epoch timestamp of when this step completed. - **serialization**: The name of the serialization format used for this step's output and error. Null if the workflow's default serializer was used. - **application_name**: The application that ran this step. - **retention_timestamp**: The epoch timestamp (in milliseconds) when this row was written. Used when applying [retention policies](../conductor/retention.md). #### dbos.notifications This table stores workflow messages/notifications. Each entry represents a different message. **Columns:** - **destination_uuid**: The ID of the workflow to which the message is sent. - **topic**: The topic to which the message is sent. - **message**: The serialized contents of the message. - **created_at_epoch_ms**: The epoch timestamp when this message was created. - **message_uuid**: The unique ID of the message. - **serialization**: The name of the serialization format used for the message. Null if the default serializer was used. - **consumed**: Whether the message has been consumed by a `DBOS.recv` call. - **consumed_by_function_id**: The ID of the step (the `DBOS.recv` call) that consumed the message, if it has been consumed. #### dbos.workflow_events This table stores workflow events. Each entry represents a different event. **Columns:** - **workflow_uuid**: The ID of the workflow that published this event. - **key**: The key of the event. - **value**: The serialized value of the event. - **serialization**: The name of the serialization format used for the event value. Null if the default serializer was used. #### dbos.workflow_events_history This table stores historic changes to workflow events over time. Each entry represents a distinct value of a workflow event during the workflow lifetime. **Columns:** - **workflow_uuid**: The ID of the workflow that published this event. - **function_id**: The monotonically increasing ID of the step that set this value. - **key**: The key of the event. - **value**: The serialized value of the event. - **serialization**: The name of the serialization format used for the event value. Null if the default serializer was used. #### dbos.streams This table stores workflow streams. Each entry represents a different message in a stream. **Columns:** - **workflow_uuid**: The ID of the workflow that wrote this stream message. - **key**: The key of the stream. - **value**: The serialized value of the message. - **offset**: The offset of the message in the stream (the first message written has offset 0, the second offset 1, and so on). - **function_id**: The monotonically increasing step ID responsible for emitting this stream. - **serialization**: The name of the serialization format used for the stream value. Null if the default serializer was used. #### dbos.application_versions This table stores registered application versions. Each time DBOS launches, it records the current application version. The latest version is determined by the highest timestamp. **Columns:** - **version_id**: A unique ID for this version. - **version_name**: The unique name of this version. - **version_timestamp**: The epoch timestamp (in milliseconds) of this version. Used to determine the latest version. - **created_at**: The epoch timestamp (in milliseconds) when this version was first registered. - **application_name**: The application that registered this version. Version names remain unique across all applications sharing a system database. #### dbos.workflow_schedules This table stores scheduled workflow definitions. Each entry represents a different scheduled workflow. **Columns:** - **schedule_id**: The unique identifier of the schedule. - **schedule_name**: The human-readable name of the schedule. Must be unique. - **workflow_name**: The name of the workflow function to execute on schedule. - **workflow_class_name**: The class name of the workflow function, if it is a class method. - **schedule**: The cron expression or schedule definition. - **status**: The status of the schedule. One of `ACTIVE` or `PAUSED`. Defaults to `ACTIVE`. - **context**: The serialized schedule context. - **last_fired_at**: The timestamp of when the schedule last fired. - **automatic_backfill**: Whether the schedule should automatically backfill missed executions on startup. - **cron_timezone**: The IANA timezone name in which the cron expression is evaluated. - **queue_name**: The name of the durable queue on which scheduled workflow invocations are enqueued, if any. - **application_name**: The application that owns this schedule and runs its workflows. #### dbos.queues This table stores [durable queue](../architecture.md#durable-queues) definitions. Each row represents a different queue. A queue is partitioned if any of its `partition_*` limits is set. **Columns:** - **queue_id**: A unique ID for this queue. - **name**: The queue's unique name. - **concurrency**: The maximum number of workflows from this queue that may run concurrently across all DBOS processes. `NULL` means unlimited. - **worker_concurrency**: The maximum number of workflows from this queue that may run concurrently on a single DBOS process. `NULL` means unlimited. - **rate_limit_max**: If a rate limit is set, the maximum number of workflows that may be started in a period. - **rate_limit_period_sec**: If a rate limit is set, the length of the period in seconds. - **partition_concurrency**: The maximum number of workflows from any one partition of this queue that may run concurrently across all DBOS processes. `NULL` means unlimited. - **partition_worker_concurrency**: The maximum number of workflows from any one partition of this queue that may run concurrently on a single DBOS process. `NULL` means unlimited. - **partition_rate_limit_max**: If a per-partition rate limit is set, the maximum number of workflows that may be started from any one partition in a period. - **partition_rate_limit_period_sec**: If a per-partition rate limit is set, the length of the period in seconds. - **polling_interval_sec**: The interval at which workers poll the database for new workflows on this queue. - **created_at**: The epoch timestamp (in milliseconds) when this queue was first registered. - **updated_at**: The epoch timestamp (in milliseconds) when this queue's configuration was last updated. - **application_name**: The application that owns this queue and dequeues workflows from it. #### dbos.dbos_migrations This table tracks which DBOS system database schema migrations have been applied. **Columns:** - **version**: The version number of the most recently applied system database schema migration. --- ## Troubleshooting & FAQ #### Where do I find the DBOS tables? DBOS checkpoints information about your workflows in an isolated _system database_ in your Postgres database server. You connect to this database through the `system_database_url` parameter in DBOS configuration. You can connect to and explore your system database with popular database clients like [psql](https://www.postgresql.org/docs/current/app-psql.html) and [DBeaver](https://dbeaver.io/). Note that the tables are in the `dbos` schema in that database, so the tables are accessible at `dbos.workflow_status`, `dbos.operation_outputs`, etc. The system database schema is documented [here](./explanations/system-tables.md). :::tip If you're using Supabase, only the `postgres` database is visible from the Supabase web console. You can use `postgres` as your system database, or you can use a different system database and connect to and explore your system database using a client like [psql](https://www.postgresql.org/docs/current/app-psql.html) or [DBeaver](https://dbeaver.io/). ::: #### What size of database should I use for DBOS? For a typical workload, processing 1000 actions (steps or workflows) per second (2 billion actions per month) with DBOS requires 4 Postgres vCPUs. This is a conservative estimate that leaves headroom for unexpected bursts or spikes. When sizing your database for DBOS, we recommend using that number (scaled to your actual workload size) as a starting point then measuring usage in practice. #### Why is my queue stuck? If a DBOS queue is stuck (workflows are not moving from `ENQUEUED` to `PENDING`), it is likely that either the number of `PENDING` workflows exceeds the queue's global "concurrency" limit or the number of queued workflows in a `PENDING` state on each worker exceeds the queue's "worker concurrency" limit. In either case, new tasks cannot be dequeued until some currently executing tasks complete or are cancelled. You can view all tasks executing on a queue from the "Queues" tab of the [DBOS Console](./conductor/workflow-management.md) If you need to, you can cancel tasks to remove them from the queue. #### Why is my workflow not finishing? The most common cause of "stuck" workflows is logic issues: infinite loops, indefinitely waiting for an event, or improper use of async in Python or TypeScript. The last is always worth checking when using those languages: any synchronous call anywhere in your program can block your event loop, preventing async operations (such as workflows) from making progress. If a worker crash or outage occurred, it may briefly delay workflow completion. In certain rare cases, you may need to allow up to 15 minutes for Conductor to begin workflow recovery. If workflows do not recover after a code upgrade, the cause is often [version mismatch](./architecture.md#upgrading-workflow-code). If you are using versioning, check that your app version matches the version of your workflow. #### How can I cancel or fork a large number of workflows in a batch? On the [DBOS Console](./conductor/workflow-management.md), filter for all workflows that meet your criteria, then select them all and apply a batch operation. Alternatively, write a script using the DBOS Client ([Python](./python/reference/client.md), [TypeScript](./typescript/reference/client.md), [Go](./golang/reference/dbos-context.md#newclient), [Java](./java/reference/client.md)) to list all the workflows that fit your criteria, then process them. #### Why am I seeing errors that objects cannot be deserialized? DBOS requires that the inputs and outputs of workflows, as well as the outputs of steps, are **serializable**. This is because DBOS checkpoints these inputs and outputs to the database to recover workflows from failures. DBOS serializes objects with [SuperJSON](https://github.com/flightcontrolhq/superjson) in TypeScript, `encoding/json` in Go, `pickle` in Python, and Jackson in Java. If your workflow needs to access an unserializable object like a database connection or API client, do not pass it into the workflow as an argument. Instead, either construct the object inside the workflow from parameters passed into the workflow, or construct it globally. #### How large can serialized step and workflow outputs be? DBOS stores two serialized fields (`inputs` and `output`) for each workflow and one `output` field for each step. Each of these is stored as a Postgres `TEXT` value which is limited by the maximum field size; currently 1GB. See [Postgres documentation](https://wiki.postgresql.org/wiki/FAQ#What_is_the_maximum_size_for_a_row.2C_a_table.2C_and_a_database.3F). #### Why am I seeing an error that function X was recorded when Y was expected? This error arises when DBOS is recovering a workflow and attempts to execute step Y, but finds a checkpoint in the database for step X instead. Typically, this occurs because the workflow function is not **deterministic**. A workflow function is deterministic if, when called multiple times with the same inputs, it invokes the same steps with the same inputs in the same order (given the same return values from those steps). If a workflow is non-deterministic, it may execute different steps during recovery than it did during its original execution. To make a workflow deterministic, make sure all non-deterministic operations (such as calling a third-party API, generating a random number, or getting the local time) are performed **in steps** instead of in the workflow function. #### Can I call a workflow from a workflow? Yes, you can call (or start, or enqueue) a workflow from inside another workflow. That workflow becomes a **child** of its caller and is by default assigned a workflow ID derived from its parent's. If you view a workflow's trace from the [DBOS Console](./conductor/workflow-management.md), it will include the workflow's children. #### Can I call a step from a step? Yes, you can call a step from another step. However, the called step becomes part of the calling step's execution rather than functioning as a separate step. #### Can I start, monitor, or cancel DBOS workflows from a non-DBOS application? Yes, your non-DBOS application can create a DBOS Client ([Python](./python/reference/client.md), [TypeScript](./typescript/reference/client.md), [Go](./golang/reference/dbos-context.md#newclient), [Java](./java/reference/client.md)) and use it to enqueue a workflow in your DBOS application and interact with it or check its status. #### What happens if you start two workflows with the same workflow ID? In DBOS, workflow IDs are unique identifiers of workflow executions. If you enqueue a workflow with the ID of a workflow that already exists, it's a no-op and a handle to the existing workflow execution is returned. #### How can I reset all my DBOS state during development? You can reset your DBOS system database and all internal DBOS state with the [`dbosctl sysdb reset`](./conductor/reference/dbosctl.md#dbosctl-sysdb-reset) command. It empties the DBOS tables and leaves the schema migrated, so nothing has to provision the database again between runs, and it works the same way whatever language your application is written in. Pass `--app` to reset just one application's state in a [shared system database](./explanations/sharing-a-system-database.md), or `--drop-database` to drop the database outright. Each SDK also ships a `dbos reset` command for its own language, which drops the database ([Python](./python/reference/cli.md#dbos-reset), [TypeScript](./typescript/reference/cli.md#npx-dbos-reset), [Go](./golang/reference/cli.md)). #### How can I reduce the number of Postgres connections DBOS uses? You can set the system database pool size in DBOS configuration ([Python](./python/reference/configuration.md), [TypeScript](./typescript/reference/configuration.md), [Go](./golang/reference/dbos-context.md), [Java](./java/reference/lifecycle.md)). Do not use values less than 5. #### Can I use DBOS with an external Postgres connection pooler? You can connect a DBOS application to its system database through a connection pooler like [PgBouncer](https://www.pgbouncer.org/), but **only in session mode**, not in transaction mode. See [this page](https://www.pgbouncer.org/features.html) for more information on the difference between session and transaction mode. #### Why do I get insufficient privilege errors when starting DBOS? DBOS creates tables for its internal state in its [system database](./explanations/system-tables.md). By default, a DBOS application automatically creates these on startup. However, in production environments, a DBOS application may not run with sufficient privilege to create databases or tables. In that case, the [`dbosctl sysdb migrate`](./conductor/reference/dbosctl.md#dbosctl-sysdb-migrate) command can be run with a privileged user to create all DBOS system tables. Then, a DBOS application can run without privilege (requiring only access to the system database). #### What database privileges does DBOS need, and how do I grant them manually? DBOS uses two distinct sets of privileges: the elevated privileges needed to **create or migrate** its schema, and the runtime privileges your **application** needs to operate on it. Separating them lets you run migrations with a privileged role and then run your application with a minimally-privileged role. **Migrating the schema.** Running the migration command (or letting a DBOS application create its tables on startup) creates the DBOS schema (`dbos` by default) and its tables, indexes, sequences, functions, and triggers in the system database. The role that runs the migration must be able to: - Create the schema if it does not already exist (the `CREATE` privilege on the database). - Create tables, indexes, sequences, functions, and triggers within that schema. - Create extensions. **Running your application.** Once the schema exists, your application does not need any DDL privileges, only read/write access to the DBOS schema. If you want to script this grant independently (rather than using the migration command), grant the following to your application's role on the DBOS schema (`dbos` by default) in the system database: ```sql GRANT USAGE ON SCHEMA "dbos" TO "your_app_role"; GRANT ALL PRIVILEGES ON ALL TABLES IN SCHEMA "dbos" TO "your_app_role"; GRANT ALL PRIVILEGES ON ALL SEQUENCES IN SCHEMA "dbos" TO "your_app_role"; GRANT EXECUTE ON ALL FUNCTIONS IN SCHEMA "dbos" TO "your_app_role"; -- So the role also has access to objects created by future migrations run by this user: ALTER DEFAULT PRIVILEGES IN SCHEMA "dbos" GRANT ALL ON TABLES TO "your_app_role"; ALTER DEFAULT PRIVILEGES IN SCHEMA "dbos" GRANT ALL ON SEQUENCES TO "your_app_role"; ALTER DEFAULT PRIVILEGES IN SCHEMA "dbos" GRANT EXECUTE ON FUNCTIONS TO "your_app_role"; ``` The [`dbosctl sysdb migrate`](./conductor/reference/dbosctl.md#dbosctl-sysdb-migrate) command does this automatically if you supply an application role with `-r`/`--app-role`. #### How does DBOS scale? The [architecture page](./architecture.md) describes how to architect a distributed DBOS application and how DBOS scales. #### When should I use the DBOS Conductor? Conductor is required for correct workflow recovery in applications that use more than one process. Additionally, Conductor offers data retention, observability, cluster-wide metrics, alerts, and organization management with RBAC. These features are designed to help engineering teams maintain long-term app health and reliability. We recommend using Conductor whenever you are running in a production environment. #### Why is my application not connecting to Conductor? The most common reason an application fails to connect to Conductor is that the name the application is registered with in its DBOS configuration does not match the name it was registered with in Conductor. Additionally, if you are [self-hosting Conductor](./conductor/self-hosting/hosting-conductor.md) with a free license, you may connect at most one executor per application to Conductor, so additional executors may see their connections rejected. To connect multiple executors, upgrade to a paid license. #### Why is my Conductor dashboard flickering? The most common cause of flickering is that you have connected multiple executors using different system databases to the same Conductor application (for example, both an executor from your dev environment and one from your prod environment), causing Conductor to receive inconsistent data. For isolation, you should set up a separate Conductor app for each environment in which you run your DBOS application. For example, you may want to have separate dev, staging, and prod Conductor apps. See [the docs](./conductor/overview.md#managing-conductor-applications) for more information. #### How are "checkpoints" calculated in Conductor pricing? Every workflow counts as one checkpoint and every step counts as one additional checkpoint. You can monitor your current usage at https://console.dbos.dev/usage You can also run the following SQL query on your [System Database](./explanations/system-tables.md) to compute your daily checkpoint count: ```SQL WITH daily_workflows AS ( SELECT DATE_TRUNC('day', TO_TIMESTAMP(created_at / 1000)) AS day, workflow_uuid FROM dbos.workflow_status ) SELECT dw.day, COUNT(DISTINCT dw.workflow_uuid) AS workflow_count, COUNT(oo.workflow_uuid) AS step_count, COUNT(DISTINCT dw.workflow_uuid) + COUNT(oo.workflow_uuid) AS total_checkpoints FROM daily_workflows dw LEFT JOIN dbos.operation_outputs oo ON dw.workflow_uuid = oo.workflow_uuid GROUP BY dw.day ORDER BY dw.day DESC; ``` --- ## Fault-Tolerant Checkout :::info This example is also available in [TypeScript](../../typescript/examples/checkout-tutorial), [Java](../../java/examples/widget-store), and [Python](../../python/examples/widget-store.md). ::: In this example, we use DBOS and Gin to build an online storefront that's resilient to any failure. You can see the application live [here](https://demo-widget-store.cloud.dbos.dev/). Try playing with it and pressing the crash button as often as you want. Within a few seconds, the app will recover and resume as if nothing happened. All source code is [available on GitHub](https://github.com/dbos-inc/dbos-demo-apps/tree/main/golang/widget-store). ![Widget store UI](./assets/widget_store_ui.png) ### Building the Checkout Workflow The heart of this application is the checkout workflow, which orchestrates the entire purchase process. This workflow is triggered whenever a customer buys a widget and handles the complete order lifecycle: 1. Creates a new order in the system 2. Reserves inventory to ensure the item is available 3. Processes payment 4. Marks the order as paid and initiates fulfillment 5. Handles failures gracefully by releasing reserved inventory and canceling orders when necessary DBOS **durably executes** this workflow. It checkpoints each step in the database so that if the app fails or is interrupted during checkout, it will automatically recover from the last completed step. This means that customers never lose their order progress, no matter what breaks. You can try this yourself! On the [live application](https://demo-widget-store.cloud.dbos.dev/), start an order and press the crash button at any time. Within seconds, your app will recover to exactly the state it was in before the crash and continue as if nothing happened. ```go func checkoutWorkflow(ctx dbos.Context, _ string) (string, error) { workflowID, err := ctx.GetWorkflowID() if err != nil { logger.Error("workflow ID retrieval failed", "error", err) return "", err } // Create a new order orderID, err := dbos.RunAsStep(ctx, func(stepCtx context.Context) (int, error) { return createOrder(stepCtx) }) if err != nil { logger.Error("order creation failed", "error", err, "wf_id", workflowID) return "", err } // Attempt to reserve inventory, cancelling the order if no inventory remains success, err := dbos.RunAsStep(ctx, func(stepCtx context.Context) (bool, error) { return reserveInventory(stepCtx) }) if err != nil || !success { logger.Warn("no inventory", "order", orderID) dbos.RunAsStep(ctx, func(stepCtx context.Context) (string, error) { return updateOrderStatus(stepCtx, UpdateOrderStatusInput{OrderID: orderID, OrderStatus: CANCELLED}) }) err = dbos.SetEvent(ctx, PAYMENT_ID, "") return "", err } err = dbos.SetEvent(ctx, PAYMENT_ID, workflowID) if err != nil { logger.Error("payment event creation failed", "error", err, "order", orderID, "payment", workflowID) return "", err } payment_status, err := dbos.Recv[string](ctx, PAYMENT_STATUS, 60*time.Second) if err != nil || payment_status != "paid" { logger.Warn("payment failed", "order", orderID, "payment", workflowID, "status", payment_status) dbos.RunAsStep(ctx, func(stepCtx context.Context) (string, error) { return undoReserveInventory(stepCtx) }) dbos.RunAsStep(ctx, func(stepCtx context.Context) (string, error) { return updateOrderStatus(stepCtx, UpdateOrderStatusInput{OrderID: orderID, OrderStatus: CANCELLED}) }) } else { logger.Info("payment success", "order", orderID, "payment", workflowID) dbos.RunAsStep(ctx, func(stepCtx context.Context) (string, error) { return updateOrderStatus(stepCtx, UpdateOrderStatusInput{OrderID: orderID, OrderStatus: PAID}) }) fmt.Println("calling dispatchOrderWorkflow") dbos.RunWorkflow(ctx, dispatchOrderWorkflow, orderID) } err = dbos.SetEvent(ctx, ORDER_ID, strconv.Itoa(orderID)) if err != nil { logger.Error("order event creation failed", "error", err, "order", orderID) return "", err } return "", nil } ``` ### The Checkout and Payment Endpoints Now let's implement the HTTP endpoints that handle customer interactions with the checkout system. The checkout endpoint is triggered when a customer clicks the "Buy Now" button. It starts the checkout workflow in the background, then waits for the workflow to generate and send it a unique payment ID. It then returns the payment ID so the browser can redirect the user to the payments page. The endpoint accepts an [idempotency key](../tutorials/workflow-tutorial.md#workflow-ids-and-idempotency) so that even if the customer presses "buy now" multiple times, only one checkout workflow is started. ```go func checkoutEndpoint(c *gin.Context, dbosCtx dbos.Context, logger *slog.Logger) { idempotencyKey := c.Param("idempotency_key") // Start the checkout workflow with the idempotency key _, err := dbos.RunWorkflow(dbosCtx, checkoutWorkflow, "", dbos.WithWorkflowID(idempotencyKey)) if err != nil { logger.Error("checkout workflow start failed", "error", err, "key", idempotencyKey) c.JSON(http.StatusInternalServerError, gin.H{"error": "Checkout failed to start"}) return } payment_id, err := dbos.GetEvent[string](dbosCtx, idempotencyKey, PAYMENT_ID, 60*time.Second) if err != nil || payment_id == "" { logger.Error("payment ID retrieval failed", "key", idempotencyKey) c.JSON(http.StatusInternalServerError, gin.H{"error": "Checkout failed"}) return } c.String(http.StatusOK, payment_id) } ``` The payment endpoint handles the communication between the payment system and the checkout workflow. It uses the payment ID to signal the checkout workflow whether the payment succeeded or failed. It then retrieves the order ID from the checkout workflow so the browser can redirect the customer to the order status page. ```go func paymentEndpoint(c *gin.Context, dbosCtx dbos.Context, logger *slog.Logger) { paymentID := c.Param("payment_id") paymentStatus := c.Param("payment_status") err := dbos.Send(dbosCtx, paymentID, paymentStatus, PAYMENT_STATUS) if err != nil { logger.Error("payment notification failed", "error", err, "payment", paymentID, "status", paymentStatus) c.JSON(http.StatusInternalServerError, gin.H{"error": "Failed to process payment"}) return } orderID, err := dbos.GetEvent[string](dbosCtx, paymentID, ORDER_ID, 60*time.Second) if err != nil || orderID == "" { logger.Error("order ID retrieval failed", "payment", paymentID) c.JSON(http.StatusInternalServerError, gin.H{"error": "Payment failed to process"}) return } c.String(http.StatusOK, orderID) } ``` ### Database Operations Now, let's implement the checkout workflow's steps. Each step performs a database operation, like updating inventory or order status. These are implemented as regular Go functions that interact with the Postgres database.
Database Operations ```go // Database operations for inventory management func reserveInventory(ctx context.Context) (bool, error) { result, err := db.Exec(ctx, "UPDATE products SET inventory = inventory - 1 WHERE product_id = $1 AND inventory > 0", WIDGET_ID) if err != nil { return false, err } return result.RowsAffected() > 0, nil } func undoReserveInventory(ctx context.Context) (string, error) { _, err := db.Exec(ctx, "UPDATE products SET inventory = inventory + 1 WHERE product_id = $1", WIDGET_ID) return "", err } // Database operations for order management func createOrder(ctx context.Context) (int, error) { var orderID int err := db.QueryRow(ctx, "INSERT INTO orders (order_status) VALUES ($1) RETURNING order_id", int(PENDING)).Scan(&orderID) return orderID, err } func updateOrderStatus(ctx context.Context, input UpdateOrderStatusInput) (string, error) { _, err := db.Exec(ctx, "UPDATE orders SET order_status = $1 WHERE order_id = $2", int(input.OrderStatus), input.OrderID) return "", err } func updateOrderProgress(ctx context.Context, orderID int) (int, error) { var progressRemaining int err := db.QueryRow(ctx, "UPDATE orders SET progress_remaining = progress_remaining - 1 WHERE order_id = $1 RETURNING progress_remaining", orderID).Scan(&progressRemaining) if err != nil { return 0, err } if progressRemaining == 0 { _, err = updateOrderStatus(ctx, UpdateOrderStatusInput{OrderID: orderID, OrderStatus: DISPATCHED}) } return progressRemaining, err } // HTTP endpoints for accessing data func getProduct(c *gin.Context, db *pgxpool.Pool, logger *slog.Logger) { var product Product err := db.QueryRow(context.Background(), "SELECT product_id, product, description, inventory, price FROM products LIMIT 1"). Scan(&product.ProductID, &product.Product, &product.Description, &product.Inventory, &product.Price) if err != nil { logger.Error("product query failed", "error", err) c.JSON(http.StatusInternalServerError, gin.H{"error": "Failed to fetch product"}) return } c.JSON(http.StatusOK, product) } func getOrders(c *gin.Context, db *pgxpool.Pool, logger *slog.Logger) { rows, err := db.Query(context.Background(), "SELECT order_id, order_status, last_update_time, progress_remaining FROM orders") if err != nil { logger.Error("orders database query failed", "error", err) c.JSON(http.StatusInternalServerError, gin.H{"error": "Failed to fetch orders"}) return } defer rows.Close() orders := []Order{} for rows.Next() { var order Order err := rows.Scan(&order.OrderID, &order.OrderStatus, &order.LastUpdateTime, &order.ProgressRemaining) if err != nil { logger.Error("order data parsing failed", "error", err) c.JSON(http.StatusInternalServerError, gin.H{"error": "Failed to process orders"}) return } orders = append(orders, order) } c.JSON(http.StatusOK, orders) } func getOrder(c *gin.Context, db *pgxpool.Pool, logger *slog.Logger) { idStr := c.Param("id") id, err := strconv.Atoi(idStr) if err != nil { logger.Warn("invalid order ID", "error", err, "id", idStr) c.JSON(http.StatusBadRequest, gin.H{"error": "Invalid order ID"}) return } var order Order err = db.QueryRow(context.Background(), "SELECT order_id, order_status, last_update_time, progress_remaining FROM orders WHERE order_id = $1", id). Scan(&order.OrderID, &order.OrderStatus, &order.LastUpdateTime, &order.ProgressRemaining) if err != nil { if err == pgx.ErrNoRows { logger.Warn("order not found", "order", id) c.JSON(http.StatusNotFound, gin.H{"error": "Order not found"}) } else { logger.Error("order database query failed", "error", err, "order", id) c.JSON(http.StatusInternalServerError, gin.H{"error": "Failed to fetch order"}) } return } c.JSON(http.StatusOK, order) } func restock(c *gin.Context, db *pgxpool.Pool, logger *slog.Logger) { _, err := db.Exec(context.Background(), "UPDATE products SET inventory = 100") if err != nil { logger.Error("inventory update failed", "error", err) c.JSON(http.StatusInternalServerError, gin.H{"error": "Failed to restock inventory"}) return } c.JSON(http.StatusOK, gin.H{"message": "Restocked successfully"}) } // Crash the app--for demonstration purposes only :) func crashApplication(c *gin.Context, logger *slog.Logger) { logger.Warn("application crash requested") c.JSON(http.StatusOK, gin.H{"message": "Crashing application..."}) // Give time for response to be sent go func() { time.Sleep(100 * time.Millisecond) logger.Error("intentional crash for demo") os.Exit(1) }() } ```
### Launching and Serving the App Finally, here's the complete main function that initializes DBOS, sets up the database connection, registers workflows, and starts the Gin HTTP server: ```go func main() { logger = slog.New(slog.NewJSONHandler(os.Stdout, &slog.HandlerOptions{ Level: slog.LevelDebug, })) dbURL := os.Getenv("DBOS_SYSTEM_DATABASE_URL") if dbURL == "" { logger.Error("DBOS_SYSTEM_DATABASE_URL required") os.Exit(1) } dbosContext, err := dbos.NewContext(context.Background(), dbos.Config{ AppName: "widget-store", ApplicationVersion: "0.1.0", DatabaseURL: os.Getenv("DBOS_SYSTEM_DATABASE_URL"), Logger: logger, ConductorAPIKey: os.Getenv("DBOS_CONDUCTOR_API_KEY"), }) if err != nil { logger.Error("DBOS initialization failed", "error", err) os.Exit(1) } dbos.RegisterWorkflow(dbosContext, checkoutWorkflow) dbos.RegisterWorkflow(dbosContext, dispatchOrderWorkflow) err = dbos.Launch(dbosContext) if err != nil { logger.Error("DBOS service start failed", "error", err) os.Exit(1) } defer dbos.Shutdown(dbosContext, 10 * time.Second) db, err = pgxpool.New(context.Background(), dbURL) if err != nil { logger.Error("database connection failed", "error", err) os.Exit(1) } defer db.Close() r := gin.Default() // Serve HTML r.StaticFile("/", "./html/app.html") // HTTP endpoints r.GET("/product", func(c *gin.Context) { getProduct(c, db, logger) }) r.GET("/orders", func(c *gin.Context) { getOrders(c, db, logger) }) r.GET("/order/:id", func(c *gin.Context) { getOrder(c, db, logger) }) r.POST("/restock", func(c *gin.Context) { restock(c, db, logger) }) r.POST("/checkout/:idempotency_key", func(c *gin.Context) { checkoutEndpoint(c, dbosContext, logger) }) r.POST("/payment_webhook/:payment_id/:payment_status", func(c *gin.Context) { paymentEndpoint(c, dbosContext, logger) }) r.POST("/crash_application", func(c *gin.Context) { crashApplication(c, logger) }) if err := r.Run(":8080"); err != nil { logger.Error("HTTP server start failed", "error", err) os.Exit(1) } } ``` ### Try it Yourself! First, clone and enter the [dbos-demo-apps](https://github.com/dbos-inc/dbos-demo-apps) repository: ```shell git clone https://github.com/dbos-inc/dbos-demo-apps.git cd golang/widget-store ``` Then follow the instructions in the README to build and run the app! --- ## Add DBOS To Your App This guide shows you how to add the open-source [DBOS Transact](https://github.com/dbos-inc/dbos-transact-golang) library to your existing application to **durably execute** it and make it resilient to any failure. #### 1. Install DBOS `go get` DBOS into your application. ```shell go get github.com/dbos-inc/dbos-transact-golang ``` DBOS requires a Postgres database. If you don't already have Postgres, you can install the DBOS Go CLI with [go install](https://pkg.go.dev/cmd/go#hdr-Compile_and_install_packages_and_dependencies) and start Postgres in a Docker container with these commands: ```shell go install github.com/dbos-inc/dbos-transact-golang/cmd/dbos@latest dbos postgres start ``` Then set the `DBOS_SYSTEM_DATABASE_URL` environment variable to your connection string (later we'll pass that value into DBOS). For example: ```shell export DBOS_SYSTEM_DATABASE_URL=postgres://postgres:dbos@localhost:5432/dbos_starter_go ``` #### 2. Add the DBOS Initializer Add these lines of code to your program's main function. They initialize a DBOS context when your program starts. ```go func main() { dbosContext, err := dbos.NewContext(context.Background(), dbos.Config{ AppName: "dbos-starter", ApplicationVersion: "0.1.0", DatabaseURL: os.Getenv("DBOS_SYSTEM_DATABASE_URL"), }) if err != nil { panic(fmt.Sprintf("Initializing DBOS failed: %v", err)) } err = dbos.Launch(dbosContext) if err != nil { panic(fmt.Sprintf("Launching DBOS failed: %v", err)) } defer dbos.Shutdown(dbosContext, 5 * time.Second) } ``` #### 3. Start Your Application Try starting your application. If everything is set up correctly, your app should run normally and log `DBOS launched` on startup. Congratulations! You've integrated DBOS into your application. #### 4. Start Building With DBOS DBOS let's you execute your functions as [workflows](./tutorials/workflow-tutorial.md) of [steps](./tutorials/step-tutorial.md). Workflows must be registered with the DBOS context before it is launched, for example: ```go dbos.RegisterWorkflow(dbosContext, workflow) ``` A workflow function must have the following signature: ```go type Workflow[P any, R any] func(ctx Context, input P) (R, error) ``` And a step this signature: ```go type Step[R any] func(ctx context.Context) (R, error) ``` DBOS durably executes workflows so if they are ever interrupted, upon restart they automatically resume from the last completed step. You can add DBOS to your application incrementally—it won't interfere with code that's already there. It's totally okay for your application to have one DBOS workflow alongside thousands of lines of non-DBOS code. To learn more about programming with DBOS, check out [the guide](./programming-guide.md). ```go func workflow(ctx dbos.Context, _ string) (string, error) { _, err := dbos.RunAsStep(ctx, stepOne) if err != nil { return "failure", err } _, err = dbos.RunAsStep(ctx, stepTwo) if err != nil { return "failure", err } return "success", err } func stepOne(ctx context.Context) (string, error) { fmt.Println("Step one completed") return "success", nil } func stepTwo(ctx context.Context) (string, error) { fmt.Println("Step two completed") return "success", nil } ``` --- ## Learn DBOS Go This guide shows you how to use DBOS to build Go apps that are **resilient to any failure**. :::tip To teach your AI coding assistant to build with DBOS, try out [skills](./prompting.md) and [MCP](../integrations/mcp.md). ::: ### 1. Setting Up Your Environment In an empty directory, initialize a new Go project and install DBOS: ```shell go mod init dbos-starter go get github.com/dbos-inc/dbos-transact-golang/dbos ``` DBOS requires a Postgres database. If you don't already have Postgres, you can install the DBOS Go CLI with [go install](https://pkg.go.dev/cmd/go#hdr-Compile_and_install_packages_and_dependencies) and start Postgres in a Docker container with these commands: ```shell go install github.com/dbos-inc/dbos-transact-golang/cmd/dbos@latest dbos postgres start ``` Then set the `DBOS_SYSTEM_DATABASE_URL` environment variable to your connection string (later we'll pass that value into DBOS). For example: ```shell export DBOS_SYSTEM_DATABASE_URL=postgres://postgres:dbos@localhost:5432/dbos_starter_go ``` ### 2. Workflows and Steps DBOS helps you add reliability to Go programs. The key feature of DBOS is **workflow functions** comprised of **steps**. DBOS automatically provides durability by checkpointing the state of your workflows and steps to its system database. If your program crashes or is interrupted, DBOS uses this saved state to recover each of your workflows from its last completed step. Thus, DBOS makes your application **resilient to any failure**. Let's create a simple DBOS program that runs a workflow of two steps. Create a file named `main.go` and add the following code to it: ```go showLineNumbers title="main.go" package main import ( "context" "fmt" "os" "time" "github.com/dbos-inc/dbos-transact-golang/dbos" ) func workflow(ctx dbos.Context, _ string) (string, error) { _, err := dbos.RunAsStep(ctx, stepOne) if err != nil { return "failure", err } _, err = dbos.RunAsStep(ctx, stepTwo) if err != nil { return "failure", err } return "success", err } func stepOne(ctx context.Context) (string, error) { fmt.Println("Step one completed") return "success", nil } func stepTwo(ctx context.Context) (string, error) { fmt.Println("Step two completed") return "success", nil } func main() { dbosContext, err := dbos.NewContext(context.Background(), dbos.Config{ AppName: "dbos-starter", ApplicationVersion: "0.1.0", DatabaseURL: os.Getenv("DBOS_SYSTEM_DATABASE_URL"), }) if err != nil { panic(fmt.Sprintf("Initializing DBOS failed: %v", err)) } dbos.RegisterWorkflow(dbosContext, workflow) err = dbos.Launch(dbosContext) if err != nil { panic(fmt.Sprintf("Launching DBOS failed: %v", err)) } defer dbos.Shutdown(dbosContext, 5 * time.Second) handle, err := dbos.RunWorkflow(dbosContext, workflow, "") if err != nil { panic(fmt.Sprintf("Error in DBOS workflow: %v", err)) } result, err := handle.GetResult() if err != nil { panic(fmt.Sprintf("Error in DBOS workflow: %v", err)) } fmt.Println("Workflow result:", result) } ``` Now, install dependencies and run this code with: ```shell go mod tidy go run main.go ``` Your program should print output like: ``` Step one completed Step two completed Workflow result: success ``` To see durable execution in action, let's modify the app to serve a DBOS workflow from an HTTP endpoint using Gin. Replace the contents of `main.go` with: ```go showLineNumbers title="main.go" package main import ( "context" "fmt" "net/http" "os" "time" "github.com/dbos-inc/dbos-transact-golang/dbos" "github.com/gin-gonic/gin" ) func workflow(ctx dbos.Context, _ string) (string, error) { _, err := dbos.RunAsStep(ctx, stepOne) if err != nil { return "failure", err } for range 5 { fmt.Println("Press Control + C to stop the app...") dbos.Sleep(ctx, time.Second) } _, err = dbos.RunAsStep(ctx, stepTwo) if err != nil { return "failure", err } return "success", err } func stepOne(ctx context.Context) (string, error) { fmt.Println("Step one completed") return "success", nil } func stepTwo(ctx context.Context) (string, error) { fmt.Println("Step two completed") return "success", nil } func main() { dbosContext, err := dbos.NewContext(context.Background(), dbos.Config{ AppName: "dbos-starter", ApplicationVersion: "0.1.0", DatabaseURL: os.Getenv("DBOS_SYSTEM_DATABASE_URL"), }) if err != nil { panic(fmt.Sprintf("Initializing DBOS failed: %v", err)) } dbos.RegisterWorkflow(dbosContext, workflow) err = dbos.Launch(dbosContext) if err != nil { panic(fmt.Sprintf("Launching DBOS failed: %v", err)) } defer dbos.Shutdown(dbosContext, 5 * time.Second) r := gin.Default() r.GET("/", func(c *gin.Context) { _, err := dbos.RunWorkflow(dbosContext, workflow, "") if err != nil { c.JSON(http.StatusInternalServerError, gin.H{"error": fmt.Sprintf("Error in DBOS workflow: %v", err)}) return } c.Status(http.StatusOK) }) r.Run(":8080") } ``` Now, install dependencies and run this code with: ```shell go mod tidy go run main.go ``` Then, visit this URL: http://localhost:8080. In your terminal, you should see an output like: ``` [GIN-debug] Listening and serving HTTP on :8080 [GIN] 2025/08/19 - 14:31:56 | 200 | 6.08315ms | ::1 | GET "/" Step one completed Press Control + C to stop the app... Press Control + C to stop the app... ``` Now, press CTRL+C stop your app. Then, run `go run main.go` to restart it. You should see an output like: ``` [GIN-debug] Listening and serving HTTP on :8080 Press Control + C to stop the app... Press Control + C to stop the app... Press Control + C to stop the app... Press Control + C to stop the app... Press Control + C to stop the app... Step two completed ``` You can see how DBOS **recovers your workflow from the last completed step**, executing step two without re-executing step one. Learn more about workflows, steps, and their guarantees [here](./tutorials/workflow-tutorial.md). ### 3. Queues and Parallelism To run many functions concurrently, use DBOS _queues_. To try them out, copy this code into `main.go`: ```go showLineNumbers title="main.go" package main import ( "context" "fmt" "net/http" "os" "time" "github.com/dbos-inc/dbos-transact-golang/dbos" "github.com/gin-gonic/gin" ) func taskWorkflow(ctx dbos.Context, i int) (int, error) { dbos.Sleep(ctx, 5*time.Second) fmt.Printf("Task %d completed\n", i) return i, nil } func queueWorkflow(ctx dbos.Context, queueName string) (int, error) { fmt.Println("Enqueuing tasks") queue, err := dbos.RetrieveQueue(ctx, queueName) if err != nil { return 0, err } handles := make([]dbos.WorkflowHandle[int], 10) for i := range 10 { handle, err := dbos.RunWorkflow(ctx, taskWorkflow, i, dbos.WithQueue(queue)) if err != nil { return 0, err } handles[i] = handle } results := make([]int, 10) for i, handle := range handles { result, err := handle.GetResult() if err != nil { return 0, err } results[i] = result } fmt.Printf("Successfully completed %d tasks\n", len(results)) return len(results), nil } func main() { dbosContext, err := dbos.NewContext(context.Background(), dbos.Config{ AppName: "dbos-starter", ApplicationVersion: "0.1.0", DatabaseURL: os.Getenv("DBOS_SYSTEM_DATABASE_URL"), }) if err != nil { panic(fmt.Sprintf("Initializing DBOS failed: %v", err)) } dbos.RegisterWorkflow(dbosContext, queueWorkflow) dbos.RegisterWorkflow(dbosContext, taskWorkflow) err = dbos.Launch(dbosContext) if err != nil { panic(fmt.Sprintf("Launching DBOS failed: %v", err)) } defer dbos.Shutdown(dbosContext, 5 * time.Second) _, err = dbos.RegisterQueue(dbosContext, "queue") if err != nil { panic(fmt.Sprintf("Registering queue failed: %v", err)) } r := gin.Default() r.GET("/", func(c *gin.Context) { _, err := dbos.RunWorkflow(dbosContext, queueWorkflow, "queue") if err != nil { c.JSON(http.StatusInternalServerError, gin.H{"error": fmt.Sprintf("Error in DBOS workflow: %v", err)}) return } c.Status(http.StatusOK) }) r.Run(":8080") } ``` When you enqueue a function by passing `dbos.WithQueue(queue)` into `dbos.RunWorkflow` (where `queue` is the handle returned by `dbos.RegisterQueue` or `dbos.RetrieveQueue`), DBOS executes it _asynchronously_, running it in the background without waiting for it to finish. `dbos.RunWorkflow` returns a handle representing the state of the enqueued function. This example enqueues ten functions, then waits for them all to finish using `.GetResult()` to wait for each of their handles. Now, restart your app with: ```shell go run main.go ``` Then, visit this URL: http://localhost:8080. Wait five seconds and you should see an output like: ``` [GIN-debug] Listening and serving HTTP on :8080 [GIN] 2025/08/19 - 14:42:14 | 200 | 6.961186ms | ::1 | GET "/" Enqueuing tasks Task 0 completed Task 2 completed Task 1 completed Task 4 completed Task 3 completed Task 5 completed Task 6 completed Task 7 completed Task 8 completed Task 9 completed Successfully completed 10 tasks ``` You can see how all ten steps run concurrently—even though each takes five seconds, they all finish at the same time. Learn more about DBOS queues [here](./tutorials/queue-tutorial.md). ### 4. Connecting to DBOS Conductor [Conductor](../conductor/overview.md) is the control plane for your durable workflows, providing distributed workflow recovery, observability, and management. Once you connect your app to Conductor, you can view and manage all its workflows and queued tasks from the [DBOS Console](https://console.dbos.dev). To connect your app to Conductor, first sign up for an account on the [DBOS Console](https://console.dbos.dev/login-redirect). Then, install [`dbosctl`](../conductor/reference/dbosctl.md), the Conductor command-line client. On Windows, [download a release binary](https://github.com/dbos-inc/dbos-ctl/releases) instead. ```shell curl -sSfL https://raw.githubusercontent.com/dbos-inc/dbos-ctl/main/install.sh | sh ``` Next, configure a `dbosctl` profile and log in. `dbosctl login` prints a URL and a code for you to approve in your browser. ```shell dbosctl config set dbos --managed dbosctl login ``` Then, register your application with Conductor and create an API key. The name you register must match the `AppName` in your DBOS configuration. The key's secret is printed once and cannot be retrieved afterwards, so copy it now. ```shell dbosctl app register dbos-starter dbosctl api-key create dbos-starter-key ``` Next, supply your API key to your app through the `ConductorAPIKey` configuration option. Update the call to `dbos.NewContext` in `main` to read the key from an environment variable: ```go dbosContext, err := dbos.NewContext(context.Background(), dbos.Config{ AppName: "dbos-starter", ApplicationVersion: "0.1.0", DatabaseURL: os.Getenv("DBOS_SYSTEM_DATABASE_URL"), ConductorAPIKey: os.Getenv("DBOS_CONDUCTOR_KEY"), }) ``` Finally, set the `DBOS_CONDUCTOR_KEY` environment variable to the key you created and restart your app: ```shell export DBOS_CONDUCTOR_KEY= go run main.go ``` Your app is now connected to Conductor! Launch a workflow by visiting http://localhost:8080, then watch it execute in real time from the [DBOS Console](https://console.dbos.dev). Learn more about Conductor [here](../conductor/overview.md). Congratulations! You've finished the DBOS Go guide. Next, you should: - Learn how to [**add DBOS to your own application**](./integrating-dbos.md). - Check out some [**example applications**](../examples/index.md). --- ## AI Model Prompting You may want assistance from an AI model in building a DBOS application. To make sure your model has the latest information on how to use DBOS, provide it with this prompt. You may also want to use the [DBOS MCP server](../integrations/mcp.md) so your model can directly access your application's workflows and steps. ### How To Use First, use the click-to-copy button in the top right of the code block to copy the full prompt to your clipboard. Then, paste into your AI tool of choice (for example OpenAI's ChatGPT or Anthropic's Claude). This adds the prompt to your AI model's context, giving it up-to-date instructions on how to build an application with DBOS. If you are using an AI-powered IDE, you can add this prompt to your project's context. For example: - Claude Code: Add the prompt, or a link to it, to your CLAUDE.md file. - Cursor: Add the prompt to [your project rules](https://docs.cursor.com/context/rules-for-ai). - Zed: Copy the prompt to a file in your project, then use the [`/file`](https://zed.dev/docs/assistant/commands?highlight=%2Ffile#file) command to add the file to your context. - GitHub Copilot: Create a [`.github/copilot-instructions.md`](https://docs.github.com/en/copilot/customizing-copilot/adding-repository-custom-instructions-for-github-copilot) file in your repository and add the prompt to it. ### Prompt ````markdown # Build Reliable Applications With DBOS ## Guidelines - Respond in a friendly and concise manner - Ask clarifying questions when requirements are ambiguous - Generate code in Golang using the DBOS library. - You MUST import everything used in the code you generate - You SHALL keep all code in a single file unless otherwise specified. - DBOS does NOT stand for anything. ## Workflow Guidelines Workflows provide durable execution so you can write programs that are resilient to any failure. Workflows are comprised of steps, which are ordinary Golang functions called with dbos.RunAsStep. When using DBOS workflows, you should call any function that performs complex operations or accesses external APIs or services as a step using dbos.RunAsStep. If a workflow is interrupted for any reason (e.g., an executor restarts or crashes), when your program restarts the workflow automatically resumes execution from the last completed step. - If asked to add DBOS to existing code, you MUST ask which function to make a workflow. Do NOT recommend any changes until they have told you what function to make a workflow. Do NOT make a function a workflow unless SPECIFICALLY requested. - When making a function a workflow, you should make all functions it calls steps. Do NOT change the functions in any way. - Do NOT make functions steps unless they are DIRECTLY called by a workflow. - If the workflow function performs a non-deterministic action, you MUST move that action to its own function and make that function a step. Examples of non-deterministic actions include accessing an external API or service, accessing files on disk, generating a random number, or getting the current time. - Do NOT start goroutines from workflows or use select in workflows. Instead, use DBOS's durable `dbos.Go` and `dbos.Select` functions which provide deterministic replay. For more complex parallel execution, use DBOS.RunWorkflow and DBOS queues. - Do NOT range over a map to call steps or start workflows: Go map iteration order is random, which breaks determinism. Sort the keys first (e.g. `slices.Sorted(maps.Keys(m))`) and iterate over the sorted slice. - DBOS workflows and steps should NOT have side effects in memory outside of their own scope. They can access global variables, but they should NOT create or update global variables or variables outside their scope. - Do NOT call DBOS context methods (DBOS.Send, DBOS.Recv, DBOS.RunWorkflow, DBOS.RunAsTransaction, DBOS.Enqueue, DBOS.Go, DBOS.Sleep, DBOS.GetEvent, DBOS.CloseStream, handle.GetResult, or workflow/schedule management writes) from a step — they return an error. Calling one step function from another is fine (it runs inline as part of the enclosing step), and DBOS.SetEvent, DBOS.WriteStream, and read/list operations are allowed from steps. ## DBOS Lifecycle Guidelines DBOS should be installed and imported from the `github.com/dbos-inc/dbos-transact-golang/dbos` package. DBOS programs MUST have a main file (typically 'main.go') that creates all objects and workflow functions during startup. Any DBOS program MUST create and launch a DBOS context in their main function. All workflows must be registered BEFORE DBOS is launched. Queues may be registered at any time, including after launch. ```go func main() { dbosContext, err := dbos.NewContext(context.Background(), dbos.Config{ AppName: "dbos-starter", ApplicationVersion: "0.1.0", DatabaseURL: os.Getenv("DBOS_SYSTEM_DATABASE_URL"), }) if err != nil { panic(fmt.Sprintf("Initializing DBOS failed: %v", err)) } dbos.RegisterWorkflow(dbosContext, workflow) err = dbos.Launch(dbosContext) if err != nil { panic(fmt.Sprintf("Launching DBOS failed: %v", err)) } defer dbos.Shutdown(dbosContext, 5 * time.Second) } ``` Here is an example main function using Gin: ```go import ( "context" "fmt" "net/http" "os" "time" "github.com/dbos-inc/dbos-transact-golang/dbos" "github.com/gin-gonic/gin" ) func main() { dbosContext, err := dbos.NewContext(context.Background(), dbos.Config{ AppName: "dbos-starter", ApplicationVersion: "0.1.0", DatabaseURL: os.Getenv("DBOS_SYSTEM_DATABASE_URL"), }) if err != nil { panic(fmt.Sprintf("Initializing DBOS failed: %v", err)) } dbos.RegisterWorkflow(dbosContext, workflow) err = dbos.Launch(dbosContext) if err != nil { panic(fmt.Sprintf("Launching DBOS failed: %v", err)) } defer dbos.Shutdown(dbosContext, 5 * time.Second) r := gin.Default() r.GET("/", func(c *gin.Context) { handle, err := dbos.RunWorkflow(dbosContext, workflow, "") if err != nil { c.JSON(http.StatusInternalServerError, gin.H{"error": fmt.Sprintf("Error in DBOS workflow: %v", err)}) return } result, err := handle.GetResult() if err != nil { c.JSON(http.StatusInternalServerError, gin.H{"error": fmt.Sprintf("Error in DBOS workflow: %v", err)}) return } c.JSON(http.StatusOK, gin.H{"result": result}) }) r.Run(":8080") } ``` ### Workflow and Steps Examples Simple example: ```go showLineNumbers title="main.go" package main import ( "context" "fmt" "os" "time" "github.com/dbos-inc/dbos-transact-golang/dbos" ) func workflow(ctx dbos.Context, _ string) (string, error) { _, err := dbos.RunAsStep(ctx, stepOne) if err != nil { return "failure", err } _, err = dbos.RunAsStep(ctx, stepTwo) if err != nil { return "failure", err } return "success", err } func stepOne(ctx context.Context) (string, error) { fmt.Println("Step one completed") return "success", nil } func stepTwo(ctx context.Context) (string, error) { fmt.Println("Step two completed") return "success", nil } func main() { dbosContext, err := dbos.NewContext(context.Background(), dbos.Config{ AppName: "dbos-starter", ApplicationVersion: "0.1.0", DatabaseURL: os.Getenv("DBOS_SYSTEM_DATABASE_URL"), }) if err != nil { panic(fmt.Sprintf("Initializing DBOS failed: %v", err)) } dbos.RegisterWorkflow(dbosContext, workflow) err = dbos.Launch(dbosContext) if err != nil { panic(fmt.Sprintf("Launching DBOS failed: %v", err)) } defer dbos.Shutdown(dbosContext, 5 * time.Second) handle, err := dbos.RunWorkflow(dbosContext, workflow, "") if err != nil { panic(fmt.Sprintf("Error in DBOS workflow: %v", err)) } result, err := handle.GetResult() if err != nil { panic(fmt.Sprintf("Error in DBOS workflow: %v", err)) } fmt.Println("Workflow result:", result) } ``` Example with Gin: ```go showLineNumbers title="main.go" package main import ( "context" "fmt" "net/http" "os" "time" "github.com/dbos-inc/dbos-transact-golang/dbos" "github.com/gin-gonic/gin" ) func workflow(ctx dbos.Context, _ string) (string, error) { _, err := dbos.RunAsStep(ctx, stepOne) if err != nil { return "failure", err } for range 5 { fmt.Println("Press Control + C to stop the app...") dbos.Sleep(ctx, time.Second) } _, err = dbos.RunAsStep(ctx, stepTwo) if err != nil { return "failure", err } return "success", err } func stepOne(ctx context.Context) (string, error) { fmt.Println("Step one completed") return "success", nil } func stepTwo(ctx context.Context) (string, error) { fmt.Println("Step two completed") return "success", nil } func main() { dbosContext, err := dbos.NewContext(context.Background(), dbos.Config{ AppName: "dbos-starter", ApplicationVersion: "0.1.0", DatabaseURL: os.Getenv("DBOS_SYSTEM_DATABASE_URL"), }) if err != nil { panic(fmt.Sprintf("Initializing DBOS failed: %v", err)) } dbos.RegisterWorkflow(dbosContext, workflow) err = dbos.Launch(dbosContext) if err != nil { panic(fmt.Sprintf("Launching DBOS failed: %v", err)) } defer dbos.Shutdown(dbosContext, 5 * time.Second) r := gin.Default() r.GET("/", func(c *gin.Context) { handle, err := dbos.RunWorkflow(dbosContext, workflow, "") if err != nil { c.JSON(http.StatusInternalServerError, gin.H{"error": fmt.Sprintf("Error in DBOS workflow: %v", err)}) return } result, err := handle.GetResult() if err != nil { c.JSON(http.StatusInternalServerError, gin.H{"error": fmt.Sprintf("Error in DBOS workflow: %v", err)}) return } c.JSON(http.StatusOK, gin.H{"result": result}) }) r.Run(":8080") } ``` Example with queues: ```go showLineNumbers title="main.go" package main import ( "context" "fmt" "net/http" "os" "time" "github.com/dbos-inc/dbos-transact-golang/dbos" "github.com/gin-gonic/gin" ) func taskWorkflow(ctx dbos.Context, i int) (int, error) { dbos.Sleep(ctx, 5*time.Second) fmt.Printf("Task %d completed\n", i) return i, nil } func queueWorkflow(ctx dbos.Context, queueName string) (int, error) { fmt.Println("Enqueuing tasks") queue, err := dbos.RetrieveQueue(ctx, queueName) if err != nil { return 0, err } handles := make([]dbos.WorkflowHandle[int], 10) for i := range 10 { handle, err := dbos.RunWorkflow(ctx, taskWorkflow, i, dbos.WithQueue(queue)) if err != nil { return 0, err } handles[i] = handle } results := make([]int, 10) for i, handle := range handles { result, err := handle.GetResult() if err != nil { return 0, err } results[i] = result } fmt.Printf("Successfully completed %d tasks\n", len(results)) return len(results), nil } func main() { dbosContext, err := dbos.NewContext(context.Background(), dbos.Config{ AppName: "dbos-starter", ApplicationVersion: "0.1.0", DatabaseURL: os.Getenv("DBOS_SYSTEM_DATABASE_URL"), }) if err != nil { panic(fmt.Sprintf("Initializing DBOS failed: %v", err)) } dbos.RegisterWorkflow(dbosContext, queueWorkflow) dbos.RegisterWorkflow(dbosContext, taskWorkflow) err = dbos.Launch(dbosContext) if err != nil { panic(fmt.Sprintf("Launching DBOS failed: %v", err)) } defer dbos.Shutdown(dbosContext, 5 * time.Second) _, err = dbos.RegisterQueue(dbosContext, "queue") if err != nil { panic(fmt.Sprintf("Registering queue failed: %v", err)) } r := gin.Default() r.GET("/", func(c *gin.Context) { handle, err := dbos.RunWorkflow(dbosContext, queueWorkflow, "queue") if err != nil { c.JSON(http.StatusInternalServerError, gin.H{"error": fmt.Sprintf("Error in DBOS workflow: %v", err)}) return } result, err := handle.GetResult() if err != nil { c.JSON(http.StatusInternalServerError, gin.H{"error": fmt.Sprintf("Error in DBOS workflow: %v", err)}) return } c.JSON(http.StatusOK, gin.H{"result": result}) }) r.Run(":8080") } ``` ### Workflow Documentation Workflows provide **durable execution** so you can write programs that are **resilient to any failure**. Workflows are comprised of steps, which wrap ordinary Go functions. If a workflow is interrupted for any reason (e.g., an executor restarts or crashes), when your program restarts the workflow automatically resumes execution from the last completed step. To write a workflow, register a Go function with `RegisterWorkflow`. Workflow registration must happen before launching the DBOS context with `dbos.Launch()` The function's signature must match: ```go type Workflow[P any, R any] func(ctx Context, input P) (R, error) ``` In other words, a workflow must take in a DBOS context and one other input of any serializable (json-encodable) type and must return one output of any serializable type and error. For example: ```go func stepOne(ctx context.Context) (string, error) { fmt.Println("Step one completed") return "success", nil } func stepTwo(ctx context.Context) (string, error) { fmt.Println("Step two completed") return "success", nil } func workflow(ctx dbos.Context, _ string) (string, error) { _, err := dbos.RunAsStep(ctx, stepOne) if err != nil { return "failure", err } _, err = dbos.RunAsStep(ctx, stepTwo) if err != nil { return "failure", err } return "success", err } func main() { ... // Create the DBOS context dbos.RegisterWorkflow(dbosContext, workflow) ... // Launch DBOS after registering all workflows } ``` Call workflows with `RunWorkflow`. This starts the workflow in the background and returns a workflow handle from which you can access information about the workflow or wait for it to complete and return its result. Here's an example: ```go func runWorkflowExample(dbosContext dbos.Context, input string) error { handle, err := dbos.RunWorkflow(dbosContext, workflow, input) if err != nil { return err } result, err := handle.GetResult() if err != nil { return err } fmt.Println("Workflow result:", result) return nil } ``` #### Workflow IDs and Idempotency Every time you execute a workflow, that execution is assigned a unique ID, by default a UUID. You can access this ID through `GetWorkflowID`, or from the handle's `GetWorkflowID` method. Workflow IDs are useful for communicating with workflows and developing interactive workflows. You can set the workflow ID of a workflow using `WithWorkflowID` when calling `RunWorkflow`. Workflow IDs must be **globally unique** for your application. An assigned workflow ID acts as an idempotency key: if a workflow is called multiple times with the same ID, it executes only once. This is useful if your operations have side effects like making a payment or sending an email. For example: ```go func exampleWorkflow(ctx dbos.Context, input string) (string, error) { workflowID, err := dbos.GetWorkflowID(ctx) if err != nil { return "", err } fmt.Printf("Running workflow with ID: %s\n", workflowID) // ... return "success", nil } func example(dbosContext dbos.Context, input string) error { myID := "unique-workflow-id-123" handle, err := dbos.RunWorkflow(dbosContext, exampleWorkflow, input, dbos.WithWorkflowID(myID)) if err != nil { log.Fatal(err) } result, err := handle.GetResult() if err != nil { log.Fatal(err) } fmt.Println("Result:", result) return nil } ``` #### Determinism Workflows are in most respects normal Go functions. They can have loops, branches, conditionals, and so on. However, a workflow function must be **deterministic**: if called multiple times with the same inputs, it should invoke the same steps with the same inputs in the same order (given the same return values from those steps). If you need to perform a non-deterministic operation like accessing the database, calling a third-party API, generating a random number, or getting the local time, you shouldn't do it directly in a workflow function. Instead, you should do all non-deterministic operations in steps. :::warning Go's goroutine scheduler and `select` operation are non-deterministic. You should use them only inside steps, or use the durable `dbos.Go` and `dbos.Select` functions instead. Go's map iteration order is also random. Don't call steps or start workflows while ranging over a map: sort the keys first (e.g. `slices.Sorted(maps.Keys(m))`) and iterate over the sorted slice. ::: For example, **don't do this**: ```go func exampleWorkflow(ctx dbos.Context, input string) (string, error) { randomChoice := rand.Intn(2) if randomChoice == 0 { return dbos.RunAsStep(ctx, stepOne) } else { return dbos.RunAsStep(ctx, stepTwo) } } ``` Instead, do this: ```go func generateChoice(ctx context.Context) (int, error) { return rand.Intn(2), nil } func exampleWorkflow(ctx dbos.Context, input string) (string, error) { randomChoice, err := dbos.RunAsStep(ctx, generateChoice) if err != nil { return "", err } if randomChoice == 0 { return dbos.RunAsStep(ctx, stepOne) } else { return dbos.RunAsStep(ctx, stepTwo) } } ``` #### Workflow Timeouts You can set a timeout for a workflow using its input `Context`. Use `WithTimeout` to obtain a cancellable `Context`, as you would with a normal `context.Context`. When the timeout expires, the workflow and all its children are cancelled. Cancelling a workflow sets its status to CANCELLED and preempts its execution at the beginning of its next step. You can detach a child workflow by passing it an uncancellable context, which you can obtain with `WithoutCancel`. Timeouts are **start-to-completion**: if a workflow is enqueued, the timeout does not begin until the workflow is dequeued and starts execution. Also, timeouts are durable: they are stored in the database and persist across restarts, so workflows can have very long timeouts. ```go func exampleWorkflow(ctx dbos.Context, input string) (string, error) {} timeoutCtx, cancelFunc := dbos.WithTimeout(dbosCtx, 12*time.Hour) handle, err := dbos.RunWorkflow(timeoutCtx, exampleWorkflow, "wait-for-cancel") ``` You can also manually cancel the workflow by calling its `cancel` function (or calling CancelWorkflow). #### Durable Sleep You can use `Sleep` to put your workflow to sleep for any period of time. This sleep is **durable**—DBOS saves the wakeup time in the database so that even if the workflow is interrupted and restarted multiple times while sleeping, it still wakes up on schedule. Sleeping is useful for scheduling a workflow to run in the future (even days, weeks, or months from now). For example: ```go func exampleWorkflow(ctx dbos.Context, input struct { TimeToSleep time.Duration Task string }) (string, error) { // Sleep for the specified duration _, err := dbos.Sleep(ctx, input.TimeToSleep) if err != nil { return "", err } // Execute the task after sleeping result, err := dbos.RunAsStep( ctx, func(stepCtx context.Context) (string, error) { return fmt.Sprintf("Completed: %s", input.Task), nil }, ) if err != nil { return "", err } return result, nil } ``` #### Concurrent Steps DBOS provides durable `Go` and `Select` functions to run multiple steps concurrently within a workflow while preserving durability guarantees. These are durable alternatives to Go's native goroutines and `select` statement. `Go` launches a step asynchronously and returns a channel that will receive the result when the step completes. `Select` waits for the first result from multiple concurrent steps. ```go func workflow(ctx dbos.Context, _ string) (string, error) { // Launch two concurrent steps ch1, err := dbos.Go(ctx, func(ctx context.Context) (string, error) { return queryServiceA(ctx) }) if err != nil { return "", err } ch2, err := dbos.Go(ctx, func(ctx context.Context) (string, error) { return queryServiceB(ctx) }) if err != nil { return "", err } // Wait for the first result result, err := dbos.Select(ctx, []<-chan dbos.StepOutcome[string]{ch1, ch2}) if err != nil { return "", err } return result, nil } ``` #### Scheduled Workflows You can schedule workflows to run on a cron expression. Schedules are stored in the database and can be created, paused, resumed, and deleted at runtime. Scheduled workflows are useful for running recurring tasks like data backups, report generation, or cleanup operations. Scheduled workflows must accept a `dbos.ScheduledWorkflowInput`, which carries the cron tick time and a user-defined `Context` value attached to the schedule: ```go type ScheduledWorkflowInput struct { ScheduledTime time.Time `json:"scheduled_time"` Context json.RawMessage `json:"context,omitempty"` } ``` The `Context` field holds the raw JSON of the value set on the schedule; decode it inside the workflow with `dbos.DecodeScheduleContext[T](input)`: ```go func DecodeScheduleContext[T any](input ScheduledWorkflowInput) (T, error) ``` Register the workflow normally, then create a schedule for it using `dbos.CreateSchedule` (or `dbos.ApplySchedules` to declaratively apply a set of schedules on start): ```go func dailyBackup(ctx dbos.Context, input dbos.ScheduledWorkflowInput) (any, error) { fmt.Printf("Running daily backup at: %s\n", input.ScheduledTime.Format(time.RFC3339)) ... // Perform daily backup operations return nil, nil } func main() { dbosContext := ... // Initialize DBOS dbos.RegisterWorkflow(dbosContext, dailyBackup) err := dbos.Launch(dbosContext) if err != nil { log.Fatal(err) } // Schedule the workflow to run daily at 2:00 AM err = dbos.CreateSchedule(dbosContext, dbos.ScheduleSpec{ ScheduleName: "daily-backup", Workflow: dailyBackup, Schedule: "0 0 2 * * *", }) if err != nil { log.Fatal(err) } } ``` Schedules can also be paused, resumed, deleted, backfilled, and triggered at runtime with `dbos.PauseSchedule`, `dbos.ResumeSchedule`, `dbos.DeleteSchedule`, `dbos.BackfillSchedule`, and `dbos.TriggerSchedule`. By default, scheduled invocations are enqueued on an internal queue; set the `QueueName` field of `dbos.ScheduleSpec` to route them to a declared queue for concurrency or rate-limit control. #### Workflow Versioning and Recovery Because DBOS recovers workflows by re-executing them using information saved in the database, a workflow cannot safely be recovered if its code has changed since the workflow was started. To guard against this, DBOS _versions_ applications and their workflows. When DBOS is launched, it computes an application version from a hash of the application source code (this can be overridden through configuration). All workflows are tagged with the application version on which they started. When DBOS tries to recover workflows, it only recovers workflows whose version matches the current application version. This prevents unsafe recovery of workflows that depend on different code. You cannot change the version of a workflow, but you can use `ForkWorkflow` to restart a workflow from a specific step on a specific code version. For more information on managing workflow recovery when self-hosting production DBOS applications, check out the guide. #### Workflow Attempts and Recovery The `Attempts` field in `WorkflowStatus` tracks how many times a workflow has been executed. - On first execution, `Attempts` is set to `1`. - If the workflow is enqueued but not yet dequeued, `Attempts` is `0`. - Each time the workflow is recovered (e.g., after a crash) or dequeued for execution, `Attempts` is incremented by `1`. You can limit the number of attempts using `WithMaxRecoveryAttempts` when registering a workflow. If `WithMaxRecoveryAttempts(n)` is set, the workflow may be attempted at most `n + 1` times (one initial execution plus `n` retries). If this limit is exceeded, the workflow's status is set to `MAX_RECOVERY_ATTEMPTS_EXCEEDED` and it will no longer be recovered automatically. You can use `ResumeWorkflow` to manually resume a workflow that has exceeded its maximum attempts after fixing the underlying issue. ```go // Register a workflow that can be attempted at most 4 times (1 initial + 3 retries) dbos.RegisterWorkflow(dbosContext, myWorkflow, dbos.WithMaxRecoveryAttempts(3)) ``` ### Steps When using DBOS workflows, you should call any function that performs complex operations or accesses external APIs or services as a _step_. If a workflow is interrupted, upon restart it automatically resumes execution from the **last completed step**. You can use `RunAsStep` to call a function as a step. For a function to be used as a step, it should return a serializable (json-encodable) value and an error and have this signature: ```go type Step[R any] func(ctx context.Context) (R, error) ``` Here's a simple example: ```go func generateRandomNumber(ctx context.Context) (int, error) { return rand.Int(), nil } func workflowFunction(ctx dbos.Context, n int) (int, error) { randomNumber, err := dbos.RunAsStep( ctx, generateRandomNumber, dbos.WithStepName("generateRandomNumber"), ) if err != nil { return 0, err } return randomNumber, nil } ``` You can pass arguments into a step by wrapping it in an anonymous function, like this: ```go func generateRandomNumber(ctx context.Context, n int) (int, error) { return rand.IntN(n), nil } func workflowFunction(ctx dbos.Context, n int) (int, error) { randomNumber, err := dbos.RunAsStep( ctx, func(stepCtx context.Context) (int, error) { return generateRandomNumber(stepCtx, n) }, dbos.WithStepName("generateRandomNumber"), ) if err != nil { return 0, err } return randomNumber, nil } ``` You should make a function a step if you're using it in a DBOS workflow and it performs a **nondeterministic** operation. A nondeterministic operation is one that may return different outputs given the same inputs. Common nondeterministic operations include: - Accessing an external API or service, like serving a file from AWS S3, calling an external API like Stripe, or accessing an external data store like Elasticsearch. - Accessing files on disk. - Generating a random number. - Getting the current time. You **cannot** call, start, or enqueue workflows from within steps. You also cannot call DBOS methods like `Send` or `Recv` from within steps. These operations should be performed from workflow functions. You can call one step from another step, but the called step becomes part of the calling step's execution rather than functioning as a separate step. #### Configurable Retries You can optionally configure a step to automatically retry any error a set number of times with exponential backoff. This is useful for automatically handling transient failures, like making requests to unreliable APIs. Retries are configurable through step options that can be passed to `RunAsStep`. Available retry configuration options include: - `WithStepName` - Custom name for the step (default to the Go runtime reflection value) - `WithStepMaxRetries` - Maximum number of times this step is automatically retried on failure (default 0) - `WithStepMaxInterval` - Maximum delay between retries (default 5s) - `WithStepBackoffFactor` - Exponential backoff multiplier between retries (default 2.0) - `WithStepBaseInterval` - Initial delay between retries (default 100ms) For example, let's configure this step to retry failures (such as if the site to be fetched is temporarily down) up to 10 times: ```go func fetchStep(ctx context.Context, url string) (string, error) { resp, err := http.Get(url) if err != nil { return "", err } defer resp.Body.Close() body, err := io.ReadAll(resp.Body) if err != nil { return "", err } return string(body), nil } func fetchWorkflow(ctx dbos.Context, inputURL string) (string, error) { return dbos.RunAsStep( ctx, func(stepCtx context.Context) (string, error) { return fetchStep(stepCtx, inputURL) }, dbos.WithStepName("fetchFunction"), dbos.WithStepMaxRetries(10), dbos.WithStepMaxInterval(30*time.Second), dbos.WithStepBackoffFactor(2.0), dbos.WithStepBaseInterval(500*time.Millisecond), ) } ``` If a step exhausts all retry attempts, it returns an error to the calling workflow. ### Transactions & Datasources A datasource is a handle to a database you own, over which DBOS can run durable transactions. A transaction run through a datasource inside a workflow commits your application writes and the DBOS durability record atomically, so it executes exactly once even across crashes and recovery (stronger than a step, which is at-least-once). Create a datasource with `NewDataSource`, passing a `*pgxpool.Pool` (Postgres/CockroachDB) or `*sql.DB` (SQLite). It may be called at any time, before or after `Launch()`, and provisions a `transaction_completion` durability table in your database unless it already exists. If the engine is the same handle as the DBOS system database (passed via `Config.SystemDBPool` or `Config.SQLiteSystemDB`), no `transaction_completion` table is created or managed: application writes and the DBOS checkpoint commit together in a single transaction (detection is by pointer identity, not connection string). ```go func NewDataSource[E Engine](ctx Context, engine E, opts ...DataSourceOption) (*DataSource, error) ``` Options: `WithDataSourceName(name string)` (logs only, default "datasource"), `WithDataSourceSchema(schema string)` (schema for the durability table, default "dbos"). Run a transaction inside a workflow with `RunAsTransaction`. The function receives a driver-agnostic `Tx` with `Exec`, `Query`, and `QueryRow` methods; DBOS commits if the function returns successfully and rolls back on error (do not call `Commit`/`Rollback` yourself). ```go func RunAsTransaction[R any](ctx Context, ds *DataSource, fn Txn[R], opts ...StepOption) (R, error) pool, _ := pgxpool.New(context.Background(), os.Getenv("APP_DATABASE_URL")) ds, err := dbos.NewDataSource(dbosContext, pool, dbos.WithDataSourceName("app")) // Inside a workflow: n, err := dbos.RunAsTransaction(ctx, ds, func(txCtx context.Context, tx dbos.Tx) (int64, error) { res, err := tx.Exec(txCtx, "INSERT INTO orders(item) VALUES ($1)", item) if err != nil { return 0, err } return res.RowsAffected() }, dbos.WithStepMaxRetries(3)) ``` Rules: - `RunAsTransaction` must be called from within a workflow; it shares the workflow's step counter with `RunAsStep` and accepts the same step options. - Serialization/deadlock conflicts are retried automatically with a fresh transaction; application errors follow the step retry policy. - Nesting a `RunAsTransaction` inside another `RunAsTransaction` or inside a `RunAsStep` is rejected with an error — run transactions from workflow code. ### Workflow Communication DBOS provides a few different ways to communicate with your workflows. You can: - Send messages to workflows - Publish events from workflows for clients to read - Stream values from workflows to clients #### Workflow Messaging and Notifications You can send messages to a specific workflow. This is useful for signaling a workflow or sending notifications to it while it's running. #### Send ```go func Send[P any](ctx Client, destinationID string, message P, topic string, opts ...SendOption) error ``` You can call `Send()` to send a message to a workflow. Messages can optionally be associated with a topic and are queued on the receiver per topic. Pass `WithIdempotencyKey(key string)` to make a retried `Send` deliver at most once. To send many messages at once, possibly to different workflows, call `SendBulk(ctx, []dbos.SendMessage{...})`. The batch is sent in a single transaction: either every message is delivered or none is. Each `SendMessage` has `DestinationID`, `Message`, `Topic`, and an optional per-message `IdempotencyKey`. #### Recv ```go func Recv[R any](ctx Context, topic string, timeout time.Duration) (R, error) ``` Workflows can call `Recv()` to receive messages sent to them, optionally for a particular topic. Each call to `Recv()` waits for and consumes the next message to arrive in the queue for the specified topic, returning an error if the wait times out. If the topic is not specified, this method only receives messages sent without a topic. #### Messages Example Messages are especially useful for sending notifications to a workflow. For example, in an e-commerce application, the checkout workflow, after redirecting customers to a secure payments service, must wait for a notification from that service that the payment has finished processing. To wait for this notification, the payments workflow uses `Recv()`, executing failure-handling code if the notification doesn't arrive in time: ```go const PaymentStatusTopic = "payment_status" func checkoutWorkflow(ctx dbos.Context, orderData OrderData) (string, error) { // Process initial checkout steps... // Wait for payment notification with a 5-minute timeout notification, err := dbos.Recv[PaymentNotification](ctx, PaymentStatusTopic, 5*time.Minute) if err != nil { ... // Handle timeout or other errors } // Handle the notification if notification.Status == "completed" { ... // Handle the notification. } else { ... // Handle a failure } } ``` A webhook waits for the payment processor to send the notification, then uses `Send()` to forward it to the workflow: ```go func paymentWebhookHandler(w http.ResponseWriter, r *http.Request) { // Parse the notification from the payment processor notification := ... // Retrieve the workflow ID from notification metadata workflowID := ... // Send the notification to the waiting workflow err := dbos.Send(dbosContext, workflowID, notification, PaymentStatusTopic) if err != nil { http.Error(w, "Failed to send notification", http.StatusInternalServerError) return } } ``` #### Reliability Guarantees All messages are persisted to the database, so if `Send` completes successfully, the destination workflow is guaranteed to be able to `Recv` it. If you're sending a message from a workflow, DBOS guarantees exactly-once delivery. #### Workflow Events Workflows can publish _events_, which are key-value pairs associated with the workflow. They are useful for publishing information about the status of a workflow or to send a result to clients while the workflow is running. #### SetEvent ```go func SetEvent[P any](ctx Context, key string, message P) error ``` Any workflow can call `SetEvent` to publish a key-value pair, or update its value if it has already been published. #### GetEvent ```go func GetEvent[R any](ctx Client, targetWorkflowID, key string, timeout time.Duration) (R, error) ``` You can call `GetEvent` to retrieve the value published by a particular workflow ID for a particular key. If the event does not yet exist, this call waits for it to be published, returning an error if the wait times out. #### Events Example Events are especially useful for writing interactive workflows that communicate information to their caller. For example, in an e-commerce application, the checkout workflow, after validating an order, directs the customer to a secure payments service to handle credit card processing. To communicate the payments URL to the customer, it uses events. The checkout workflow emits the payments URL using `SetEvent()`: ```go const PaymentURLKey = "payment_url" func checkoutWorkflow(ctx dbos.Context, orderData OrderData) (string, error) { // Process order validation... paymentsURL := ... err := dbos.SetEvent(ctx, PaymentURLKey, paymentsURL) if err != nil { return "", fmt.Errorf("failed to set payment URL event: %w", err) } // Continue with checkout process... } ``` The HTTP handler that originally started the workflow uses `GetEvent()` to await this URL, then redirects the customer to it: ```go func webCheckoutHandler(dbosContext dbos.Context, w http.ResponseWriter, r *http.Request) { orderData := parseOrderData(r) // Parse order from request handle, err := dbos.RunWorkflow(dbosContext, checkoutWorkflow, orderData) if err != nil { http.Error(w, "Failed to start checkout", http.StatusInternalServerError) return } // Wait up to 30 seconds for the payment URL event url, err := dbos.GetEvent[string](dbosContext, handle.GetWorkflowID(), PaymentURLKey, 30*time.Second) if err != nil { // Handle a timeout } // Redirect the customer } ``` #### Reliability Guarantees All events are persisted to the database, so the latest version of an event is always retrievable. Additionally, if `GetEvent` is called in a workflow, the retrieved value is persisted in the database so workflow recovery can use that value, even if the event is later updated. #### Workflow Streaming Workflows can stream data in real time to clients. This is useful for streaming results from a long-running workflow or LLM call, or for monitoring and progress reporting. ##### Writing to Streams ```go func WriteStream[P any](ctx Context, key string, value P) error ``` You can write values to a stream from a workflow or its steps. A workflow may have any number of streams, each identified by a unique key. When you are done writing to a stream, you should close it with `CloseStream`. Otherwise, streams are automatically closed when the workflow terminates. ```go func CloseStream(ctx Context, key string) error ``` DBOS streams are immutable and append-only. Writes to a stream from a workflow happen exactly-once. Writes to a stream from a step happen at-least-once; if a step fails and is retried, it may write to the stream multiple times. ##### Reading from Streams ```go func ReadStream[R any](ctx Client, workflowID string, key string, opts ...ReadStreamOption) ([]R, bool, error) ``` You can read values from a stream from anywhere. This function reads all values from a stream identified by a workflow ID and key. It blocks until the stream is closed or the workflow becomes inactive (status is not `PENDING` or `ENQUEUED`). It returns the values, whether the stream is closed, and any error. To read without blocking, pass `WithReadStreamSnapshot()`, which returns as soon as all currently-available values have been drained; pair it with `WithReadStreamFromOffset(offset int)` to poll a stream incrementally. You can also read from a stream asynchronously, which returns a channel: ```go func ReadStreamAsync[R any](ctx Client, workflowID string, key string) (<-chan StreamValue[R], error) ``` ```go type StreamValue[R any] struct { Value R // The stream value (zero value if error/closed) Err error // Error if one occurred (nil otherwise) Closed bool // Whether the stream is closed } ``` ##### Streaming Example ```go func producerWorkflow(ctx dbos.Context, _ string) (string, error) { err := dbos.WriteStream(ctx, "progress", "step 1 complete") if err != nil { return "", err } err = dbos.WriteStream(ctx, "progress", "step 2 complete") if err != nil { return "", err } err = dbos.CloseStream(ctx, "progress") if err != nil { return "", err } return "done", nil } // Blocking read values, closed, err := dbos.ReadStream[string](ctx, workflowID, "progress") // Async read: process values as they arrive ch, err := dbos.ReadStreamAsync[string](ctx, workflowID, "progress") if err != nil { return err } for streamValue := range ch { if streamValue.Err != nil { return streamValue.Err } if streamValue.Closed { break } fmt.Printf("Received: %s\n", streamValue.Value) } ``` You can also read from a stream from outside a DBOS application with a DBOS Client using `ReadStream` or `ReadStreamAsync`. ### Cross-Language Interoperability DBOS supports multiple languages (Python, TypeScript, Go, Java). A client in one language can interact with workflows in another language using the **portable JSON** serialization format. #### Portable Workflows Use `WithPortableWorkflow()` when calling `RunWorkflow` to serialize inputs, outputs, and errors in portable JSON: ```go handle, err := dbos.RunWorkflow(dbosContext, processOrder, "order-123", dbos.WithPortableWorkflow(), ) ``` #### Portable Communication Use portable options on `Send`, `SetEvent`, and `WriteStream` for cross-language messaging: ```go // Send a message readable by any language dbos.Send(ctx, "workflow-123", map[string]any{"status": "complete"}, "updates", dbos.WithPortableSend(), ) // Set an event readable by any language dbos.SetEvent(ctx, "progress", map[string]any{"percent": 75}, dbos.WithPortableSetEvent(), ) // Write to a stream readable by any language dbos.WriteStream(ctx, "results", map[string]any{"item": "processed"}, dbos.WithPortableWriteStream(), ) ``` #### Enqueueing Cross-Language Workflows To enqueue a workflow on an application written in another language from a Go client, pass `PortableWorkflowArgs` as the input (this automatically uses portable JSON): ```go args := dbos.PortableWorkflowArgs{ PositionalArgs: []any{"order-123", 42}, } handle, err := dbos.Enqueue[any]( client, "orders", "process_order", args, dbos.WithEnqueueClassName("OrderProcessor"), // Required for Python/TS/Java targets ) ``` #### Portable Errors Return a `PortableWorkflowError` from a portable workflow to pass structured error info cross-language: ```go return nil, &dbos.PortableWorkflowError{ Name: "ValidationError", Message: "invalid input", Code: 400, Data: map[string]any{"field": "email"}, } ``` ### Alerting If you are using DBOS Conductor, you can register an alert handler to receive alerts when failure conditions are met. The handler must be registered before calling `Launch()`. Only one handler is allowed per application. ```go dbos.SetAlertHandler(dbosContext, func(ruleType string, message string, metadata map[string]string) { slog.Warn(fmt.Sprintf("Alert received: %s - %s", ruleType, message)) for key, value := range metadata { slog.Warn(fmt.Sprintf(" %s: %s", key, value)) } }) ``` The handler receives: - **ruleType**: One of `WorkflowFailure`, `SlowQueue`, or `UnresponsiveApplication`. - **message**: The alert message. - **metadata**: Key-value string pairs with additional alert context. ### Queues You can use queues to run many workflows at once with managed concurrency. Queues provide _flow control_, letting you manage how many workflows run at once or how often workflows are started. To create a queue, register it with `RegisterQueue`: ```go queue, err := dbos.RegisterQueue(dbosContext, "example_queue") ``` `RegisterQueue` persists the queue's configuration to the system database. It can be called at any time, including after `Launch()`, and the queue's configuration can be changed at runtime. `RegisterQueue` returns a `Queue` handle. Keep it somewhere your code can reach it (for example, a package-level variable)—enqueueing a workflow requires the handle, not the queue name. If you only have the name, fetch the handle with `RetrieveQueue`. You can then enqueue any workflow by passing the handle to `WithQueue` when calling `RunWorkflow`. Enqueuing a function submits it for execution and returns a handle to it. Queued tasks are started in first-in, first-out (FIFO) order. ```go func processTask(ctx dbos.Context, task string) (string, error) { // Process the task... return fmt.Sprintf("Processed: %s", task), nil } func example(dbosContext dbos.Context, queue dbos.Queue) error { // Enqueue a workflow task := "example_task" handle, err := dbos.RunWorkflow(dbosContext, processTask, task, dbos.WithQueue(queue)) if err != nil { return err } // Get the result result, err := handle.GetResult() if err != nil { return err } fmt.Println("Task result:", result) return nil } ``` #### Queue Example Here's an example of a workflow using a queue to process tasks concurrently: ```go func taskWorkflow(ctx dbos.Context, task string) (string, error) { // Process the task... return fmt.Sprintf("Processed: %s", task), nil } func queueWorkflow(ctx dbos.Context, queueName string) ([]string, error) { // Look up the queue handle by name queue, err := dbos.RetrieveQueue(ctx, queueName) if err != nil { return nil, err } // Enqueue each task so all tasks are processed concurrently tasks := []string{"task1", "task2", "task3", "task4", "task5"} var handles []dbos.WorkflowHandle[string] for _, task := range tasks { handle, err := dbos.RunWorkflow(ctx, taskWorkflow, task, dbos.WithQueue(queue)) if err != nil { return nil, fmt.Errorf("failed to enqueue task %s: %w", task, err) } handles = append(handles, handle) } // Wait for each task to complete and retrieve its result var results []string for i, handle := range handles { result, err := handle.GetResult() if err != nil { return nil, fmt.Errorf("task %d failed: %w", i, err) } results = append(results, result) } return results, nil } func example(dbosContext dbos.Context) error { handle, err := dbos.RunWorkflow(dbosContext, queueWorkflow, "example_queue") if err != nil { return err } results, err := handle.GetResult() if err != nil { return err } for _, result := range results { fmt.Println(result) } return nil } ``` #### Enqueueing from Another Application Often, you want to enqueue a workflow from outside your DBOS application. For example, let's say you have an API server and a data processing service. You're using DBOS to build a durable data pipeline in the data processing service. When the API server receives a request, it should enqueue the data pipeline for execution on the data processing service. You can use the DBOS Client to enqueue workflows from outside your DBOS application by connecting directly to your DBOS application's system database. Since the DBOS Client is designed to be used from outside your DBOS application, workflow and queue metadata must be specified explicitly. For example, this code enqueues the `dataPipeline` workflow on the `pipelineQueue` queue with a `ProcessInput` argument: ```go type ProcessInput struct { TaskID string Data string } type ProcessOutput struct { Result string Status string } config := dbos.ClientConfig{ DatabaseURL: os.Getenv("DBOS_SYSTEM_DATABASE_URL"), } client, err := dbos.NewClient(context.Background(), config) if err != nil { log.Fatal(err) } defer dbos.Shutdown(client, 5*time.Second) handle, err := dbos.Enqueue[ProcessOutput]( client, "pipelineQueue", "dataPipeline", ProcessInput{TaskID: "task-123", Data: "data"}, ) if err != nil { log.Fatal(err) } ``` #### Managing Concurrency You can control how many workflows from a queue run simultaneously by configuring concurrency limits. This helps prevent resource exhaustion when workflows consume significant memory or processing power. ##### Worker Concurrency Worker concurrency sets the maximum number of workflows from a queue that can run concurrently on a single DBOS process. This is particularly useful for resource-intensive workflows to avoid exhausting the resources of any process. For example, this queue has a worker concurrency of 5, so each process will run at most 5 workflows from this queue simultaneously: ```go queue, err := dbos.RegisterQueue(dbosContext, "example_queue", dbos.WithWorkerConcurrency(5)) ``` ##### Global Concurrency Global concurrency limits the total number of workflows from a queue that can run concurrently across all DBOS processes in your application. For example, this queue will have a maximum of 10 workflows running simultaneously across your entire application. :::warning Worker concurrency limits are recommended for most use cases. Take care when using a global concurrency limit as any `PENDING` workflow on the queue counts toward the limit, including workflows from previous application versions ::: ```go queue, err := dbos.RegisterQueue(dbosContext, "example_queue", dbos.WithGlobalConcurrency(10)) ``` #### Rate Limiting You can set _rate limits_ for a queue, limiting the number of functions that it can start in a given period. Rate limits are global across all DBOS processes using this queue. For example, this queue has a limit of 100 workflows with a period of 60 seconds, so it may not start more than 100 workflows in 60 seconds: ```go queue, err := dbos.RegisterQueue(dbosContext, "example_queue", dbos.WithRateLimiter(&dbos.RateLimiter{ Limit: 100, Period: 60 * time.Second, // 60 seconds })) ``` Rate limits are especially useful when working with a rate-limited API, such as many LLM APIs. #### Reconfiguring Queues Because queue configuration is persisted to the system database, you can change a queue's configuration at runtime without redeploying or restarting your workers. Workers pick up the new configuration on their next polling iteration. Use `RetrieveQueue` to fetch a queue, then call its `Set*` methods: ```go queue, err := dbos.RetrieveQueue(dbosContext, "example_queue") if err != nil { return err } concurrency := 50 err = queue.SetGlobalConcurrency(dbosContext, &concurrency) ``` If your application calls `RegisterQueue` on startup, the next process to start can overwrite settings you applied at runtime via `Set*` methods. Either update the `RegisterQueue` call to match the new configuration, or pass `WithQueueOnConflict(dbos.QueueConflictNeverUpdate)` to preserve the runtime changes. #### Deduplication You can set a deduplication ID for an enqueued workflow using `WithDeduplicationID` when calling `RunWorkflow`. At any given time, only one workflow with a specific deduplication ID can be enqueued in the specified queue. If a workflow with a deduplication ID is currently enqueued or actively executing (status `ENQUEUED` or `PENDING`), subsequent workflow enqueue attempts with the same deduplication ID in the same queue will return an error. Alternatively, use `WithDeduplicationPolicy(dbos.DeduplicationPolicyReturnExisting)` to instead return a handle to the existing workflow holding the deduplication ID. For example, this is useful if you only want to have one workflow active at a time per user—set the deduplication ID to the user's ID. **Example syntax:** ```go func taskWorkflow(ctx dbos.Context, task string) (string, error) { // Process the task... return "completed", nil } func example(dbosContext dbos.Context, queue dbos.Queue) error { task := "example_task" deduplicationID := "user_12345" // Use user ID for deduplication handle, err := dbos.RunWorkflow( dbosContext, taskWorkflow, task, dbos.WithQueue(queue), dbos.WithDeduplicationID(deduplicationID)) if err != nil { // Handle deduplication error or other failures return fmt.Errorf("failed to enqueue workflow: %w", err) } result, err := handle.GetResult() if err != nil { return fmt.Errorf("workflow failed: %w", err) } fmt.Printf("Workflow completed: %s\n", result) return nil } ``` #### Priority You can set a priority for an enqueued workflow using `WithPriority` when calling `RunWorkflow`. Workflows with the same priority are dequeued in **FIFO (first in, first out)** order. Priority values can range from `1` to `2,147,483,647`, where **a low number indicates a higher priority**. If using priority, you must set `WithPriorityEnabled` on your queue. :::tip Workflows without assigned priorities have the highest priority and are dequeued before workflows with assigned priorities. ::: To use priorities in a queue, you must enable it when creating the queue: ```go queue, err := dbos.RegisterQueue(dbosContext, "example_queue", dbos.WithPriorityEnabled()) ``` **Example syntax:** ```go func taskWorkflow(ctx dbos.Context, task string) (string, error) { // Process the task... return "completed", nil } func example(dbosContext dbos.Context, queue dbos.Queue) error { task := "example_task" priority := uint(10) // Lower number = higher priority handle, err := dbos.RunWorkflow(dbosContext, taskWorkflow, task, dbos.WithQueue(queue), dbos.WithPriority(priority)) if err != nil { return err } result, err := handle.GetResult() if err != nil { return fmt.Errorf("workflow failed: %w", err) } fmt.Printf("Workflow completed: %s\n", result) return nil } ``` #### Partitioned Queues You can partition queues to distribute work across dynamically created queue partitions. A queue is partitioned if you register it with any per-partition limit: `WithPartitionConcurrency` (maximum workflows from any one partition running at once across all processes), `WithPartitionWorkerConcurrency` (the same, on a single process), or `WithPartitionRateLimiter` (maximum workflows started from any one partition in a given period). When you enqueue a workflow on a partitioned queue, you must supply a queue partition key; a workflow enqueued without one is never dequeued. A partitioned queue enforces its partition limits and its queue-wide limits (`WithGlobalConcurrency`, `WithWorkerConcurrency`, `WithRateLimiter`) at the same time. Each per-partition concurrency limit must be less than or equal to its queue-wide counterpart. For example, to allow each user to run at most one task at a time, while running at most 10 tasks on any single process: ```go partitionedQueue, err := dbos.RegisterQueue(dbosContext, "user-tasks", dbos.WithPartitionConcurrency(1), dbos.WithWorkerConcurrency(10), ) // Enqueue workflows with partition keys // At most one task per user runs at once, but tasks from different users run concurrently handle, err := dbos.RunWorkflow(dbosContext, processTask, taskData, dbos.WithQueue(partitionedQueue), dbos.WithQueuePartitionKey(userID), ) ``` `WithPartitionQueue()` is deprecated: under it the queue-wide limits apply per partition instead. Use the partition limits. #### Delayed Execution You can delay an enqueued workflow's execution using `WithDelay`. The workflow is initially placed in `DELAYED` status and does not execute. After the delay expires, it transitions to `ENQUEUED` status and may be dequeued and executed. ```go remindersQueue, err := dbos.RegisterQueue(dbosContext, "reminders") if err != nil { return err } // Send a reminder in one hour handle, err := dbos.RunWorkflow(dbosContext, sendReminder, userID, dbos.WithQueue(remindersQueue), dbos.WithDelay(1 * time.Hour), ) ``` When enqueueing from a Client, use `WithEnqueueDelay` instead. You can dynamically update or shorten the delay of a `DELAYED` workflow using `SetWorkflowDelay`: ```go // Shorten the delay to 10 seconds from now err := dbos.SetWorkflowDelay(ctx, handle.GetWorkflowID(), dbos.WithDelayDuration(10*time.Second)) // Or set an absolute deadline err = dbos.SetWorkflowDelay(ctx, handle.GetWorkflowID(), dbos.WithDelayUntil(time.Now().Add(time.Minute))) ``` #### Debouncing Debouncing delays a workflow's execution until some time has passed since it was last called. This is useful when rapid successive triggers should be coalesced into a single workflow execution. ```go func processInput(ctx dbos.Context, input string) (string, error) { fmt.Printf("Processing input: %s\n", input) return "processed", nil } func main() { dbosContext, _ := dbos.NewContext(context.Background(), dbos.Config{ AppName: "debounce-example", ApplicationVersion: "0.1.0", DatabaseURL: os.Getenv("DBOS_SYSTEM_DATABASE_URL"), }) dbos.RegisterWorkflow(dbosContext, processInput) // Create a debouncer with a maximum timeout of 30 seconds debouncer, err := dbos.NewDebouncer(dbosContext, processInput, dbos.WithDebouncerTimeout(30*time.Second)) if err != nil { log.Fatal(err) } dbos.Launch(dbosContext) defer dbos.Shutdown(dbosContext, 5*time.Second) // Each call to Debounce pushes back the workflow start time by the delay. // The workflow runs with the most recent input once the delay expires. handle, err := debouncer.Debounce(dbosContext, "user-123", 5*time.Second, "first input") if err != nil { log.Fatal(err) } // If this call arrives within 5 seconds, the delay resets and the input updates handle, err = debouncer.Debounce(dbosContext, "user-123", 5*time.Second, "updated input") if err != nil { log.Fatal(err) } result, err := handle.GetResult() fmt.Println("Result:", result) // Processed with "updated input" } ``` To debounce a workflow method of a configured instance (registered with `WithInstance`), pass the instance with `WithDebouncerInstance`: ```go debouncer, err := dbos.NewDebouncer(ctx, slack.Send, dbos.WithDebouncerInstance(slack)) ``` Debouncers can be created at any time, including after `Launch()`. #### ListenQueues You can configure which queues the current DBOS process should listen to for workflow execution using `ListenQueues`. By default, all registered queues are listened to. This allows multiple DBOS processes to share the same queues but listen to different subsets. ```go dbos.RegisterQueue(ctx, "queue-1") dbos.RegisterQueue(ctx, "queue-2") dbos.RegisterQueue(ctx, "queue-3") // This process only listens to queue-1 and queue-2. dbos.ListenQueues(ctx, "queue-1", "queue-2") dbos.Launch(ctx) ``` Queues are identified by name; each call to `ListenQueues` replaces the whole listen set (an empty set listens to every queue), and the set may be changed at any time, including after `Launch()`. `ListenQueues` only controls what workflows are dequeued, not what workflows can be enqueued. ## Reference DBOS has two interfaces: `Client` and `Context`, where `Context` embeds (extends) `Client`. A `Client` connects to your application's system database and can enqueue and manage workflows, queues, and schedules from any process, including from outside a DBOS application (create one with `NewClient`). A `Context` is at the center of a DBOS-enabled application: it is everything a `Client` is, plus workflow registration and durable execution (create one with `NewContext`). `Context` extends Go's `context.Context` interface and carries essential state across workflow execution. Workflows and steps receive a new `Context` spun out of the root `Context` you manage. In addition, a `Context` can be used to set workflow timeouts. DBOS operations are package-level functions whose first parameter tells you who can call them: - A function taking a **`Client`** (e.g. `Enqueue`, `Send`, `ListWorkflows`, `RegisterQueue`, `CreateSchedule`) accepts a standalone client or any `Context`, because every `Context` **is** a `Client`. - A function taking a **`Context`** (e.g. `RunWorkflow`, `RunAsStep`, `Recv`, `SetEvent`, `Sleep`) requires a DBOS context; workflow-scoped functions must receive the `Context` passed into the workflow function. ### Lifecycle #### Initialization You can create a DBOS context using `NewContext`, which takes a `Config` object where `AppName` and one of `DatabaseURL`, `SystemDBPool`, or `SQLiteSystemDB` are mandatory. ```go func NewContext(ctx context.Context, inputConfig Config) (Context, error) ``` ```go type Config struct { AppName string // Application name for identification (required) DatabaseURL string // Connection string to your system database. May be a PostgreSQL (postgres://...) or SQLite (sqlite:...) URL. Exactly one of DatabaseURL, SystemDBPool, or SQLiteSystemDB is required. SystemDBPool *pgxpool.Pool // A custom Postgres/CockroachDB connection pool for your system database. Optional; takes precedence over DatabaseURL. Mutually exclusive with SQLiteSystemDB. SQLiteSystemDB *sql.DB // A custom SQLite handle (e.g. from modernc.org/sqlite) to use as your system database. Optional; takes precedence over DatabaseURL. Mutually exclusive with SystemDBPool. DatabaseSchema string // Database schema name (defaults to "dbos"; Postgres only) Logger *slog.Logger // Custom logger instance (defaults to a new slog logger) ConductorURL string // DBOS conductor service URL (optional) ConductorAPIKey string // DBOS conductor API key (optional) ConductorExecutorMetadata map[string]any // Metadata used to identify this executor on the Conductor dashboard (optional, must be JSON-serializable) ApplicationVersion string // Application version (optional) ExecutorID string // Executor ID (optional) EnablePatching bool // Enable the patching system for Patch/DeprecatePatch (default: false) Serializer Serializer[any] // Custom serializer for workflow inputs, outputs, and events (defaults to a JSON serializer) SchedulerPollingInterval time.Duration // How often database-backed schedules are reconciled (default: 30s) SystemDBStartupTimeout time.Duration // Maximum time for system database connection and migrations (default: 2 minutes) } ``` For example: ```go dbosContext, err := dbos.NewContext(context.Background(), dbos.Config{ AppName: "dbos-starter", ApplicationVersion: "0.1.0", DatabaseURL: os.Getenv("DBOS_SYSTEM_DATABASE_URL"), }) if err != nil { panic(err) } ``` The newly created Context must be launched with `Launch()` before use and should be shut down with Shutdown() at program termination. DBOS can back its system database with either Postgres (recommended for production; pass a `postgres://` `DatabaseURL` or a `*pgxpool.Pool` as `SystemDBPool`) or SQLite (useful for local development, testing, and single-node deployments; pass a `sqlite:` `DatabaseURL` or a `*sql.DB` as `SQLiteSystemDB`). SQLite support is not linked into your binary by default: to use SQLite, register the driver with one blank import anywhere in your binary: `import _ "github.com/dbos-inc/dbos-transact-golang/dbos/driver/sqlite"`. Without it, `NewContext` (or `NewClient`) fails at startup with an error naming this import. The import registers `modernc.org/sqlite`, a pure-Go driver requiring no cgo. SQLite `DatabaseURL` examples: `"sqlite:dbos.db"` (relative file) or `"sqlite:/var/lib/dbos.db"` (absolute file). `DatabaseSchema` applies to Postgres only. #### launch ```go dbos.Launch(ctx Context) error ``` Launch the following resources managed by a `Context`: - A system database connection pool - A workflow scheduler - A workflow queue runner - (Optionally) a Conductor connection In addition, `Launch()` may perform workflow recovery. `Launch()` should be called by your program during startup before running any workflows. #### Shutdown ```go func Shutdown(c Client, timeout time.Duration) error ``` Gracefully shutdown the DBOS runtime, waiting for workflows to complete and cleaning up resources. Accepts either a `Context` or a standalone `Client`. When you shutdown a `Context`, the underlying `context.Context` will be cancelled, which signals all DBOS resources they should stop executing, including workflows and steps. Shutting down a standalone client releases its system database connection pool. **Parameters:** - **timeout**: The time to wait for DBOS resources to gracefully terminate. ### Context management #### WithTimeout ```go func WithTimeout(ctx Context, timeout time.Duration) (Context, context.CancelFunc) ``` `WithTimeout` returns a copy of the DBOS context with a timeout. The returned context will be canceled after the specified duration. See workflow timeouts for usage. #### WithoutCancel ```go func WithoutCancel(ctx Context) Context ``` `WithoutCancel` returns a copy of the DBOS context that is not canceled when the parent context is canceled. This is useful to detach child workflows from their parent's timeout. #### WithCancel ```go func WithCancel(ctx Context) (Context, context.CancelFunc) ``` `WithCancel` returns a copy of the DBOS context that can be manually canceled, along with a `CancelFunc`. Cancelling propagates to workflows and steps running under the returned context. Call the returned `CancelFunc` when the derived context is no longer needed to release its resources. `WithCancelCause` is a variant that returns a `context.CancelCauseFunc`, letting you supply an error describing why the context was canceled (retrievable with `context.Cause`). ### Context metadata #### GetApplicationVersion ```go func GetApplicationVersion(ctx Context) string ``` `GetApplicationVersion` returns the application version for this context. #### GetExecutorID ```go func GetExecutorID(ctx Context) string ``` `GetExecutorID` returns the executor ID for this context. ### Workflow Communication #### GetEvent ```go func GetEvent[R any](ctx Client, targetWorkflowID, key string, timeout time.Duration) (R, error) ``` Retrieve the latest value of an event published by the workflow identified by `targetWorkflowID` to the key `key`. If the event does not yet exist, wait for it to be published, returning an error if the wait times out. **Parameters:** - **ctx**: The DBOS client or context. - **targetWorkflowID**: The identifier of the workflow whose events to retrieve. - **key**: The key of the event to retrieve. - **timeout**: A timeout. If the wait times out, return an error. #### SetEvent ```go func SetEvent[P any](ctx Context, key string, message P, opts ...SetEventOption) error ``` Create and associate with this workflow an event with key `key` and value `value`. If the event already exists, update its value. Can only be called from within a workflow. Use `WithPortableSetEvent()` for cross-language event consumption. **Parameters:** - **ctx**: The DBOS context. - **key**: The key of the event. - **message**: The value of the event. Must be serializable. - **opts**: Optional `SetEventOption` functions (e.g., `WithPortableSetEvent()`). #### Send ```go func Send[P any](ctx Client, destinationID string, message P, topic string, opts ...SendOption) error ``` Send a message to the workflow identified by `destinationID`. Messages can optionally be associated with a topic. Use `WithPortableSend()` for cross-language messaging. **Parameters:** - **ctx**: The DBOS client or context. - **destinationID**: The workflow to which to send the message. - **message**: The message to send. Must be serializable. - **topic**: A topic with which to associate the message. Messages are enqueued per-topic on the receiver. - **opts**: Optional `SendOption` functions (e.g., `WithPortableSend()`, `WithIdempotencyKey(key string)`). Pass `WithIdempotencyKey(key string)` to make a `Send` deliver at most once: the key is combined with the destination workflow ID to form the message's primary key, so retrying a `Send` with the same key inserts the message only once. #### SendBulk ```go func SendBulk(ctx Client, messages []SendMessage, opts ...SendOption) error type SendMessage struct { DestinationID string // The workflow to which to send the message Message any // The message to send. Must be serializable. Topic string // Optional topic IdempotencyKey string // Optional; delivers the message at most once per destination } ``` Send many messages in a single transaction; each message carries its own destination. The batch is atomic: if any destination does not exist, no message is sent. Inside a workflow the whole batch is one durable step. At most `dbos.MaxSendBulkMessages` (10,000) messages per call, and two messages in one call may not share an idempotency key. `WithSendTransaction` and `WithPortableSend` apply to the whole batch; `WithIdempotencyKey` is rejected, set `SendMessage.IdempotencyKey` instead. #### Recv ```go func Recv[R any](ctx Context, topic string, timeout time.Duration) (R, error) ``` Receive and return a message sent to this workflow. Can only be called from within a workflow. Messages are dequeued first-in, first-out from a queue associated with the topic. Calls to `Recv` wait for the next message in the queue, returning an error if the wait times out. **Parameters:** - **ctx**: The DBOS context. - **topic**: A topic queue on which to wait. - **timeout**: A timeout duration. If the wait times out, return an error. ### Streams #### WriteStream ```go func WriteStream[P any](ctx Context, key string, value P, opts ...WriteStreamOption) error ``` Write a value to a durable stream. May only be called from within a workflow or step. Writes from a workflow are exactly-once; writes from a step are at-least-once. Use `WithPortableWriteStream()` for cross-language stream reading. #### CloseStream ```go func CloseStream(ctx Context, key string) error ``` Close a durable stream. After closing, no more values can be written to the stream. Streams are also automatically closed when the workflow terminates. #### ReadStream ```go func ReadStream[R any](ctx Client, workflowID string, key string, opts ...ReadStreamOption) ([]R, bool, error) ``` Read all values from a durable stream. Blocks until the stream is closed or the workflow becomes inactive. Pass `WithReadStreamSnapshot()` to instead return immediately once all currently-available values have been drained; pair it with `WithReadStreamFromOffset(offset int)` to poll incrementally. #### ReadStreamAsync ```go func ReadStreamAsync[R any](ctx Client, workflowID string, key string) (<-chan StreamValue[R], error) ``` Read values from a durable stream asynchronously. Returns immediately with a channel that receives values as they are written to the stream. #### Sleep ```go func Sleep(ctx Context, duration time.Duration) (time.Duration, error) ``` Sleep for the given duration. May only be called from within a workflow. This sleep is durable—it records its intended wake-up time in the database so if it is interrupted and recovers, it still wakes up at the intended time. **Parameters:** - **ctx**: The DBOS context. - **duration**: The duration to sleep. #### RetrieveWorkflow ```go func RetrieveWorkflow[R any](ctx Client, workflowID string) (WorkflowHandle[R], error) ``` Retrieve the handle of a workflow. **Parameters**: - **ctx**: The DBOS client or context. - **workflowID**: The ID of the workflow whose handle to retrieve. ### Workflow Management Methods #### ListWorkflows ```go func ListWorkflows(ctx Client, opts ...ListWorkflowsOption) ([]WorkflowStatus, error) ``` Retrieve a list of `WorkflowStatus` of all workflows matching specified criteria. **Example usage:** ```go // List all successful workflows from the last 24 hours workflows, err := dbos.ListWorkflows(ctx, dbos.WithFilterStatus(dbos.WorkflowStatusSuccess), dbos.WithFilterCreatedAfter(time.Now().Add(-24*time.Hour)), dbos.WithFilterLimit(100)) if err != nil { log.Fatal(err) } // List workflows by specific IDs without loading input/output data workflows, err := dbos.ListWorkflows(ctx, dbos.WithFilterWorkflowIDs("workflow1", "workflow2"), dbos.WithFilterLoadInput(false), dbos.WithFilterLoadOutput(false)) if err != nil { log.Fatal(err) } ``` ##### WithFilterAppVersion ```go func WithFilterAppVersion(appVersion ...string) ListWorkflowsOption ``` Retrieve workflows tagged with this application version. ##### WithFilterCreatedBefore ```go func WithFilterCreatedBefore(endTime time.Time) ListWorkflowsOption ``` Retrieve workflows started before this timestamp. ##### WithFilterLimit ```go func WithFilterLimit(limit int) ListWorkflowsOption ``` Retrieve up to this many workflows. ##### WithFilterLoadInput ```go func WithFilterLoadInput(loadInput bool) ListWorkflowsOption ``` WithFilterLoadInput controls whether to load workflow input data (default: true). ##### WithFilterLoadOutput ```go func WithFilterLoadOutput(loadOutput bool) ListWorkflowsOption ``` WithFilterLoadOutput controls whether to load workflow output data (default: true). ##### WithFilterName ```go func WithFilterName(names ...string) ListWorkflowsOption ``` Filter workflows by the specified workflow function name. ##### WithFilterOffset ```go func WithFilterOffset(offset int) ListWorkflowsOption ``` Skip this many workflows from the results returned (for pagination). ##### WithFilterSortDesc ```go func WithFilterSortDesc() ListWorkflowsOption ``` Sort the results in descending order by workflow start time (ascending is the default). ##### WithFilterCreatedAfter ```go func WithFilterCreatedAfter(startTime time.Time) ListWorkflowsOption ``` Retrieve workflows started after this timestamp. ##### WithFilterStatus ```go func WithFilterStatus(status ...WorkflowStatusType) ListWorkflowsOption ``` Filter workflows by status. Multiple statuses can be specified. ##### WithFilterUser ```go func WithFilterUser(user ...string) ListWorkflowsOption ``` Filter workflows run by this authenticated user. ##### WithFilterWorkflowIDs ```go func WithFilterWorkflowIDs(workflowIDs ...string) ListWorkflowsOption ``` Filter workflows by specific workflow IDs. ##### WithFilterWorkflowIDPrefix ```go func WithFilterWorkflowIDPrefix(prefix ...string) ListWorkflowsOption ``` Filter workflows whose IDs start with the specified prefix. ##### WithFilterQueuesOnly ```go func WithFilterQueuesOnly() ListWorkflowsOption ``` Return only workflows that are currently in a queue (queue name is not null, status is `ENQUEUED` or `PENDING`). ##### WithFilterCompletedAfter ```go func WithFilterCompletedAfter(completedAfter time.Time) ListWorkflowsOption ``` Retrieve workflows that reached a terminal state (`SUCCESS`, `ERROR`, or `CANCELLED`) at or after this timestamp. ##### WithFilterCompletedBefore ```go func WithFilterCompletedBefore(completedBefore time.Time) ListWorkflowsOption ``` Retrieve workflows that reached a terminal state (`SUCCESS`, `ERROR`, or `CANCELLED`) at or before this timestamp. ##### WithFilterDequeuedAfter ```go func WithFilterDequeuedAfter(dequeuedAfter time.Time) ListWorkflowsOption ``` Retrieve workflows that started executing at or after this timestamp. ##### WithFilterDequeuedBefore ```go func WithFilterDequeuedBefore(dequeuedBefore time.Time) ListWorkflowsOption ``` Retrieve workflows that started executing at or before this timestamp. ##### WithFilterWasForkedFrom ```go func WithFilterWasForkedFrom(wasForkedFrom bool) ListWorkflowsOption ``` Filter workflows by whether they have been forked from (true) or not (false). ##### WithFilterHasParent ```go func WithFilterHasParent(hasParent bool) ListWorkflowsOption ``` Filter workflows by whether they have a parent workflow (true) or not (false). #### GetWorkflowSteps ```go func GetWorkflowSteps(ctx Client, workflowID string, opts ...GetWorkflowStepsOption) ([]StepInfo, error) ``` GetWorkflowSteps retrieves the execution steps of a workflow. This is a list of `StepInfo` objects, with the following structure: ```go type StepInfo struct { StepID int // The sequential ID of the step within the workflow StepName string // The name of the step function Output any // The output returned by the step (if any) Error error // The error returned by the step (if any) ChildWorkflowID string // If the step starts or retrieves the result of a workflow, its ID } ``` **Parameters:** - **ctx**: The DBOS client or context. - **workflowID**: The ID of the workflow whose steps to retrieve. - **opts**: Optional configuration, documented below. ##### WithStepsLoadOutput ```go func WithStepsLoadOutput(loadOutput bool) GetWorkflowStepsOption ``` Control whether to load step output data. When unset, output is loaded only if the DBOS context has been launched. ##### WithStepsLimit ```go func WithStepsLimit(limit int) GetWorkflowStepsOption ``` Limit the number of steps returned, ordered by step ID ascending. ##### WithStepsOffset ```go func WithStepsOffset(offset int) GetWorkflowStepsOption ``` Skip the given number of steps before returning results. Combine with `WithStepsLimit` to paginate through a workflow's steps. #### CancelWorkflow ```go func CancelWorkflow(ctx Client, workflowID string, opts ...CancelWorkflowOption) error ``` Cancel a workflow. This sets its status to `CANCELLED`, removes it from its queue (if it is enqueued) and preempts its execution (interrupting it at the beginning of its next step, or waking it immediately if it is in a durable sleep). Pass `WithCancelChildren()` to also cancel all the workflow's child workflows, recursively. To cancel many workflows in a single database round-trip, use `CancelWorkflows(ctx, workflowIDs []string, opts ...CancelWorkflowOption)`. **Parameters:** - **ctx**: The DBOS client or context. - **workflowID**: The ID of the workflow to cancel. - **opts**: Optional configuration (e.g., `WithCancelChildren()`). #### SetWorkflowAttributes ```go func SetWorkflowAttributes(ctx Client, workflowID string, attributes map[string]any) error ``` Replace the custom attributes attached to an existing workflow. Pass a `nil` attributes map to clear all attributes. Attach attributes at creation with the `WithWorkflowAttributes(map[string]any)` workflow option, and search workflows by attributes with the `WithFilterAttributes(map[string]any)` ListWorkflows option (Postgres only). #### ResumeWorkflow ```go func ResumeWorkflow[R any](ctx Client, workflowID string, opts ...ResumeWorkflowOption) (WorkflowHandle[R], error) ``` Resume a workflow. This immediately starts it from its last completed step. You can use this to resume workflows that are cancelled or have exceeded their maximum recovery attempts. You can also use this to start an enqueued workflow immediately, bypassing its queue. **Parameters:** - **ctx**: The DBOS client or context. - **workflowID**: The ID of the workflow to resume. - **opts**: Optional configuration. ##### WithResumeQueue ```go func WithResumeQueue(queueName string) ResumeWorkflowOption ``` Re-enqueue the resumed workflow on the specified queue instead of starting it immediately. #### ResumeWorkflows ```go func ResumeWorkflows[R any](ctx Client, workflowIDs []string, opts ...ResumeWorkflowOption) ([]WorkflowHandle[R], error) ``` Resume multiple workflows in a single database round-trip. Each workflow that exists and is not in a terminal state is re-enqueued; completed or missing workflows are skipped. Unlike `ResumeWorkflow`, this function does not return an error when some IDs are missing. Accepts the same options as `ResumeWorkflow` (e.g., `WithResumeQueue`). #### ForkWorkflow ```go func ForkWorkflow[R any](ctx Client, input ForkWorkflowInput) (WorkflowHandle[R], error) ``` Start a new execution of a workflow from a specific step. The input step ID (`startStep`) must match the step number of the step returned by workflow introspection. The specified `startStep` is the step from which the new workflow will start, so any steps whose ID is less than `startStep` will not be re-executed. **Parameters:** - **ctx**: The DBOS client or context. - **input**: A `ForkWorkflowInput` struct where `OriginalWorkflowID` is mandatory. ```go type ForkWorkflowInput struct { OriginalWorkflowID string // Required: The UUID of the original workflow to fork from ForkedWorkflowID string // Optional: Custom workflow ID for the forked workflow (auto-generated if empty) StartStep uint // Optional: Step to start the forked workflow from (default: 0) ApplicationVersion string // Optional: Application version for the forked workflow (inherits from original if empty) QueueName string // Optional: Queue to enqueue the forked workflow on (defaults to starting immediately) } ``` #### SetWorkflowDelay ```go func SetWorkflowDelay(ctx Client, workflowID string, opts ...SetWorkflowDelayOption) error ``` Set or update the delay on a `DELAYED` workflow. Provide exactly one of `WithDelayDuration` (relative) or `WithDelayUntil` (absolute). Only affects workflows currently in the `DELAYED` status. ```go func WithDelayDuration(d time.Duration) SetWorkflowDelayOption func WithDelayUntil(t time.Time) SetWorkflowDelayOption ``` #### Workflow Status Some workflow introspection and management methods return a `WorkflowStatus`. This object has the following definition: ```go type WorkflowStatus struct { ID string `json:"workflow_uuid"` // Unique identifier for the workflow Status WorkflowStatusType `json:"status"` // Current execution status Name string `json:"name"` // Function name of the workflow AuthenticatedUser *string `json:"authenticated_user"` // User who initiated the workflow (if applicable) AssumedRole *string `json:"assumed_role"` // Role assumed during execution (if applicable) AuthenticatedRoles *string `json:"authenticated_roles"` // Roles available to the user (if applicable) Output any `json:"output"` // Workflow output (available after completion) Error error `json:"error"` // Error information (if status is ERROR) ExecutorID string `json:"executor_id"` // ID of the executor running this workflow CreatedAt time.Time `json:"created_at"` // When the workflow was created UpdatedAt time.Time `json:"updated_at"` // When the workflow status was last updated ApplicationVersion string `json:"application_version"` // Version of the application that created this workflow ApplicationID string `json:"application_id"` // Application identifier Attempts int `json:"attempts"` // Number of execution attempts QueueName string `json:"queue_name"` // Queue name (if workflow was enqueued) Timeout time.Duration `json:"timeout"` // Workflow timeout duration Deadline time.Time `json:"deadline"` // Absolute deadline for workflow completion StartedAt time.Time `json:"started_at"` // When the workflow execution actually started CompletedAt time.Time `json:"completed_at"` // When the workflow reached a terminal state (SUCCESS, ERROR, or CANCELLED) ForkedFrom string `json:"forked_from"` // ID of the original workflow if this is a fork WasForkedFrom bool `json:"was_forked_from"` // Whether this workflow has been forked from ParentWorkflowID string `json:"parent_workflow_id"` // ID of the parent workflow if this is a child DeduplicationID string `json:"deduplication_id"` // Deduplication identifier (if applicable) Input any `json:"input"` // Input parameters passed to the workflow Priority int `json:"priority"` // Execution priority (lower numbers have higher priority) DelayUntil time.Time `json:"delay_until"` // Time before which a DELAYED workflow should not be dequeued Attributes map[string]any `json:"attributes"` // Custom key-value attributes attached to the workflow } ``` ##### WorkflowStatusType The `WorkflowStatusType` represents the execution status of a workflow: ```go type WorkflowStatusType string const ( WorkflowStatusPending WorkflowStatusType = "PENDING" // Workflow is running or ready to run WorkflowStatusEnqueued WorkflowStatusType = "ENQUEUED" // Workflow is queued and waiting for execution WorkflowStatusDelayed WorkflowStatusType = "DELAYED" // Workflow is delayed and will transition to ENQUEUED after the delay expires WorkflowStatusSuccess WorkflowStatusType = "SUCCESS" // Workflow completed successfully WorkflowStatusError WorkflowStatusType = "ERROR" // Workflow completed with an error WorkflowStatusCancelled WorkflowStatusType = "CANCELLED" // Workflow was cancelled (manually or due to timeout) WorkflowStatusMaxRecoveryAttemptsExceeded WorkflowStatusType = "MAX_RECOVERY_ATTEMPTS_EXCEEDED" // Workflow exceeded maximum retry attempts ) ``` ### DBOS Variables #### GetWorkflowID ```go func GetWorkflowID(ctx Context) (string, error) ``` Return the ID of the current workflow, if in a workflow. Returns an error if not called from within a workflow context. **Parameters:** - **ctx**: The DBOS context. #### GetStepID ```go func GetStepID(ctx Context) (int, error) ``` Return the unique ID of the current step within a workflow. Returns an error if not called from within a step context. **Parameters:** - **ctx**: The DBOS context. Workflow queues allow you to ensure that workflow functions will be run, without starting them immediately. Queues are useful for controlling the number of workflows run in parallel, or the rate at which they are started. Queue configuration is persisted to the system database, so any DBOS process connected to the same system database can register, retrieve, and reconfigure queues. #### RegisterQueue ```go func RegisterQueue(ctx Client, name string, options ...QueueOption) (Queue, error) ``` Register a queue and persist its configuration to the system database, returning a `Queue`. If a queue with the same name already exists in the database, the `WithQueueOnConflict` option controls whether its configuration is overwritten. Queues may be registered at any time, including after `Launch()`; live workers periodically reload queue configuration, so changes take effect without a restart. You can enqueue a workflow by passing the returned `Queue` handle to the `WithQueue` option of `RunWorkflow`. **Parameters:** - **ctx**: The DBOS client or context. - **name**: The name of the queue. Must be unique among all queues in the application. - **options**: Functional options for the queue, documented below. **Example Syntax:** ```go queue, err := dbos.RegisterQueue(ctx, "email-queue", dbos.WithWorkerConcurrency(5), dbos.WithRateLimiter(&dbos.RateLimiter{ Limit: 100, Period: 60 * time.Second, // 100 workflows per minute }), dbos.WithPriorityEnabled(), ) // Enqueue workflows to this queue by passing its handle to WithQueue: handle, err := dbos.RunWorkflow(ctx, SendEmailWorkflow, emailData, dbos.WithQueue(queue)) ``` The returned `Queue` interface has `Get*` methods reflecting the queue's configuration as of the most recent read from the database, and `Set*` methods that update the configuration in the database at runtime: ```go type Queue interface { GetName() string GetGlobalConcurrency() *int GetWorkerConcurrency() *int GetRateLimit() *RateLimiter GetPartitionConcurrency() *int GetPartitionWorkerConcurrency() *int GetPartitionRateLimit() *RateLimiter GetPriorityEnabled() bool GetPartitionQueue() bool GetPollingInterval() time.Duration SetGlobalConcurrency(ctx Client, value *int) error SetWorkerConcurrency(ctx Client, value *int) error SetRateLimit(ctx Client, value *RateLimiter) error SetPartitionConcurrency(ctx Client, value *int) error SetPartitionWorkerConcurrency(ctx Client, value *int) error SetPartitionRateLimit(ctx Client, value *RateLimiter) error SetPriorityEnabled(ctx Client, value bool) error SetPartitionQueue(ctx Client, value bool) error SetPollingInterval(ctx Client, value time.Duration) error } ``` ##### WithQueueOnConflict ```go func WithQueueOnConflict(policy QueueConflictResolution) QueueOption const ( QueueConflictUpdateIfLatestVersion QueueConflictResolution = "update_if_latest_version" QueueConflictAlwaysUpdate QueueConflictResolution = "always_update" QueueConflictNeverUpdate QueueConflictResolution = "never_update" ) ``` Set how `RegisterQueue` behaves when a queue with the same name already exists in the system database: - **QueueConflictUpdateIfLatestVersion** (default): overwrite the existing configuration only if the running application is the latest registered application version. - **QueueConflictAlwaysUpdate**: always overwrite the existing configuration. - **QueueConflictNeverUpdate**: leave the existing configuration unchanged. #### RetrieveQueue ```go func RetrieveQueue(ctx Client, name string) (Queue, error) ``` Retrieve a queue by name from the system database. If no queue with that name has been registered, returns an error matching `dbos.ErrQueueNotFound`. #### ListQueues ```go func ListQueues(ctx Client) ([]Queue, error) ``` Return all queues registered in the system database. #### DeleteQueue ```go func DeleteQueue(ctx Client, name string) error ``` Delete a queue from the system database. No-op if no queue with that name exists. Workflows already enqueued on a deleted queue can no longer be dequeued, executed, or recovered — unless a queue with the same name is later registered, in which case it will dequeue the leftover workflows. Do not rely on this behavior: cancel or drain pending workflows on the queue before deleting it. ##### WithWorkerConcurrency ```go func WithWorkerConcurrency(concurrency int) QueueOption ``` Set the maximum number of workflows from this queue that may run concurrently within a single DBOS process. ##### WithGlobalConcurrency ```go func WithGlobalConcurrency(concurrency int) QueueOption ``` Set the maximum number of workflows from this queue that may run concurrently. Defaults to 0 (no limit). This concurrency limit is global across all DBOS processes using this queue. ##### WithPriorityEnabled ```go func WithPriorityEnabled() QueueOption ``` Enable setting priority for workflows on this queue. ##### WithRateLimiter ```go func WithRateLimiter(limiter *RateLimiter) QueueOption ``` ```go type RateLimiter struct { Limit int // Maximum number of workflows to start within the period Period time.Duration // Time period for the rate limit } ``` A limit on the maximum number of functions which may be started in a given period. ##### WithPartitionConcurrency ```go func WithPartitionConcurrency(concurrency int) QueueOption ``` Set the maximum number of workflows from any one partition of this queue that may run concurrently across all DBOS processes. Must be at least 1 and less than or equal to the queue's global concurrency. Setting any partition limit makes the queue partitioned: workflows must then be enqueued with `WithQueuePartitionKey`. ##### WithPartitionWorkerConcurrency ```go func WithPartitionWorkerConcurrency(concurrency int) QueueOption ``` Set the maximum number of workflows from any one partition of this queue that may run concurrently within a single DBOS process. Must be at least 1 and less than or equal to the queue's partition concurrency, worker concurrency, and global concurrency. ##### WithPartitionRateLimiter ```go func WithPartitionRateLimiter(limiter *RateLimiter) QueueOption ``` A limit on the maximum number of workflows which may be started from any one partition in a given period, applied to each partition separately. ##### WithPartitionQueue ```go func WithPartitionQueue() QueueOption ``` Deprecated: use a partition limit instead. Enables the legacy partitioned mode, under which the queue's global concurrency, worker concurrency, and rate limit each apply per partition. Cannot be combined with a partition limit. #### RegisterWorkflow ```go func RegisterWorkflow[P any, R any](ctx Context, fn Workflow[P, R], opts ...WorkflowRegistrationOption) ``` Register a function as a DBOS workflow. All workflows must be registered before the context is launched. Workflow functions must be compatible with the following signature: ```go type Workflow[P any, R any] func(ctx Context, input P) (R, error) ``` **Parameters:** - **ctx**: The Context. - **fn**: The workflow function to register. - **opts**: Functional options for workflow registration, documented below. ##### WithMaxRecoveryAttempts ```go func WithMaxRecoveryAttempts(maxRetries int) WorkflowRegistrationOption ``` Configure the maximum number of times execution of a workflow may be attempted. This acts as a dead letter queue so that a buggy workflow that crashes its application (for example, by running it out of memory) does not do so infinitely. If a workflow exceeds this limit, its status is set to `MAX_RECOVERY_ATTEMPTS_EXCEEDED` and it may no longer be executed. ##### WithWorkflowName ```go func WithWorkflowName(name string) WorkflowRegistrationOption ``` Register a workflow with a custom name. If not provided, the name of the workflow function is used. ##### WithInstance ```go func WithInstance(instance ConfiguredInstance) WorkflowRegistrationOption ``` Register a workflow method bound to a specific configured instance. Method values bound to different receivers (e.g. `a.Run` and `b.Run`) share a function name, so each instance's method must be registered under a per-instance key, derived from the instance's config name. The instance must implement the `ConfiguredInstance` interface: ```go type ConfiguredInstance interface { ConfigName() string } ``` `ConfigName` must return a stable, unique name for the instance: it is durably recorded so recovery runs the workflow on the correct instance. Instances must be registered with the same config name on every process start, before `Launch()`. ```go dbos.RegisterWorkflow(ctx, slack.Send, dbos.WithInstance(slack)) dbos.RegisterWorkflow(ctx, email.Send, dbos.WithInstance(email)) ``` Run a workflow registered with `WithInstance` using the matching `WithRunInstance` option. #### RunWorkflow ```go func RunWorkflow[P any, R any](ctx Context, fn Workflow[P, R], input P, opts ...WorkflowOption) (WorkflowHandle[R], error) ``` Execute a workflow function. The workflow may execute immediately or be enqueued for later execution based on options. Returns a WorkflowHandle that can be used to check the workflow's status or wait for its completion and retrieve its results. **Parameters:** - **ctx**: The Context. - **fn**: The workflow function to execute. - **input** The input to the workflow function. - **opts**: Functional options for workflow execution, documented below. **Example Syntax**: ```go func workflow(ctx dbos.Context, input string) (string, error) { return "success", nil } func example(input string) error { handle, err := dbos.RunWorkflow(dbosContext, workflow, input) if err != nil { return err } result, err := handle.GetResult() if err != nil { return err } fmt.Println("Workflow result:", result) return nil } ``` ##### WithWorkflowID ```go func WithWorkflowID(id string) WorkflowOption ``` Run the workflow with a custom workflow ID. If not specified, a UUID workflow ID is generated. ##### WithRunInstance ```go func WithRunInstance(instance ConfiguredInstance) WorkflowOption ``` Run a workflow method registered with `WithInstance`. The instance's config name selects the per-instance registration, so the workflow executes on (and recovers to) the correct instance. ```go handle, err := dbos.RunWorkflow(ctx, slack.Send, input, dbos.WithRunInstance(slack)) ``` ##### WithQueue ```go func WithQueue(queue Queue) WorkflowOption ``` Enqueue the workflow to the given queue instead of executing it immediately. Queued workflows will be dequeued and executed according to the queue's configuration. The queue must be a non-nil `Queue` handle returned by `RegisterQueue`, `RetrieveQueue`, or `ListQueues`; passing `nil` makes the enclosing `RunWorkflow` call return an error. To enqueue by name instead (for example, from a standalone client), use `Enqueue`. ##### WithDeduplicationID ```go func WithDeduplicationID(id string) WorkflowOption ``` Set a deduplication ID for this workflow. Should be used alongside `WithQueue`. At any given time, only one workflow with a specific deduplication ID can be enqueued in a given queue. ##### WithDeduplicationPolicy ```go func WithDeduplicationPolicy(policy DeduplicationPolicy) WorkflowOption ``` Set how a colliding deduplication ID is handled for a queued workflow. Must be used alongside `WithQueue` and `WithDeduplicationID`. With the default `DeduplicationPolicyReject`, a colliding enqueue fails with a `ErrorCodeQueueDeduplicated` error; with `DeduplicationPolicyReturnExisting`, it instead returns a handle to the existing workflow. ##### WithPriority ```go func WithPriority(priority uint) WorkflowOption ``` Set a queue priority for the workflow. Should be used alongside `WithQueue`. Workflows with the same priority are dequeued in **FIFO (first in, first out)** order. Priority values can range from `1` to `2,147,483,647`, where **a low number indicates a higher priority**. Workflows without assigned priorities have the highest priority and are dequeued before workflows with assigned priorities. ##### WithQueuePartitionKey ```go func WithQueuePartitionKey(partitionKey string) WorkflowOption ``` Set a queue partition key for the workflow. Use if and only if the queue is partitioned (registered with at least one partition limit, such as `WithPartitionConcurrency`). ##### WithDelay ```go func WithDelay(delay time.Duration) WorkflowOption ``` Delay execution of a queued workflow by the specified duration. Must be used together with `WithQueue`. The workflow is initially placed in `DELAYED` status and does not execute until the delay expires, at which point it transitions to `ENQUEUED` and may be dequeued. The delay can later be updated via `SetWorkflowDelay`. ##### WithApplicationVersion ```go func WithApplicationVersion(version string) WorkflowOption ``` Set the application version for this workflow, overriding the version in Context. ##### WithAuthenticatedUser ```go func WithAuthenticatedUser(user string) WorkflowOption ``` Associate the workflow execution with a user name. Useful to define workflow identity. Child workflows automatically inherit their parent's authentication information (authenticated user, assumed role, and authenticated roles) unless explicitly overridden. ##### WithWorkflowAttributes ```go func WithWorkflowAttributes(attributes map[string]any) WorkflowOption ``` Attach custom key-value attributes to the workflow. Attributes are recorded in the workflow status at creation, must be JSON-serializable, and are not inherited by child workflows. On Postgres they can be searched with the `WithFilterAttributes(map[string]any)` ListWorkflows option, and replaced later with `SetWorkflowAttributes`. #### RunAsStep ```go func RunAsStep[R any](ctx Context, fn Step[R], opts ...StepOption) (R, error) ``` Execute a function as a step in a durable workflow. **Parameters:** - **ctx**: The Context. - **fn**: The step to execute, typically wrapped in an anonymous function. Syntax shown below. - **opts**: Functional options for step execution, documented below. **Example Syntax:** Any Go function can be a step as long as it outputs one json-encodable value and an error. To pass inputs into a function being called as a step, wrap it in an anonymous function as shown below: ```go func step(ctx context.Context, input string) (string, error) { output := ... return output } func workflow(ctx dbos.Context, input string) (string, error) { output, err := dbos.RunAsStep( ctx, func(stepCtx context.Context) (string, error) { return step(stepCtx, input) } ) } ``` ##### WithStepName ```go func WithStepName(name string) StepOption ``` Set a custom name for a step. ##### WithStepMaxRetries ```go func WithStepMaxRetries(maxRetries int) StepOption ``` Set the maximum number of times this step is automatically retired on failure. A value of 0 (the default) indicates no retries. ##### WithStepMaxInterval ```go func WithStepMaxInterval(interval time.Duration) StepOption ``` WithStepMaxInterval sets the maximum delay between retries. Default value is 5s. ##### WithStepBackoffFactor ```go func WithStepBackoffFactor(factor float64) StepOption ``` WithStepBackoffFactor sets the exponential backoff multiplier between retries. Default value is 2.0. ##### WithStepBaseInterval ```go func WithStepBaseInterval(interval time.Duration) StepOption ``` WithStepBaseInterval sets the initial delay between retries. Default value is 100ms. #### Go ```go func Go[R any](ctx Context, fn Step[R], opts ...StepOption) (<-chan StepOutcome[R], error) ``` Launch a step asynchronously and return a receive-only channel that will receive the result when the step completes. This is a durable alternative to Go's native goroutines. Can only be called from within a workflow (not from inside a step). ```go type StepOutcome[R any] struct { Result R Err error } ``` #### Select ```go func Select[R any](ctx Context, channels []<-chan StepOutcome[R]) (R, error) ``` Wait for and return the first result from multiple channels obtained from `Go`. This is a durable alternative to Go's native `select` statement. Can only be called from within a workflow. #### WorkflowHandle ```go type WorkflowHandle[R any] interface { GetResult(opts ...GetResultOption) (R, error) GetStatus() (WorkflowStatus, error) GetWorkflowID() string } ``` WorkflowHandle provides methods to interact with a running or completed workflow. The type parameter `R` represents the expected return type of the workflow. Handles can be used to wait for workflow completion, check status, and retrieve results. ##### WorkflowHandle.GetResult ```go WorkflowHandle.GetResult(opts ...GetResultOption) (R, error) ``` Wait for the workflow to complete and return its result. ##### WorkflowHandle.GetStatus ```go WorkflowHandle.GetStatus() (WorkflowStatus, error) ``` Retrieve the WorkflowStatus of the workflow. ##### WorkflowHandle.GetWorkflowID ```go WorkflowHandle.GetWorkflowID() string ``` Retrieve the ID of the workflow. `Client` provides a programmatic way to interact with your DBOS application from external code. Because every `Context` **is** a `Client` (the `Context` interface embeds `Client`), all the package-level functions documented above whose first parameter is a `Client` work identically with a standalone client. Use them by passing your client where you would pass a DBOS context; only functions requiring a `Context` (workflow registration and execution, workflow-scoped operations) are unavailable on a standalone client. This is the `Client` interface: ```go type Client interface { context.Context // Workflow operations Enqueue(_ Client, queueName string, workflowName string, input any, opts ...EnqueueOption) (WorkflowHandle[any], error) Send(_ Client, destinationID string, message any, topic string, opts ...SendOption) error GetEvent(_ Client, targetWorkflowID string, key string, timeout time.Duration) (any, error) ReadStream(_ Client, workflowID string, key string, opts ...ReadStreamOption) ([]any, bool, error) ReadStreamAsync(_ Client, workflowID string, key string) (<-chan StreamValue[any], error) // Workflow management RetrieveWorkflow(_ Client, workflowID string) (WorkflowHandle[any], error) CancelWorkflow(_ Client, workflowID string, opts ...CancelWorkflowOption) error CancelWorkflows(_ Client, workflowIDs []string, opts ...CancelWorkflowOption) error SetWorkflowAttributes(_ Client, workflowID string, attributes map[string]any) error SetWorkflowDelay(_ Client, workflowID string, opts ...SetWorkflowDelayOption) error ResumeWorkflow(_ Client, workflowID string, opts ...ResumeWorkflowOption) (WorkflowHandle[any], error) ResumeWorkflows(_ Client, workflowIDs []string, opts ...ResumeWorkflowOption) ([]WorkflowHandle[any], error) ForkWorkflow(_ Client, input ForkWorkflowInput) (WorkflowHandle[any], error) ForkWorkflows(_ Client, input ForkWorkflowsInput) ([]WorkflowHandle[any], error) ListWorkflows(_ Client, opts ...ListWorkflowsOption) ([]WorkflowStatus, error) GetWorkflowSteps(_ Client, workflowID string, opts ...GetWorkflowStepsOption) ([]StepInfo, error) GetWorkflowAggregates(_ Client, input GetWorkflowAggregatesInput) ([]WorkflowAggregateRow, error) GetStepAggregates(_ Client, input GetStepAggregatesInput) ([]StepAggregateRow, error) DeleteWorkflows(_ Client, workflowIDs []string, opts ...DeleteWorkflowOption) error // Queue management RegisterQueue(_ Client, name string, options ...QueueOption) (Queue, error) RetrieveQueue(_ Client, name string) (Queue, error) ListQueues(_ Client) ([]Queue, error) DeleteQueue(_ Client, name string) error // Schedule management CreateSchedule(_ Client, spec ScheduleSpec) error ApplySchedules(_ Client, schedules []ScheduleSpec) error PauseSchedule(_ Client, scheduleName string) error ResumeSchedule(_ Client, scheduleName string) error DeleteSchedule(_ Client, scheduleName string) error GetSchedule(_ Client, scheduleName string) (WorkflowSchedule, error) ListSchedules(_ Client, opts ...ListSchedulesOption) ([]WorkflowSchedule, error) BackfillSchedule(_ Client, scheduleName string, start, end time.Time) ([]string, error) TriggerSchedule(_ Client, scheduleName string) (WorkflowHandle[any], error) // Application version management ListApplicationVersions(_ Client) ([]VersionInfo, error) GetLatestApplicationVersion(_ Client) (VersionInfo, error) SetLatestApplicationVersion(_ Client, versionName string) error Shutdown(_ Client, timeout time.Duration) error } ``` Prefer the generic package-level functions over calling interface methods directly: they return typed handles or values (`RetrieveWorkflow[R]`, `ResumeWorkflow[R]`, `ResumeWorkflows[R]`, `ForkWorkflow[R]`, `TriggerSchedule[R]`, `GetEvent[R]`, `ReadStream[R]`, `ReadStreamAsync[R]`, `Enqueue[R]`), while the interface methods return `any`. #### Constructor ```go func NewClient(ctx context.Context, config ClientConfig) (Client, error) ``` **Parameters:** - `ctx`: A context for initialization operations - `config`: A `ClientConfig` object with connection and application settings ```go type ClientConfig struct { DatabaseURL string // Connection string to your system database. May be a PostgreSQL (postgres://...) or SQLite (sqlite:...) URL. Exactly one of DatabaseURL, SystemDBPool, or SQLiteSystemDB is required. SQLite URLs additionally require the driver import: import _ "github.com/dbos-inc/dbos-transact-golang/dbos/driver/sqlite" SystemDBPool *pgxpool.Pool // A custom Postgres/CockroachDB pool. Optional; takes precedence over DatabaseURL. Mutually exclusive with SQLiteSystemDB. SQLiteSystemDB *sql.DB // A custom SQLite handle (e.g. from modernc.org/sqlite). Optional; takes precedence over DatabaseURL. Mutually exclusive with SystemDBPool. DatabaseSchema string // Database schema name (defaults to "dbos"; Postgres only) Logger *slog.Logger // Optional custom logger Serializer Serializer[any] // Optional custom serializer (defaults to JSON) SystemDBStartupTimeout time.Duration // Maximum time for system database connection and migrations (default: 2 minutes) } ``` **Returns:** - A new `Client` instance or an error if initialization fails **Example syntax:** This DBOS client connects to the system database specified in the configuration: ```go config := dbos.ClientConfig{ DatabaseURL: os.Getenv("DBOS_SYSTEM_DATABASE_URL"), } client, err := dbos.NewClient(context.Background(), config) if err != nil { log.Fatal(err) } defer dbos.Shutdown(client, 5*time.Second) ``` A client manages a connection pool to the DBOS system database. Shut it down with the same unified `dbos.Shutdown(client, timeout)` function documented above, which releases the connection pool. ### Workflow Interaction Methods #### Enqueue ```go func Enqueue[R any, P any]( ctx Client, queueName string, workflowName string, input P, opts ...EnqueueOption ) (WorkflowHandle[R], error) ``` The result type parameter `R` comes first so you can name only it and let the input type be inferred: `dbos.Enqueue[MyOutput](client, ...)`. Enqueue a workflow for processing and return a handle to it, similar to RunWorkflow with the WithQueue option. Returns a WorkflowHandle. When enqueuing a workflow from the DBOS client, you must specify the name of the workflow to enqueue (rather than passing a workflow function as with `RunWorkflow`.) Required parameters: * `ctx`: The DBOS client (or context) * `queueName`: The name of the queue on which to enqueue the workflow * `workflowName`: The name of the workflow function being enqueued * `input`: The input to pass to the workflow Optional configuration via `EnqueueOption`: * `WithEnqueueWorkflowID(id string)`: The unique ID for the enqueued workflow. If left undefined, DBOS Client will generate a UUID. Please see Workflow IDs and Idempotency for more information. * `WithEnqueueApplicationVersion(version string)`: The version of your application that should process this workflow. If left undefined, it will use the current application version. * `WithEnqueueTimeout(timeout time.Duration)`: Set a timeout for the enqueued workflow. When the timeout expires, the workflow **and all its children** are cancelled (except if the child's context has been made uncancellable using `WithoutCancel`). The timeout does not begin until the workflow is dequeued and starts execution. * `WithEnqueueDeduplicationID(id string)`: At any given time, only one workflow with a specific deduplication ID can be enqueued in the specified queue. If a workflow with a deduplication ID is currently enqueued or actively executing (status `ENQUEUED` or `PENDING`), subsequent workflow enqueue attempts with the same deduplication ID in the same queue will fail. * `WithEnqueueDeduplicationPolicy(policy DeduplicationPolicy)`: Set how a colliding deduplication ID is handled. Requires `WithEnqueueDeduplicationID`. With the default `DeduplicationPolicyReject`, a colliding enqueue fails with a `ErrorCodeQueueDeduplicated` error; with `DeduplicationPolicyReturnExisting`, it instead returns a handle to the existing workflow. * `WithEnqueuePriority(priority uint)`: The priority of the enqueued workflow in the specified queue. Workflows with the same priority are dequeued in **FIFO (first in, first out)** order. Priority values can range from `1` to `2,147,483,647`, where **a low number indicates a higher priority**. Workflows without assigned priorities have the highest priority and are dequeued before workflows with assigned priorities. * `WithEnqueueDelay(delay time.Duration)`: Delay execution of the enqueued workflow by the specified duration. The workflow is initially placed in `DELAYED` status and transitions to `ENQUEUED` after the delay expires. The delay can later be updated via `SetWorkflowDelay`. * `WithEnqueueQueuePartitionKey(partitionKey string)`: Set the queue partition key. Required if and only if the target queue is partitioned (registered with at least one partition limit, such as `WithPartitionConcurrency`). **Example syntax:** ```go type ProcessInput struct { TaskID string Data string } type ProcessOutput struct { Result string Status string } handle, err := dbos.Enqueue[ProcessOutput]( client, "process_queue", "ProcessWorkflow", ProcessInput{TaskID: "task-123", Data: "data"}, dbos.WithEnqueueTimeout(30 * time.Minute), dbos.WithEnqueuePriority(5), ) if err != nil { log.Fatal(err) } result, err := handle.GetResult() if err != nil { log.Printf("Workflow failed: %v", err) } else { log.Printf("Result: %+v", result) } ``` All other workflow interaction, management, streaming, queue, and schedule functions are the same package-level functions documented above (`Send`, `GetEvent`, `RetrieveWorkflow`, `ListWorkflows`, `GetWorkflowSteps`, `CancelWorkflow`/`CancelWorkflows`, `ResumeWorkflow`/`ResumeWorkflows`, `ForkWorkflow`/`ForkWorkflows`, `SetWorkflowAttributes`, `SetWorkflowDelay`, `DeleteWorkflows`, `ReadStream`, `ReadStreamAsync`, `RegisterQueue`, `RetrieveQueue`, `ListQueues`, `DeleteQueue`, `CreateSchedule`, `ApplySchedules`, and the rest): they take `ctx Client` as their first parameter, so pass your client where a DBOS application would pass its context. Note that with `ListWorkflows` and `GetWorkflowSteps` from a standalone client, workflow inputs, outputs, and step outputs are not loaded or decoded by default; pass `WithFilterLoadInput(true)` / `WithFilterLoadOutput(true)` / `WithStepsLoadOutput(true)` to opt in. #### NewDebouncerClient ```go func NewDebouncerClient[R any, P any](workflowName string, client Client, opts ...DebouncerOption) *DebouncerClient[R, P] ``` Both type parameters must be named explicitly (the workflow is referenced by name, so neither can be inferred): `dbos.NewDebouncerClient[MyOutput, MyInput]("workflowName", client)`. Create a new debouncer client for use from outside a DBOS application. Similar to `NewDebouncer` but uses a Client instead of a Context and takes a workflow name string instead of a function reference. To debounce a workflow registered on a configured instance, pass the instance's config name with `WithDebouncerConfigName(configName string)`. ### Cross-Language Portable Types #### WithPortableWorkflow ```go func WithPortableWorkflow() WorkflowOption ``` Mark a workflow to use portable JSON serialization for cross-language interoperability. #### WithPortableSend / WithPortableSetEvent / WithPortableWriteStream ```go func WithPortableSend() SendOption func WithPortableSetEvent() SetEventOption func WithPortableWriteStream() WriteStreamOption ``` Use portable JSON for cross-language Send, SetEvent, and WriteStream operations. #### PortableWorkflowArgs ```go type PortableWorkflowArgs struct { PositionalArgs []any `json:"positional_args,omitempty"` NamedArgs map[string]any `json:"named_args,omitempty"` } ``` Cross-language envelope for workflow inputs. When passed as the input to `Enqueue`, portable JSON serialization is used automatically. #### PortableWorkflowError ```go type PortableWorkflowError struct { Name string Message string Code any Data any } ``` Structured error type for portable workflows. Return this from a workflow to pass structured error info cross-language. #### WithEnqueueClassName / WithEnqueueConfigName ```go func WithEnqueueClassName(className string) EnqueueOption func WithEnqueueConfigName(configName string) EnqueueOption ``` Set the class/namespace and config/instance name when enqueueing to Python, TypeScript, or Java targets. `WithEnqueueConfigName` is also required when enqueueing to a Go workflow registered on a configured instance with `WithInstance`; the value must match the instance's config name. ### Alerting #### SetAlertHandler ```go func SetAlertHandler(ctx Context, handler AlertHandler) ``` ```go type AlertHandler func(name string, message string, metadata map[string]string) ``` Register a handler to receive alerts from Conductor. Must be called before `Launch()`. Only one handler per application. If no handler is registered, alerts are logged automatically. ```` --- ## DBOS CLI import {RedirectToGitHub} from '@site/src/components/RedirectToGitHub'; [Click here to view the DBOS CLI documentation →](https://pkg.go.dev/github.com/dbos-inc/dbos-transact-golang/cmd/dbos) --- ## Configuration ### Configuring DBOS To configure DBOS, pass a `Config` object to [`NewContext`](./dbos-context.md#newcontext). `AppName` and one of `DatabaseURL`, `SystemDBPool`, or `SQLiteSystemDB` are mandatory. ```go type Config struct { AppName string // Application name for identification (required) DatabaseURL string // Connection string to your system database. May be a PostgreSQL/CockroachDB URL (postgres://...), a key=value DSN, or a SQLite (sqlite:...) URL. Exactly one of DatabaseURL, SystemDBPool, or SQLiteSystemDB is required. SystemDBPool *pgxpool.Pool // A custom Postgres/CockroachDB connection pool DBOS can use to access your system database. Optional; takes precedence over DatabaseURL. Mutually exclusive with SQLiteSystemDB. SQLiteSystemDB *sql.DB // A custom SQLite handle (e.g. from modernc.org/sqlite) DBOS can use as your system database. Optional; takes precedence over DatabaseURL. Mutually exclusive with SystemDBPool. DatabaseSchema string // Database schema name (defaults to "dbos"; Postgres/CockroachDB only, ignored by SQLite) Logger *slog.Logger // Custom logger instance (defaults to a new slog logger) ConductorURL string // DBOS conductor service URL (optional) ConductorAPIKey string // DBOS conductor API key (optional) ConductorExecutorMetadata map[string]any // Metadata used to identify this executor on the Conductor dashboard (optional, must be JSON-serializable) ApplicationVersion string // Application version (optional) ExecutorID string // Executor ID (optional) EnablePatching bool // Enable the patching system for Patch/DeprecatePatch (default: false) Serializer Serializer[any] // Custom serializer for workflow inputs, outputs, and events (defaults to a JSON serializer). See the serialization reference. SchedulerPollingInterval time.Duration // How often database-backed schedules are reconciled (default: 30s) SystemDBStartupTimeout time.Duration // Maximum time for system database connection and schema migration or verification (default: 2 minutes) SkipMigrations bool // Verify the system database schema on startup instead of creating and migrating it (default: false) AdminServer bool // Deprecated: run the HTTP admin server for workflow management operations (default: false) AdminServerPort int // Deprecated: port for the admin server (default: 3001) } ``` :::warning Deprecated `AdminServer` and `AdminServerPort` are deprecated and will be removed in v1.5.0. Use [DBOS Conductor](../../conductor/overview.md) for remote workflow management instead. ::: `ApplicationVersion` and `ExecutorID` are overridden by the `DBOS__APPVERSION` and `DBOS__VMID` environment variables, respectively, when set. `AppName` identifies your application. It must be between 3 and 256 characters long and contain only lowercase letters, numbers, dashes, and underscores. An application connecting to [Conductor](../../conductor/overview.md) (with `ConductorAPIKey` set) fails to start with a name outside that rule, because Conductor refuses to register it; a self-hosted application logs a warning and starts. Multiple applications (potentially in different languages) may [share a system database](../../explanations/sharing-a-system-database.md), in which case each must have a distinct name: the name identifies which application owns each workflow, queue, schedule, and application version, and applications only run their own workflows. If you rename an application, transfer ownership of its data with [`RenameApplication`](./methods.md#renameapplication) or the `dbos rename-application` CLI command. For example: ```go dbosContext, err := dbos.NewContext(context.Background(), dbos.Config{ AppName: "dbos-starter", ApplicationVersion: "0.1.0", DatabaseURL: os.Getenv("DBOS_SYSTEM_DATABASE_URL"), }) if err != nil { panic(err) } ``` To supply a custom serializer through `Config.Serializer`, see the [serialization reference](./workflows-steps.md#serialization). #### Using SQLite SQLite support is not linked into your binary by default — Postgres-only applications do not compile or link a SQLite driver. To use a SQLite system database (a `sqlite:` URL or `SQLiteSystemDB`), register the driver with one blank import, anywhere in your binary: ```go import _ "github.com/dbos-inc/dbos-transact-golang/dbos/driver/sqlite" ``` Without it, `NewContext` (or `NewClient`) fails at startup with an error naming this import. The import registers [modernc.org/sqlite](https://pkg.go.dev/modernc.org/sqlite), a pure-Go driver requiring no cgo. ### System database startup `NewContext` is the call that connects to your system database: it creates and validates the connection pool, creates the database if it does not exist, runs any pending [schema migrations](../../explanations/system-tables.md), and pings the database. If the database is unreachable or migrations fail, `NewContext` returns an initialization error — a failed `NewContext` (like a failed `Launch`) is terminal; create a fresh context for each attempt. The entire startup window is bounded by `Config.SystemDBStartupTimeout` (default: 2 minutes). On expiry, the returned error names the startup phase that timed out (connecting, running migrations, pinging, …), wraps `context.DeadlineExceeded`, and includes a diagnostic hint. In particular, if the connection pool had no free connections when the timeout expired — common when passing a shared `SystemDBPool` whose connections are checked out by the application — the error reports the pool's acquired/max connection counts and suggests increasing pool capacity or releasing checked-out connections. By default, DBOS creates its own pool: for Postgres/CockroachDB, at most 20 connections (1 hour max connection lifetime, 5 minutes idle timeout, 10 seconds connect timeout); for SQLite, at most 8 open connections. For Postgres/CockroachDB, a `pool_max_conns` parameter in `DatabaseURL` (e.g. `postgres://...?pool_max_conns=7`) overrides the default maximum pool size. To use different pool settings, construct the pool yourself and pass it via `Config.SystemDBPool` (Postgres/CockroachDB) or `Config.SQLiteSystemDB` (SQLite); DBOS uses it as-is. **Migrations** are versioned and recorded in the `dbos_migrations` table of your system database schema. `NewContext` applies only migrations newer than the recorded version, so startup against an up-to-date database performs no schema work, and re-running it is a no-op. Set `Config.SkipMigrations` to make `NewContext` **verify** the schema instead of creating and migrating it. Use it for a process whose database role cannot run DDL, or a deployment that migrates out of band with the `dbos migrate` [CLI command](./cli.md). `NewContext` then fails if the system database is missing or its schema is behind the version this build of DBOS requires. A [standalone client](./dbos-context.md#newclient) always behaves this way: it never creates or migrates the system database. After a successful `NewContext`, `Launch` and subsequent runtime operations do not fail fast on database outages: transient database errors are retried (indefinitely, until the context is cancelled or shut down). --- ## Datasources A datasource is a handle to a database **you** own, over which DBOS can run durable transactions. Use [`RunAsTransaction`](#runastransaction) to run a function inside a database transaction where your application writes and the DBOS durability record commit atomically, guaranteeing the transaction executes exactly once even across crashes and recovery. See the [Transactions & Datasources tutorial](../tutorials/transaction-tutorial.md) for a full walkthrough. #### NewDataSource ```go func NewDataSource[E Engine](ctx Context, engine E, opts ...DataSourceOption) (*DataSource, error) type Engine interface { *pgxpool.Pool | *sql.DB } ``` Build a durable data source over a user-provided database engine. The engine type is constrained at compile time to `*pgxpool.Pool` (Postgres/CockroachDB) or `*sql.DB` (SQLite). The returned handle is ready to use immediately: `NewDataSource` detects whether the engine is the DBOS system database, resolves the dialect (CockroachDB is auto-detected), and creates the `transaction_completion` durability table in your database if it does not already exist. It may be called at any time, before or after `Launch()`, and there is no registry: create as many data sources as you need. If the engine is the **same handle** as the DBOS system database (the pool you passed as [`Config.SystemDBPool` or `Config.SQLiteSystemDB`](./configuration.md)), DBOS does not create or manage a `transaction_completion` table at all: transactions on this data source commit your application writes and the DBOS checkpoint together in a single transaction against the system database. See [Sharing the System Database Engine](#sharing-the-system-database-engine). **Parameters:** - **ctx**: The Context. - **engine**: A `*pgxpool.Pool` or `*sql.DB` connecting to your database. - **opts**: Functional options, documented below. **Example Syntax:** ```go pool, err := pgxpool.New(context.Background(), os.Getenv("APP_DATABASE_URL")) if err != nil { log.Fatal(err) } ds, err := dbos.NewDataSource(ctx, pool, dbos.WithDataSourceName("app")) if err != nil { log.Fatal(err) } ``` ##### WithDataSourceName ```go func WithDataSourceName(name string) DataSourceOption ``` Set the data source's name (default `"datasource"`). It is used only in logs and error messages; names need not be unique. ##### WithDataSourceSchema ```go func WithDataSourceSchema(schema string) DataSourceOption ``` Override the schema that holds the `transaction_completion` table (default `"dbos"`). Ignored by SQLite, which has no schemas. ##### Provisioning and Permissions `NewDataSource` checks for the `transaction_completion` table first and skips all DDL if it already exists: - **First-time creation** needs a role with `CREATE` on the database (for the schema) and `CREATE` on the schema (for the table). - **Steady state** needs only `SELECT, INSERT` on `transaction_completion`, plus whatever your application tables require. Because the engine is user-provided, you choose the role: give the data source a role that can run the DDL, or pre-create the table in your own migrations and connect with a DML-only role. If the table is missing and the role cannot create it, `NewDataSource` fails fast with an actionable error. #### DataSource.Name ```go func (ds *DataSource) Name() string ``` Return the data source's name. #### RunAsTransaction ```go func RunAsTransaction[R any](ctx Context, ds *DataSource, fn Txn[R], opts ...StepOption) (R, error) type Txn[R any] func(ctx context.Context, tx Tx) (R, error) ``` Durably execute `fn` as a transaction against the data source `ds`. `fn` receives a portable [`Tx`](#the-tx-interface); within it your application can write its own tables, and DBOS atomically records a durability row in the same transaction, so the function runs exactly once even across crashes and recovery. `RunAsTransaction` must be called from within a workflow; calling it at top level returns an error. It cannot be nested: calling it inside another `RunAsTransaction` or inside a [`RunAsStep`](./workflows-steps.md#runasstep) is rejected with an error. It shares the per-workflow step counter with `RunAsStep`, so transactions and steps can be freely interleaved. Standard [step options](./workflows-steps.md#withstepname) apply (`WithStepName`, `WithStepMaxRetries`, retry intervals, retry predicate). Serialization and deadlock conflicts are retried internally with a fresh transaction; application errors follow your step retry policy. Additionally, `WithTxIsolation` sets the transaction's isolation level (default `ReadCommitted`): ```go func WithTxIsolation(level IsoLevel) StepOption ``` `IsoLevel` is one of `dbos.IsoLevelDefault`, `dbos.IsoLevelReadCommitted`, `dbos.IsoLevelRepeatableRead`, or `dbos.IsoLevelSerializable`. **Example Syntax:** ```go n, err := dbos.RunAsTransaction(ctx, ds, func(ctx context.Context, tx dbos.Tx) (int64, error) { res, err := tx.Exec(ctx, "INSERT INTO orders(item) VALUES ($1)", item) if err != nil { return 0, err } return res.RowsAffected() }, dbos.WithStepMaxRetries(3)) ``` ##### Durability and Recovery A top-level `RunAsTransaction` first commits your application writes together with a row in the `transaction_completion` table (one atomic transaction in your database), then checkpoints the result in the DBOS system database. Recovery checks the system database first, then `transaction_completion`, and replays the stored output without re-running `fn` if the transaction already committed—covering the crash window between the two commits. ##### Sharing the System Database Engine If the data source's engine is the very same handle as the DBOS system database (you passed your pool as [`Config.SystemDBPool` or `Config.SQLiteSystemDB`](./configuration.md) and reuse it here), DBOS does not create or manage a `transaction_completion` table for it. There is no need for one: `RunAsTransaction` collapses onto a single transaction in which application writes and the DBOS checkpoint commit together, and recovery replays from the system database alone. ```go pool, _ := pgxpool.New(context.Background(), databaseURL) ctx, _ := dbos.NewContext(context.Background(), dbos.Config{ AppName: "my-app", SystemDBPool: pool, }) // Same handle as the system database: single-transaction durability, // no transaction_completion table. ds, err := dbos.NewDataSource(ctx, pool) ``` Detection is by pointer identity, not connection string: a second pool to the same physical database is treated as a separate database and uses the regular two-transaction path with a `transaction_completion` table. #### The Tx Interface The transaction function receives a driver-agnostic `Tx`: ```go type Tx interface { Exec(ctx context.Context, query string, args ...any) (Result, error) Query(ctx context.Context, query string, args ...any) (Rows, error) QueryRow(ctx context.Context, query string, args ...any) Row Commit(ctx context.Context) error Rollback(ctx context.Context) error } ``` Queries are passed to the underlying driver verbatim, so use your database's placeholder style: ```go // Postgres / CockroachDB (*pgxpool.Pool): $1, $2, ... placeholders _, err := tx.Exec(ctx, "INSERT INTO orders(item, quantity) VALUES ($1, $2)", item, quantity) // SQLite (*sql.DB): ? placeholders _, err := tx.Exec(ctx, "INSERT INTO orders(item, quantity) VALUES (?, ?)", item, quantity) ``` Do not call `Commit` or `Rollback` yourself inside `RunAsTransaction`: DBOS commits the transaction if `fn` returns successfully and rolls it back if `fn` returns an error. --- ## DBOS Context & Client A DBOS Context is at the center of a DBOS-enabled application. Use it to register [workflows](../tutorials/workflow-tutorial.md), [queues](../tutorials/queue-tutorial.md) and perform [workflow management](../tutorials/workflow-management.md) tasks. DBOS defines two interfaces: - [`Client`](#client) owns every DBOS operation that only needs a connection to the system database, for instance listing workflows. Create a standalone client with [`NewClient`](#newclient) to perform these operations from outside a DBOS application. - [`Context`](#context) **extends `Client`** and adds what requires a DBOS runtime, like running workflows and steps. Launching `Context` starts all background resources a DBOS process needs, like the queue runner. Create one with [`NewContext`](#newcontext). It also extends Go's [`context.Context`](https://pkg.go.dev/context#Context) and carries essential state across workflow execution: workflows and steps receive a new `Context` spun out of the root `Context` you manage, and a `Context` can be used to set [workflow timeouts](../tutorials/workflow-tutorial.md#workflow-timeouts). Because every `Context` **is** a `Client`, anywhere a `Client` is accepted you can pass either a standalone client or a (launched or unlaunched) `Context`. ### Client ```go type Client interface { context.Context // Workflow operations Enqueue(_ Client, queueName string, workflowName string, input any, opts ...EnqueueOption) (WorkflowHandle[any], error) Send(_ Client, destinationID string, message any, topic string, opts ...SendOption) error SendBulk(_ Client, messages []SendMessage, opts ...SendOption) error GetEvent(_ Client, targetWorkflowID string, key string, timeout time.Duration) (any, error) ReadStream(_ Client, workflowID string, key string, opts ...ReadStreamOption) ([]any, bool, error) ReadStreamAsync(_ Client, workflowID string, key string) (<-chan StreamValue[any], error) // Workflow management RetrieveWorkflow(_ Client, workflowID string) (WorkflowHandle[any], error) CancelWorkflow(_ Client, workflowID string, opts ...CancelWorkflowOption) error CancelWorkflows(_ Client, workflowIDs []string, opts ...CancelWorkflowOption) error SetWorkflowAttributes(_ Client, workflowID string, attributes map[string]any) error SetWorkflowDelay(_ Client, workflowID string, opts ...SetWorkflowDelayOption) error ResumeWorkflow(_ Client, workflowID string, opts ...ResumeWorkflowOption) (WorkflowHandle[any], error) ResumeWorkflows(_ Client, workflowIDs []string, opts ...ResumeWorkflowOption) ([]WorkflowHandle[any], error) ForkWorkflow(_ Client, input ForkWorkflowInput) (WorkflowHandle[any], error) ForkWorkflows(_ Client, input ForkWorkflowsInput) ([]WorkflowHandle[any], error) ListWorkflows(_ Client, opts ...ListWorkflowsOption) ([]WorkflowStatus, error) GetWorkflowSteps(_ Client, workflowID string, opts ...GetWorkflowStepsOption) ([]StepInfo, error) GetWorkflowAggregates(_ Client, input GetWorkflowAggregatesInput) ([]WorkflowAggregateRow, error) GetStepAggregates(_ Client, input GetStepAggregatesInput) ([]StepAggregateRow, error) DeleteWorkflows(_ Client, workflowIDs []string, opts ...DeleteWorkflowOption) error // Queue management RegisterQueue(_ Client, name string, options ...QueueOption) (Queue, error) RetrieveQueue(_ Client, name string) (Queue, error) ListQueues(_ Client, opts ...ListQueuesOption) ([]Queue, error) DeleteQueue(_ Client, name string) error // Schedule management CreateSchedule(_ Client, spec ScheduleSpec) error ApplySchedules(_ Client, schedules []ScheduleSpec) error PauseSchedule(_ Client, scheduleName string) error ResumeSchedule(_ Client, scheduleName string) error DeleteSchedule(_ Client, scheduleName string) error GetSchedule(_ Client, scheduleName string) (WorkflowSchedule, error) ListSchedules(_ Client, opts ...ListSchedulesOption) ([]WorkflowSchedule, error) BackfillSchedule(_ Client, scheduleName string, start, end time.Time) ([]string, error) TriggerSchedule(_ Client, scheduleName string) (WorkflowHandle[any], error) // Application management ListApplicationVersions(_ Client) ([]VersionInfo, error) GetLatestApplicationVersion(_ Client) (VersionInfo, error) SetLatestApplicationVersion(_ Client, versionName string) error RenameApplication(_ Client, input RenameApplicationInput) (ApplicationRowCounts, error) Shutdown(_ Client, timeout time.Duration) error } ``` ### Context ```go type Context interface { Client // Context Lifecycle Launch() error // Workflow operations RunAsStep(_ Context, fn StepFunc, opts ...StepOption) (any, error) RunAsTransaction(_ Context, ds *DataSource, fn TxnFunc, opts ...StepOption) (any, error) RunWorkflow(_ Context, fn WorkflowFunc, input any, opts ...WorkflowOption) (WorkflowHandle[any], error) Go(_ Context, fn StepFunc, opts ...StepOption) (<-chan StepOutcome[any], error) Select(_ Context, channels []<-chan StepOutcome[any]) (any, error) Recv(_ Context, topic string, timeout time.Duration) (any, error) SetEvent(_ Context, key string, message any, opts ...SetEventOption) error WriteStream(_ Context, key string, value any, opts ...WriteStreamOption) error CloseStream(_ Context, key string) error Sleep(_ Context, duration time.Duration) (time.Duration, error) Patch(_ Context, patchName string) (bool, error) DeprecatePatch(_ Context, patchName string) error GetWorkflowID() (string, error) // Only available within workflows GetStepID() (int, error) // Only available within workflows // Registration ListRegisteredWorkflows(_ Context) []WorkflowRegistryEntry ListenQueues(_ Context, names ...string) ListenedQueues(_ Context) []string // Accessors GetApplicationVersion() string GetExecutorID() string GetApplicationID() string // Context management From(_ Context, ctx context.Context) Context WithoutCancel(_ Context) Context WithTimeout(_ Context, timeout time.Duration) (Context, context.CancelFunc) WithValue(key, val any) Context WithCancel() (Context, context.CancelFunc) WithCancelCause() (Context, context.CancelCauseFunc) // Alert handling SetAlertHandler(handler AlertHandler) // Must be called before Launch } ``` ### Who can do what - **Every method in the `Client` interface** works from a standalone client and from any `Context`, before or after `Launch()`. These operations talk directly to the system database. One exception: on the `Context` a step body receives, mutating operations return an error — see [Calling DBOS operations from steps](./workflows-steps.md#calling-dbos-operations-from-steps). When called from workflow code, these operations are checkpointed as steps (`DBOS.*` step names in the workflow's step list). - **`Launch()`, workflow registration ([`RegisterWorkflow`](./workflows-steps.md#registerworkflow)), and `SetAlertHandler`** require a `Context`; registration and `SetAlertHandler` must happen before `Launch()`. - **Starting workflows (with [`RunWorkflow`](./workflows-steps.md#runworkflow))** requires a launched `Context`. - **Workflow-scope methods** ([`RunAsStep`](./workflows-steps.md#runasstep), [`Recv`](./methods.md#recv), [`SetEvent`](./methods.md#setevent), [`WriteStream`](./methods.md#writestream), [`CloseStream`](./methods.md#closestream), [`Sleep`](./methods.md#sleep), [`GetWorkflowID`](./methods.md#getworkflowid), [`GetStepID`](./methods.md#getstepid), [`Patch`](./workflows-steps.md#patch)) can only be called on the `Context` a workflow function receives. In practice, never call the interface methods directly — call their mirror package-level functions instead. These are strongly typed (generic) and the interface methods are not. ### Lifecycle #### NewContext You can create a DBOS context using `NewContext`, which takes a [`Config`](./configuration.md) object where `AppName` and one of `DatabaseURL`, `SystemDBPool`, or `SQLiteSystemDB` are mandatory. ```go func NewContext(ctx context.Context, inputConfig Config) (Context, error) ``` For example: ```go dbosContext, err := dbos.NewContext(context.Background(), dbos.Config{ AppName: "dbos-starter", ApplicationVersion: "0.1.0", DatabaseURL: os.Getenv("DBOS_SYSTEM_DATABASE_URL"), }) if err != nil { panic(err) } ``` `NewContext` connects to your system database and runs any pending schema migrations — see [System database startup](./configuration.md#system-database-startup). The newly created Context must be launched with `Launch()` before running workflows and should be shut down with `Shutdown()` at program termination. Before launch, a `Context` can already be used for every [`Client`](#client) operation. `AppName` is the application on whose behalf this context acts. Always set `AppName` if multiple applications [share a system database](../../explanations/sharing-a-system-database.md). #### Launch ```go dbos.Launch(ctx Context) error ``` Launch the following resources managed by a `Context`: - A [system database connection pool](../../explanations/system-tables.md) - A [workflow scheduler](../tutorials/scheduled-workflows.md) - A [workflow queue runner](../tutorials/queue-tutorial.md) - (Optionally) a Conductor connection In addition, `Launch()` may perform [workflow recovery](../../architecture.md#how-workflow-recovery-works). `Launch()` should be called by your program during startup before running any workflows. `Launch()` must be called on the `Context` returned by `NewContext`. Calling it on a context derived from that one (with [`WithTimeout`](#withtimeout), [`WithValue`](#withvalue), [`From`](#from), and so on) returns an error. #### NewClient ```go func NewClient(ctx context.Context, config ClientConfig) (Client, error) ``` Create a standalone `Client`, to interact with a DBOS application from external code — a process that registers no workflows and never calls `Launch()`. **Parameters:** - `ctx`: A context for initialization operations - `config`: A `ClientConfig` object with connection and application settings ```go type ClientConfig struct { DatabaseURL string // Connection string to your system database. May be a PostgreSQL (postgres://...) or SQLite (sqlite:...) URL. Exactly one of DatabaseURL, SystemDBPool, or SQLiteSystemDB is required. AppName string // The application this client acts on behalf of (optional) SystemDBPool *pgxpool.Pool // A custom Postgres/CockroachDB connection pool. Optional; takes precedence over DatabaseURL. Mutually exclusive with SQLiteSystemDB. SQLiteSystemDB *sql.DB // A custom SQLite handle (e.g. from modernc.org/sqlite). Optional; takes precedence over DatabaseURL. Mutually exclusive with SystemDBPool. DatabaseSchema string // Database schema name (defaults to "dbos") Logger *slog.Logger // Optional custom logger Serializer Serializer[any] // Optional custom serializer (defaults to JSON). See the serialization reference. SystemDBStartupTimeout time.Duration // Maximum time for system database connection and schema verification (default: 2 minutes) } ``` `NewClient` connects to the system database and starts a notification listener (or a poller on backends without listen/notify support), so every client operation — including blocking ones like `GetEvent` — works without launching the DBOS runtime. Startup follows the same rules as `NewContext`, including `SystemDBStartupTimeout` — see [System database startup](./configuration.md#system-database-startup). Unlike `NewContext`, a client never creates or migrates the system database: it verifies the schema instead, and `NewClient` fails if the system database is missing or behind the version this build of DBOS requires. Migrate it from the application (`NewContext`) or out of band with the `dbos migrate` [CLI command](./cli.md). Like `NewContext`, using a SQLite system database requires registering the SQLite driver with a blank import — see [Using SQLite](./configuration.md#using-sqlite). Because workflows are not registered with a client, operations that take a workflow function reference on a `Context` take a workflow **name** (string) from a client — for example [`Enqueue`](./methods.md#enqueue), or the `WorkflowName` field of [`ScheduleSpec`](./methods.md#schedulespec). `AppName` is the application on whose behalf this client acts: workflows the client enqueues and queues and schedules it registers are owned by that application, and the client's listing operations default to that application's rows. A client with no `AppName` sees every application's rows, but everything it creates is owned by no application. Always set `AppName` if multiple applications [share a system database](../../explanations/sharing-a-system-database.md). **Example syntax:** ```go config := dbos.ClientConfig{ DatabaseURL: os.Getenv("DBOS_SYSTEM_DATABASE_URL"), } client, err := dbos.NewClient(context.Background(), config) if err != nil { log.Fatal(err) } defer dbos.Shutdown(client, 5*time.Second) ``` #### Shutdown ```go dbos.Shutdown(c Client, timeout time.Duration) error ``` Gracefully shut down a `Context` or a standalone `Client`, waiting for resources to stop and cleaning up. Returns a non-nil error if the timeout expired before all resources stopped. When you shut down a `Context`, the underlying `context.Context` will be cancelled, which signals all DBOS resources they should stop executing, including workflows and steps. When you shut down a standalone `Client`, its system database connection pool and notification listener are released. Like `Launch()`, `Shutdown()` must be called on the `Context` returned by `NewContext` (or the `Client` returned by `NewClient`); calling it on a derived context returns an error. **Parameters:** - **timeout**: The time to wait for DBOS resources to gracefully terminate. ### Context management #### WithTimeout ```go func WithTimeout(ctx Context, timeout time.Duration) (Context, context.CancelFunc) ``` `WithTimeout` returns a copy of the DBOS context with a timeout. The returned context will be canceled after the specified duration. See [workflow timeouts](../tutorials/workflow-tutorial.md#workflow-timeouts) for usage. #### WithoutCancel ```go func WithoutCancel(ctx Context) Context ``` `WithoutCancel` returns a copy of the DBOS context that is not canceled when the parent context is canceled. This is useful to detach child workflows from their parent's timeout. #### WithCancel ```go func WithCancel(ctx Context) (Context, context.CancelFunc) ``` `WithCancel` returns a copy of the DBOS context that can be manually canceled, along with a `CancelFunc`. Cancelling propagates to workflows and steps running under the returned context. You must call the returned `CancelFunc` (e.g. with `defer`) when the derived context is no longer needed to release its resources. #### WithCancelCause ```go func WithCancelCause(ctx Context) (Context, context.CancelCauseFunc) ``` `WithCancelCause` behaves like [`WithCancel`](#withcancel) but returns a [`context.CancelCauseFunc`](https://pkg.go.dev/context#CancelCauseFunc), letting you supply an error describing why the context was canceled. The cause can later be retrieved with [`context.Cause`](https://pkg.go.dev/context#Cause). #### WithValue ```go func WithValue(ctx Context, key, val any) Context ``` `WithValue` returns a copy of the DBOS context with the given key-value pair, like [`context.WithValue`](https://pkg.go.dev/context#WithValue) but preserving DBOS context capabilities. #### From ```go func From(dbosCtx Context, ctx context.Context) Context ``` `From` returns a copy of `dbosCtx` whose embedded `context.Context` is replaced by `ctx`. The returned Context takes its deadline, cancellation, and values entirely from `ctx`; `dbosCtx` contributes only the DBOS runtime state (system database, registries, configuration, logger). `ctx` must descend from a `context.Context` provided by DBOS (e.g., the first argument of a workflow or step function), because DBOS metadata such as the current workflow state travels in context values. Returns nil if either argument is nil. ### Context metadata #### GetApplicationVersion ```go func GetApplicationVersion(ctx Context) string ``` `GetApplicationVersion` returns the application version for this context. #### GetExecutorID ```go func GetExecutorID(ctx Context) string ``` `GetExecutorID` returns the executor ID for this context. #### GetApplicationID ```go func GetApplicationID(ctx Context) string ``` `GetApplicationID` returns the application ID for this context (set in DBOS Cloud; empty otherwise). #### ListRegisteredWorkflows ```go func ListRegisteredWorkflows(ctx Context) []WorkflowRegistryEntry ``` ```go type WorkflowRegistryEntry struct { MaxRetries int // Maximum recovery attempts before dead-lettering (set via WithMaxRecoveryAttempts); not step retries Name string FQN string // Fully qualified name of the workflow function. For configured instances, qualified with the config name. ClassName string // Receiver type name for configured instance workflows ConfigName string // Config name for configured instance workflows } ``` `ListRegisteredWorkflows` Lists the context's workflow registry. --- ## DBOS Methods & Variables This page documents the package-level functions for interacting with workflows: communication, streams, management, schedules, and application versions. The first parameter of each function tells you who can call it — see [Who can do what](./dbos-context.md#who-can-do-what): - A function taking a **`Client`** accepts a [standalone client](./dbos-context.md#newclient) or any [`Context`](./dbos-context.md#context) (launched or not). - A function taking a **`Context`** requires a DBOS context; some (like [`Recv`](#recv) or [`SetEvent`](#setevent)) can only be called within a workflow. ### Enqueueing Workflows #### Enqueue ```go func Enqueue[R any, P any]( ctx Client, queueName string, workflowName string, input P, opts ...EnqueueOption ) (WorkflowHandle[R], error) ``` Enqueue a workflow for processing and return a [WorkflowHandle](./workflows-steps.md#workflowhandle) to it, similar to [RunWorkflow with the WithQueue option](./workflows-steps.md#withqueue). The workflow is identified by **name** rather than by function reference, so the enqueueing process does not need to have the workflow registered — this is how you enqueue workflows from a [standalone client](./dbos-context.md#newclient). Required parameters: * `ctx`: The DBOS client or context * `queueName`: The name of the [queue](./queues.md) on which to enqueue the workflow * `workflowName`: The name of the workflow function being enqueued * `input`: The input to pass to the workflow Optional configuration via `EnqueueOption`, documented below. :::tip Cross-Language Enqueue To enqueue a workflow on a target application written in another language, pass a [`PortableWorkflowArgs`](#portableworkflowargs) as the input. This automatically uses portable JSON serialization. See [Cross-Language Interaction](../../explanations/portable-workflows.md) for details. ::: **Example syntax:** ```go type ProcessInput struct { TaskID string Data string } type ProcessOutput struct { Result string Status string } handle, err := dbos.Enqueue[ProcessOutput]( client, "process_queue", "ProcessWorkflow", ProcessInput{TaskID: "task-123", Data: "data"}, dbos.WithEnqueueTimeout(30 * time.Minute), dbos.WithEnqueuePriority(5), ) if err != nil { log.Fatal(err) } result, err := handle.GetResult() if err != nil { log.Printf("Workflow failed: %v", err) } else { log.Printf("Result: %+v", result) } ``` ##### WithEnqueueWorkflowID ```go func WithEnqueueWorkflowID(id string) EnqueueOption ``` The unique ID for the enqueued workflow. If left undefined, DBOS will generate a [UUID](https://en.wikipedia.org/wiki/Universally_unique_identifier). Please see [Workflow IDs and Idempotency](../tutorials/workflow-tutorial.md#workflow-ids-and-idempotency) for more information. ##### WithEnqueueApplicationVersion ```go func WithEnqueueApplicationVersion(version string) EnqueueOption ``` The version of your application that should process this workflow. If left undefined, a `Context` enqueueing to its own application uses its current application version; otherwise (from a [standalone client](./dbos-context.md#newclient), or when enqueueing to another application with [`WithEnqueueApplicationName`](#withenqueueapplicationname)) the version is left unset and the workflow is dequeued at the owning application's latest registered version. ##### WithEnqueueApplicationName ```go func WithEnqueueApplicationName(name string) EnqueueOption ``` The application that owns the enqueued workflow, which dequeues and runs it. Defaults to the enqueueing context's own application. Use this to run another application's workflows when multiple applications [share a system database](../../explanations/sharing-a-system-database.md); if the applications are written in different languages, also pass a [`PortableWorkflowArgs`](#portableworkflowargs) as the input so the target application can read the arguments. ##### WithEnqueueTimeout ```go func WithEnqueueTimeout(timeout time.Duration) EnqueueOption ``` Set a timeout for the enqueued workflow. When the timeout expires, the workflow **and all its children** are cancelled (except if the child's context has been made uncancellable using [`WithoutCancel`](./dbos-context.md#withoutcancel)). The timeout does not begin until the workflow is dequeued and starts execution. ##### WithEnqueueDeduplicationID ```go func WithEnqueueDeduplicationID(id string) EnqueueOption ``` At any given time, only one workflow with a specific deduplication ID can be enqueued in the specified queue. If a workflow with a deduplication ID is currently enqueued or actively executing (status `ENQUEUED` or `PENDING`), subsequent workflow enqueue attempts with the same deduplication ID in the same queue will fail. This behavior can be changed with [`WithEnqueueDeduplicationPolicy`](#withenqueuededuplicationpolicy). ##### WithEnqueueDeduplicationPolicy ```go func WithEnqueueDeduplicationPolicy(policy DeduplicationPolicy) EnqueueOption ``` Set how a colliding deduplication ID is handled. Requires [`WithEnqueueDeduplicationID`](#withenqueuededuplicationid). With the default `DeduplicationPolicyReject`, a colliding enqueue fails with a `ErrorCodeQueueDeduplicated` error; with `DeduplicationPolicyReturnExisting`, it instead returns a handle to the existing workflow. See [`WithDeduplicationPolicy`](./workflows-steps.md#withdeduplicationpolicy). ##### WithEnqueuePriority ```go func WithEnqueuePriority(priority uint) EnqueueOption ``` The priority of the enqueued workflow in the specified queue. Workflows with the same priority are dequeued in **FIFO (first in, first out)** order. Priority values can range from `1` to `2,147,483,647`, where **a low number indicates a higher priority**. Workflows without assigned priorities have the highest priority and are dequeued before workflows with assigned priorities. ##### WithEnqueueClassName ```go func WithEnqueueClassName(className string) EnqueueOption ``` The class/namespace name for the target workflow. Required when enqueueing to Python, TypeScript, or Java targets, which dispatch workflows by (class_name, workflow_name) pair. ##### WithEnqueueConfigName ```go func WithEnqueueConfigName(configName string) EnqueueOption ``` The config/instance name for the target workflow. Required when enqueueing to a workflow registered on a configured instance: a Go workflow registered with [`WithInstance`](./workflows-steps.md#withinstance), or a Python, TypeScript, or Java class instance workflow (e.g., Python's [`DBOSConfiguredInstance`](../../python/tutorials/classes.md), TypeScript's [`ConfiguredInstance`](../../typescript/tutorials/instantiated-objects.md)). The value must match the instance name used by the target application. ##### WithEnqueueDelay ```go func WithEnqueueDelay(delay time.Duration) EnqueueOption ``` Delay execution of the enqueued workflow by the specified duration. The workflow is initially placed in `DELAYED` status and transitions to `ENQUEUED` after the delay expires. The delay can later be updated via [`SetWorkflowDelay`](#setworkflowdelay). ##### WithEnqueueQueuePartitionKey ```go func WithEnqueueQueuePartitionKey(partitionKey string) EnqueueOption ``` The partition key to enqueue under when the target queue is a [partitioned queue](../tutorials/queue-tutorial.md#partitioning-queues). Required if and only if the queue is partitioned. The queue's partition limits apply to each partition separately. ##### WithEnqueueAttributes ```go func WithEnqueueAttributes(attributes map[string]any) EnqueueOption ``` Attach custom key-value [attributes](./workflows-steps.md#withworkflowattributes) to the enqueued workflow. Attributes are recorded in the workflow status at creation, must be JSON-serializable, and can be searched with [`WithFilterAttributes`](#withfilterattributes) on Postgres. ##### WithEnqueueAuthenticatedUser ```go func WithEnqueueAuthenticatedUser(user string) EnqueueOption ``` Associate the enqueued workflow with a user name. ##### WithEnqueueAssumedRole ```go func WithEnqueueAssumedRole(role string) EnqueueOption ``` Set the assumed role for the enqueued workflow. ##### WithEnqueueAuthenticatedRoles ```go func WithEnqueueAuthenticatedRoles(roles ...string) EnqueueOption ``` Set the authenticated roles for the enqueued workflow. ##### WithEnqueueTransaction ```go func WithEnqueueTransaction(tx any) EnqueueOption ``` Enqueue the workflow on a transaction you own instead of one DBOS opens, so the enqueue commits **atomically** with your own database writes: either both are committed or both are rolled back. `tx` must be a `pgx.Tx`, a `*sql.Tx`, or a [`Tx`](./datasources.md#the-tx-interface), and must run against your DBOS system database. You own the transaction: `Enqueue` does not begin, commit, or roll back it, and does not retry on database errors. The returned [`WorkflowHandle`](./workflows-steps.md#workflowhandle) is created immediately, but the workflow is not enqueued until you commit, so do not call `GetResult` on the handle until after the transaction commits. If the enqueue fails, the transaction is left in an aborted state: roll it back rather than retrying the call on it. This option is available from a [standalone client](./dbos-context.md#newclient) or a context outside a workflow. It cannot be used inside a workflow, where an enqueue is checkpointed as a step, nor together with [`WithEnqueueDeduplicationPolicy`](#withenqueuededuplicationpolicy) set to `DeduplicationPolicyReturnExisting`, which retries the insert on collision and so would abort your transaction. ```go tx, err := pool.Begin(ctx) if err != nil { return err } defer tx.Rollback(ctx) // Perform your own writes on tx here, in the same transaction... _, err = tx.Exec(ctx, "INSERT INTO orders (id) VALUES ($1)", orderID) if err != nil { return err } handle, err := dbos.Enqueue[ProcessOutput](client, "process_queue", "ProcessWorkflow", orderID, dbos.WithEnqueueTransaction(tx)) if err != nil { return err } // Until this commits, the workflow does not exist. If you roll back instead, it never does. if err := tx.Commit(ctx); err != nil { return err } result, err := handle.GetResult() ``` ### Workflow Communication #### GetEvent ```go func GetEvent[R any](ctx Client, targetWorkflowID, key string, timeout time.Duration) (R, error) ``` Retrieve the latest value of an event published by the workflow identified by `targetWorkflowID` to the key `key`. If the event does not yet exist, wait for it to be published, returning an error if the wait times out. **Parameters:** - **ctx**: The DBOS client or context. - **targetWorkflowID**: The identifier of the workflow whose events to retrieve. - **key**: The key of the event to retrieve. - **timeout**: A timeout. If the wait times out, return an error. #### SetEvent ```go func SetEvent[P any](ctx Context, key string, message P, opts ...SetEventOption) error ``` Create and associate with this workflow an event with key `key` and value `value`. If the event already exists, update its value. May only be called from within a workflow or step. Writes from a workflow are exactly-once; writes from a step are at-least-once, attributed to the enclosing step. **Parameters:** - **ctx**: The DBOS context. - **key**: The key of the event. - **message**: The value of the event. Must be serializable. - **opts**: Optional [SetEventOption](#withportablesetevent) functions. #### Send ```go func Send[P any](ctx Client, destinationID string, message P, topic string, opts ...SendOption) error ``` Send a message to the workflow identified by `destinationID`. Messages can optionally be associated with a topic. **Parameters:** - **ctx**: The DBOS client or context. - **destinationID**: The workflow to which to send the message. - **message**: The message to send. Must be serializable. - **topic**: A topic with which to associate the message. Messages are enqueued per-topic on the receiver. - **opts**: Optional `SendOption` functions ([`WithIdempotencyKey`](#withidempotencykey), [`WithSendTransaction`](#withsendtransaction), [`WithPortableSend`](#withportablesend)). ##### WithIdempotencyKey ```go func WithIdempotencyKey(key string) SendOption ``` Make a `Send` deliver at most once. The key is combined with the destination workflow ID to form the message's primary key, so retrying a `Send` with the same key (after a crash, timeout, or network failure) inserts the message only once. Keys are scoped per destination. Without a key, every `Send` delivers a new message. This option is not valid on [`SendBulk`](#sendbulk): set the `IdempotencyKey` field of each [`SendMessage`](#sendmessage) instead. ```go err := dbos.Send(ctx, destinationID, payload, "payments", dbos.WithIdempotencyKey("payment-123")) ``` ##### WithSendTransaction ```go func WithSendTransaction(tx any) SendOption ``` Send the message on a transaction you own instead of one DBOS opens, so the message commits **atomically** with your own database writes. `tx` must be a `pgx.Tx`, a `*sql.Tx`, or a [`Tx`](./datasources.md#the-tx-interface), and must run against your DBOS system database. You own the transaction: `Send` does not begin, commit, or roll back it, and does not retry on database errors. The message is not visible to the destination workflow until you commit. If the send fails, the transaction is left in an aborted state: roll it back rather than retrying the call on it. This option is available from a [standalone client](./dbos-context.md#newclient) or a context outside a workflow. It cannot be used inside a workflow, where a send is checkpointed as a step. It applies to [`SendBulk`](#sendbulk) the same way, committing the whole batch with your writes. ```go tx, err := pool.Begin(ctx) if err != nil { return err } defer tx.Rollback(ctx) _, err = tx.Exec(ctx, "INSERT INTO orders (id) VALUES ($1)", orderID) if err != nil { return err } err = dbos.Send(client, workflowID, orderID, "orders", dbos.WithSendTransaction(tx)) if err != nil { return err } return tx.Commit(ctx) ``` #### SendBulk ```go func SendBulk(ctx Client, messages []SendMessage, opts ...SendOption) error ``` Send many messages in a single transaction. Each message carries its own destination, so a batch may target many workflows. The batch is atomic: if any message cannot be delivered (for example, its destination workflow does not exist), no message is sent. A batch may carry at most `dbos.MaxSendBulkMessages` (10,000) messages. Inside a workflow, the whole batch is checkpointed as one step, so it is delivered exactly once. **Parameters:** - **ctx**: The DBOS client or context. - **messages**: The messages to send, as [`SendMessage`](#sendmessage) values. Two messages in the same call may not share an idempotency key. - **opts**: Optional `SendOption` functions applied to the whole batch ([`WithSendTransaction`](#withsendtransaction), [`WithPortableSend`](#withportablesend)). [`WithIdempotencyKey`](#withidempotencykey) is per-message and is rejected here: set `SendMessage.IdempotencyKey` instead. ```go err := dbos.SendBulk(ctx, []dbos.SendMessage{ {DestinationID: orderWorkflowID, Message: "confirmed", Topic: "orders"}, {DestinationID: inventoryWorkflowID, Message: order, Topic: "reserve", IdempotencyKey: "reserve-42"}, }) ``` ##### SendMessage ```go type SendMessage struct { DestinationID string // The workflow to which to send the message Message any // The message to send. Must be serializable. Topic string // A topic with which to associate the message. Messages are enqueued per-topic on the receiver. IdempotencyKey string // If set, the message is delivered at most once per destination no matter how many times it is submitted with this key. } ``` One entry in a [`SendBulk`](#sendbulk) batch. An empty `Topic` sends the message without a topic, as `Send` does. `IdempotencyKey` behaves like [`WithIdempotencyKey`](#withidempotencykey): it is combined with `DestinationID` to form the message's primary key, so a retried batch inserts the message only once. #### Recv ```go func Recv[R any](ctx Context, topic string, timeout time.Duration) (R, error) ``` Receive and return a message sent to this workflow. Can only be called from within a workflow. Messages are dequeued first-in, first-out from a queue associated with the topic. Calls to `recv` wait for the next message in the queue, returning an error if the wait times out. **Parameters:** - **ctx**: The DBOS context. - **topic**: A topic queue on which to wait. - **timeout**: A `time.Duration` to wait. If the wait times out, return an error. ### Streams Workflows can stream data to clients in real time. Streams are durable, append-only, and ordered by offset. See the [streaming tutorial](../tutorials/workflow-communication.md#workflow-streaming) for usage examples. #### WriteStream ```go func WriteStream[P any](ctx Context, key string, value P, opts ...WriteStreamOption) error ``` Write a value to a durable stream. May only be called from within a workflow or step. Writes from a workflow are exactly-once; writes from a step are at-least-once. **Parameters:** - **ctx**: The DBOS context. - **key**: The stream key. A workflow can have multiple streams, each identified by a unique key. - **value**: The value to write. Must be serializable (json-encodable). - **opts**: Optional [WriteStreamOption](#withportablewritestream) functions. #### CloseStream ```go func CloseStream(ctx Context, key string) error ``` Close a durable stream. May only be called from within a workflow (not from inside a step). After closing, no more values can be written to the stream. Streams are also automatically closed when the workflow terminates. **Parameters:** - **ctx**: The DBOS context. - **key**: The stream key to close. #### ReadStream ```go func ReadStream[R any](ctx Client, workflowID string, key string, opts ...ReadStreamOption) ([]R, bool, error) ``` Read all values from a durable stream. By default, blocks until the stream is closed or the workflow becomes inactive (status is not `PENDING` or `ENQUEUED`). Pass [`WithReadStreamSnapshot`](#withreadstreamsnapshot) to instead return immediately once all currently-available values have been drained. **Parameters:** - **ctx**: The DBOS client or context. - **workflowID**: The ID of the workflow whose stream to read. - **key**: The stream key to read. - **opts**: Optional [ReadStreamOption](#withreadstreamsnapshot) functions. **Returns:** - The values read from the stream. - Whether the stream is closed. - Any error that occurred. #### ReadStreamAsync ```go func ReadStreamAsync[R any](ctx Client, workflowID string, key string) (<-chan StreamValue[R], error) ``` Read values from a durable stream asynchronously. Returns immediately with a channel that receives values as they are written to the stream. The channel is closed when the stream is closed or an error occurs. **Parameters:** - **ctx**: The DBOS client or context. - **workflowID**: The ID of the workflow whose stream to read. - **key**: The stream key to read. **Returns:** - A receive-only channel of [`StreamValue[R]`](#streamvalue). - Any error that occurred during setup. #### StreamValue ```go type StreamValue[R any] struct { Value R // The stream value (zero value if error/closed) Err error // Error if one occurred (nil otherwise) Closed bool // Whether the stream is closed } ``` `StreamValue` holds a value, error, or closed status from an async stream read operation. When reading from the channel returned by `ReadStreamAsync`, check `Err` and `Closed` before using `Value`. ### Sleep #### Sleep ```go func Sleep(ctx Context, duration time.Duration) (time.Duration, error) ``` Sleep for the given duration. May only be called from within a workflow. This sleep is durable—it records its intended wake-up time in the database so if it is interrupted and recovers, it still wakes up at the intended time. If the workflow's context is cancelled (e.g., its [durable timeout](../tutorials/workflow-tutorial.md#workflow-timeouts) expires), the sleep wakes immediately, returning the elapsed duration and the context's error. **Parameters:** - **ctx**: The DBOS context. - **duration**: The duration to sleep. ### Workflow Management Methods #### RetrieveWorkflow ```go func RetrieveWorkflow[R any](ctx Client, workflowID string) (WorkflowHandle[R], error) ``` Retrieve the [handle](./workflows-steps.md#workflowhandle) of a workflow. The generic `RetrieveWorkflow` returns a typed handle whose `GetResult` decodes the workflow output into type `R`. **Parameters**: - **ctx**: The DBOS client or context. - **workflowID**: The ID of the workflow whose handle to retrieve. #### ListWorkflows ```go func ListWorkflows(ctx Client, opts ...ListWorkflowsOption) ([]WorkflowStatus, error) ``` Retrieve a list of [`WorkflowStatus`](#workflow-status) of all workflows matching specified criteria. If multiple applications [share a system database](../../explanations/sharing-a-system-database.md), only workflows owned by the calling context's application (plus workflows owned by no application) are listed by default; use [`WithFilterApplicationName`](#withfilterapplicationname) to list other applications' workflows. A [standalone client](./dbos-context.md#newclient) with no `AppName` lists every application's workflows. **Example usage:** ```go // List all successful workflows from the last 24 hours workflows, err := dbos.ListWorkflows(ctx, dbos.WithFilterStatus(dbos.WorkflowStatusSuccess), dbos.WithFilterCreatedAfter(time.Now().Add(-24*time.Hour)), dbos.WithFilterLimit(100)) if err != nil { log.Fatal(err) } // List workflows by specific IDs without loading input/output data workflows, err := dbos.ListWorkflows(ctx, dbos.WithFilterWorkflowIDs("workflow1", "workflow2"), dbos.WithFilterLoadInput(false), dbos.WithFilterLoadOutput(false)) if err != nil { log.Fatal(err) } ``` ##### WithFilterApplicationName ```go func WithFilterApplicationName(applicationName ...string) ListWorkflowsOption ``` List workflows owned by these applications (workflows owned by no application are always included). ##### WithFilterAppVersion ```go func WithFilterAppVersion(appVersion ...string) ListWorkflowsOption ``` Retrieve workflows tagged with any of these application versions. ##### WithFilterCreatedBefore ```go func WithFilterCreatedBefore(endTime time.Time) ListWorkflowsOption ``` Retrieve workflows started before this timestamp. ##### WithFilterAttributes ```go func WithFilterAttributes(attributes map[string]any) ListWorkflowsOption ``` Retrieve workflows whose [attributes](./workflows-steps.md#withworkflowattributes) contain all the given key-value pairs (JSONB containment). Requires a Postgres system database; listing fails with an error on SQLite. ##### WithFilterLimit ```go func WithFilterLimit(limit int) ListWorkflowsOption ``` Retrieve up to this many workflows. ##### WithFilterLoadInput ```go func WithFilterLoadInput(loadInput bool) ListWorkflowsOption ``` WithFilterLoadInput controls whether to load workflow input data (default: true on a launched `Context`, false on an unlaunched context or standalone client). ##### WithFilterLoadOutput ```go func WithFilterLoadOutput(loadOutput bool) ListWorkflowsOption ``` WithFilterLoadOutput controls whether to load workflow output data (default: true on a launched `Context`, false on an unlaunched context or standalone client). ##### WithFilterName ```go func WithFilterName(names ...string) ListWorkflowsOption ``` Filter workflows by the specified workflow function name. ##### WithFilterOffset ```go func WithFilterOffset(offset int) ListWorkflowsOption ``` Skip this many workflows from the results returned (for pagination). ##### WithFilterSortDesc ```go func WithFilterSortDesc() ListWorkflowsOption ``` Sort the results in descending order by workflow start time (ascending is the default). ##### WithFilterCreatedAfter ```go func WithFilterCreatedAfter(startTime time.Time) ListWorkflowsOption ``` Retrieve workflows started after this timestamp. ##### WithFilterStatus ```go func WithFilterStatus(status ...WorkflowStatusType) ListWorkflowsOption ``` Filter workflows by [status](#workflowstatustype). Multiple statuses can be specified. ##### WithFilterUser ```go func WithFilterUser(user ...string) ListWorkflowsOption ``` Filter workflows run by any of these authenticated users. ##### WithFilterWorkflowIDs ```go func WithFilterWorkflowIDs(workflowIDs ...string) ListWorkflowsOption ``` Filter workflows by specific workflow IDs. ##### WithFilterWorkflowIDPrefix ```go func WithFilterWorkflowIDPrefix(prefix ...string) ListWorkflowsOption ``` Filter workflows whose IDs start with any of the specified prefixes. ##### WithFilterQueuesOnly ```go func WithFilterQueuesOnly() ListWorkflowsOption ``` Return only workflows that are currently in a queue (queue name is not null, status is `ENQUEUED`, `PENDING`, or `DELAYED`). ##### WithFilterQueueName ```go func WithFilterQueueName(queueName ...string) ListWorkflowsOption ``` Filter workflows enqueued on any of these queues. ##### WithFilterExecutorIDs ```go func WithFilterExecutorIDs(executorIDs ...string) ListWorkflowsOption ``` Filter workflows by the executor IDs that ran them. ##### WithFilterForkedFrom ```go func WithFilterForkedFrom(forkedFrom ...string) ListWorkflowsOption ``` Filter workflows forked from any of these workflow IDs. ##### WithFilterParentWorkflowID ```go func WithFilterParentWorkflowID(parentWorkflowID ...string) ListWorkflowsOption ``` Filter child workflows spawned by any of these parent workflow IDs. ##### WithFilterDeduplicationID ```go func WithFilterDeduplicationID(deduplicationID ...string) ListWorkflowsOption ``` Filter workflows by their queue deduplication IDs. ##### WithFilterCompletedAfter ```go func WithFilterCompletedAfter(completedAfter time.Time) ListWorkflowsOption ``` Retrieve workflows that reached a terminal state (`SUCCESS`, `ERROR`, or `CANCELLED`) at or after this timestamp. ##### WithFilterCompletedBefore ```go func WithFilterCompletedBefore(completedBefore time.Time) ListWorkflowsOption ``` Retrieve workflows that reached a terminal state (`SUCCESS`, `ERROR`, or `CANCELLED`) at or before this timestamp. ##### WithFilterDequeuedAfter ```go func WithFilterDequeuedAfter(dequeuedAfter time.Time) ListWorkflowsOption ``` Retrieve workflows that started executing at or after this timestamp. ##### WithFilterDequeuedBefore ```go func WithFilterDequeuedBefore(dequeuedBefore time.Time) ListWorkflowsOption ``` Retrieve workflows that started executing at or before this timestamp. ##### WithFilterWasForkedFrom ```go func WithFilterWasForkedFrom(wasForkedFrom bool) ListWorkflowsOption ``` Filter workflows by whether they have been forked from (true) or not (false). ##### WithFilterHasParent ```go func WithFilterHasParent(hasParent bool) ListWorkflowsOption ``` Filter workflows by whether they have a parent workflow (true) or not (false). ##### WithFilterIsDebounced ```go func WithFilterIsDebounced(isDebounced bool) ListWorkflowsOption ``` Filter workflows by whether they are pending [debounced](./queues.md#debouncer) invocations (true) or not (false). ##### WithFilterScheduleName ```go func WithFilterScheduleName(scheduleName ...string) ListWorkflowsOption ``` Filter workflows by the name(s) of the [schedule](#workflow-schedules) that enqueued them. Only workflows enqueued by a named schedule match. #### GetWorkflowSteps ```go func GetWorkflowSteps(ctx Client, workflowID string, opts ...GetWorkflowStepsOption) ([]StepInfo, error) ``` GetWorkflowSteps retrieves the execution steps of a workflow. This is a list of `StepInfo` objects, with the following structure: ```go type StepInfo struct { StepID int // The sequential ID of the step within the workflow StepName string // The name of the step function Output any // The output returned by the step (if any) Error error // The error returned by the step (if any) ChildWorkflowID string // If the step starts or retrieves the result of a workflow, its ID StartedAt time.Time // When the step execution started CompletedAt time.Time // When the step execution completed } ``` **Parameters:** - **ctx**: The DBOS client or context. - **workflowID**: The ID of the workflow whose steps to retrieve. - **opts**: Optional configuration, documented below. ##### WithStepsLoadOutput ```go func WithStepsLoadOutput(loadOutput bool) GetWorkflowStepsOption ``` Control whether to load step output data. When unset, output is loaded only if the DBOS context has been launched. ##### WithStepsLimit ```go func WithStepsLimit(limit int) GetWorkflowStepsOption ``` Limit the number of steps returned, ordered by step ID ascending. ##### WithStepsOffset ```go func WithStepsOffset(offset int) GetWorkflowStepsOption ``` Skip the given number of steps before returning results. Combine with `WithStepsLimit` to paginate through a workflow's steps. #### GetWorkflowAggregates ```go func GetWorkflowAggregates(ctx Client, input GetWorkflowAggregatesInput) ([]WorkflowAggregateRow, error) ``` Return aggregates of workflows grouped by one or more columns and/or by `created_at` time bucket. At least one `GroupBy*` flag must be set, or `TimeBucketSize` must be greater than zero. At least one `Select*` flag must be set. Filter fields narrow which workflows are aggregated before grouping. ```go type GetWorkflowAggregatesInput struct { GroupByStatus bool GroupByName bool GroupByQueueName bool GroupByExecutorID bool GroupByApplicationVersion bool GroupByApplicationName bool // Select* flags choose which aggregates to compute. At least one must be true. SelectCount bool SelectMinCreatedAt bool SelectMaxQueueWaitMs bool SelectMaxTotalLatencyMs bool // When non-zero, groups results by created_at time bucket of this size. TimeBucketSize time.Duration // Filters Status []WorkflowStatusType StartTime time.Time EndTime time.Time CompletedAfter time.Time CompletedBefore time.Time DequeuedAfter time.Time DequeuedBefore time.Time Name []string ApplicationVersion []string ExecutorID []string QueueName []string WorkflowIDPrefix []string WorkflowIDs []string AuthenticatedUser []string ForkedFrom []string ParentWorkflowID []string ApplicationName []string WasForkedFrom *bool HasParent *bool Attributes map[string]any } ``` The result is one [`WorkflowAggregateRow`](#workflowaggregaterow) per non-empty group. The `Group` map contains an entry per enabled grouping column (`"status"`, `"name"`, `"queue_name"`, `"executor_id"`, `"application_version"`, `"application_name"`, `"time_bucket"`). Like [`ListWorkflows`](#listworkflows), when the `ApplicationName` filter is unset, only workflows owned by the calling context's application (plus workflows owned by no application) are aggregated. `Count`, `MinCreatedAt`, `MaxQueueWaitMs`, and `MaxTotalLatencyMs` are populated only for the corresponding enabled `Select*` flag. **Parameters:** - **ctx**: The DBOS client or context. - **input**: A `GetWorkflowAggregatesInput` describing the grouping columns, aggregates, time bucket, and filters. **Example:** ```go rows, err := dbos.GetWorkflowAggregates(ctx, dbos.GetWorkflowAggregatesInput{ GroupByStatus: true, SelectCount: true, StartTime: time.Now().Add(-24 * time.Hour), }) if err != nil { log.Fatal(err) } for _, r := range rows { fmt.Printf("status=%s count=%d\n", *r.Group["status"], *r.Count) } ``` ##### WorkflowAggregateRow ```go type WorkflowAggregateRow struct { Group map[string]*string // One entry per enabled grouping column; nil values represent NULL Count *int64 // Number of workflows in this group (nil if SelectCount is false) MinCreatedAt *int64 // Earliest created_at in this group, as an epoch-ms timestamp (nil if SelectMinCreatedAt is false) MaxQueueWaitMs *int64 // Max time workflows in this group spent enqueued, in milliseconds (nil if SelectMaxQueueWaitMs is false) MaxTotalLatencyMs *int64 // Max total latency in this group, in milliseconds (nil if SelectMaxTotalLatencyMs is false) } ``` #### GetStepAggregates ```go func GetStepAggregates(ctx Client, input GetStepAggregatesInput) ([]StepAggregateRow, error) ``` Return aggregate counts and/or max durations of steps grouped by function name and/or status, optionally bucketed by `completed_at` time. At least one `GroupBy*` flag must be set, or `TimeBucketSize` must be greater than zero. At least one `Select*` flag must be set. Step status is derived from the step's recorded outcome: steps with no recorded error are `SUCCESS`, otherwise `ERROR`. ```go type GetStepAggregatesInput struct { GroupByFunctionName bool GroupByStatus bool SelectCount bool SelectMaxDurationMs bool // When non-zero, groups results by completed_at time bucket of this size. TimeBucketSize time.Duration // Filters Status []string FunctionName []string WorkflowIDPrefix []string CompletedAfter time.Time CompletedBefore time.Time ApplicationName []string } ``` The result is one [`StepAggregateRow`](#stepaggregaterow) per non-empty group. The `Group` map contains an entry per enabled grouping column (`"function_name"`, `"status"`, `"time_bucket"`). Like [`ListWorkflows`](#listworkflows), when the `ApplicationName` filter is unset, only steps owned by the calling context's application (plus steps owned by no application) are aggregated. `Count` and `MaxDurationMs` are populated only for the corresponding enabled `Select*` flag. **Parameters:** - **ctx**: The DBOS client or context. - **input**: A `GetStepAggregatesInput` describing the grouping columns, aggregates, time bucket, and filters. **Example:** ```go rows, err := dbos.GetStepAggregates(ctx, dbos.GetStepAggregatesInput{ GroupByFunctionName: true, SelectCount: true, SelectMaxDurationMs: true, CompletedAfter: time.Now().Add(-24 * time.Hour), }) if err != nil { log.Fatal(err) } for _, r := range rows { fmt.Printf("step=%s count=%d max_duration_ms=%d\n", *r.Group["function_name"], *r.Count, *r.MaxDurationMs) } ``` ##### StepAggregateRow ```go type StepAggregateRow struct { Group map[string]*string // One entry per enabled grouping column; nil values represent NULL Count *int64 // Number of steps in this group (nil if SelectCount is false) MaxDurationMs *int64 // Max step duration in this group (nil if SelectMaxDurationMs is false) } ``` #### CancelWorkflow ```go func CancelWorkflow(ctx Client, workflowID string, opts ...CancelWorkflowOption) error ``` Cancel a workflow. This sets its status to `CANCELLED` and removes it from its queue (if it is enqueued). A running execution is not interrupted mid-step: it stops at the start of its next durable operation (step, sleep, `Send`/`Recv`, child workflow, …), which returns an error matching `dbos.ErrWorkflowCancelled`. You can also cancel a running workflow directly by cancelling its context: start it under [`WithCancel`](./dbos-context.md#withcancel) (or [`WithTimeout`](./dbos-context.md#withtimeout)) and call the returned cancel function. Calling this cancel function will trigger a durable cancel and enable cooperative cancellation: an executing step receives the cancellation through its `context.Context` and can select on `ctx.Done()` to return early instead of running to completion. See [cancellation behavior](../tutorials/workflow-management.md#cancelling-workflows) for how cancellation interacts with executing steps, durable sleeps, and awaiting workflows. **Parameters:** - **ctx**: The DBOS client or context. - **workflowID**: The ID of the workflow to cancel. - **opts**: Optional configuration, documented below. ##### WithCancelChildren ```go func WithCancelChildren() CancelWorkflowOption ``` Also cancel all the workflow's child workflows, recursively. ```go err := dbos.CancelWorkflow(ctx, workflowID, dbos.WithCancelChildren()) ``` #### CancelWorkflows ```go func CancelWorkflows(ctx Client, workflowIDs []string, opts ...CancelWorkflowOption) error ``` Cancel multiple workflows in a single database round-trip. Each workflow that exists and is not already in a terminal state (`SUCCESS`, `ERROR`, `CANCELLED`) is moved to `CANCELLED` and removed from its queue. Unlike [`CancelWorkflow`](#cancelworkflow), this function does not return an error when some IDs are missing. Accepts the same options as [`CancelWorkflow`](#cancelworkflow) (e.g., [`WithCancelChildren`](#withcancelchildren)). **Parameters:** - **ctx**: The DBOS client or context. - **workflowIDs**: The IDs of the workflows to cancel. - **opts**: Optional configuration. #### ResumeWorkflow ```go func ResumeWorkflow[R any](ctx Client, workflowID string, opts ...ResumeWorkflowOption) (WorkflowHandle[R], error) ``` Resume a workflow. This immediately starts it from its last completed step. You can use this to resume workflows that are cancelled or have exceeded their maximum recovery attempts. You can also use this to start an enqueued workflow immediately, bypassing its queue. **Parameters:** - **ctx**: The DBOS client or context. - **workflowID**: The ID of the workflow to resume. - **opts**: Optional configuration, documented below. ##### WithResumeQueue ```go func WithResumeQueue(queueName string) ResumeWorkflowOption ``` Re-enqueue the resumed workflow on the specified queue instead of starting it immediately. #### ResumeWorkflows ```go func ResumeWorkflows[R any](ctx Client, workflowIDs []string, opts ...ResumeWorkflowOption) ([]WorkflowHandle[R], error) ``` Resume multiple workflows in a single database round-trip. Each workflow that exists and is not in a terminal state is re-enqueued; completed or missing workflows are skipped. Unlike [`ResumeWorkflow`](#resumeworkflow), this function does not return an error when some IDs are missing. Accepts the same options as [`ResumeWorkflow`](#resumeworkflow) (e.g., [`WithResumeQueue`](#withresumequeue)). **Parameters:** - **ctx**: The DBOS client or context. - **workflowIDs**: The IDs of the workflows to resume. - **opts**: Optional configuration. #### ForkWorkflow ```go func ForkWorkflow[R any](ctx Client, input ForkWorkflowInput) (WorkflowHandle[R], error) ``` Start a new execution of a workflow from a specific step. The input step ID (`startStep`) must match the step number of the step returned by workflow introspection. The specified `startStep` is the step from which the new workflow will start, so any steps whose ID is less than `startStep` will not be re-executed. **Parameters:** - **ctx**: The DBOS client or context. - **input**: A `ForkWorkflowInput` struct where `OriginalWorkflowID` is mandatory. ```go type ForkWorkflowInput struct { OriginalWorkflowID string // Required: The UUID of the original workflow to fork from ForkedWorkflowID string // Optional: Custom workflow ID for the forked workflow (auto-generated if empty) StartStep uint // Optional: Step to start the forked workflow from (default: 0) ApplicationVersion string // Optional: Application version for the forked workflow (inherits from original if empty) QueueName string // Optional: Queue to enqueue the forked workflow on (defaults to starting immediately) QueuePartitionKey string // Optional: Partition key when enqueueing onto a partitioned queue (requires QueueName) Timeout time.Duration // Optional: Maximum execution time for the forked workflow (default: no timeout) ReplacementChildren map[string]string // Optional: Maps original child workflow IDs to replacement child workflow IDs } ``` If `QueueName` is set, the forked workflow is enqueued on the specified queue instead of starting immediately. Set `QueuePartitionKey` together with `QueueName` to enqueue the forked workflow onto a specific partition of a [partitioned queue](../tutorials/queue-tutorial.md#partitioning-queues). If `Timeout` is set, the forked workflow is cancelled if it runs longer than that duration; the clock starts when the fork begins executing (when it is dequeued, if enqueued). The original workflow's timeout is not inherited. If `ReplacementChildren` is set, it maps original child workflow IDs to replacement child workflow IDs. When the forked workflow encounters a copied step that started a child workflow matching an original ID, it substitutes the replacement ID instead. This is useful when you need to fork a parent workflow that depends on the results of child workflows that have also been forked: ```go // Fork the child first, then fork the parent so its checkpoints point at the new child. childHandle, err := dbos.ForkWorkflow[ChildResult](ctx, dbos.ForkWorkflowInput{ OriginalWorkflowID: "old-child-id", ForkedWorkflowID: "new-child-id", StartStep: 2, }) parentHandle, err := dbos.ForkWorkflow[ParentResult](ctx, dbos.ForkWorkflowInput{ OriginalWorkflowID: "parent-workflow-id", StartStep: 6, ReplacementChildren: map[string]string{"old-child-id": "new-child-id"}, }) ``` #### ForkWorkflows ```go func ForkWorkflows[R any](ctx Client, input ForkWorkflowsInput) ([]WorkflowHandle[R], error) ``` Fork a batch of workflows in a single database round-trip. Each forked workflow gets a new UUID (unless a custom `ForkedWorkflowID` is provided) and executes from its specified `StartStep`, reusing the operation outputs of steps `0` to `StartStep-1` copied from the original workflow. The returned handles are in the same order as `input.Workflows`. **Parameters:** - **ctx**: The DBOS client or context. - **input**: A `ForkWorkflowsInput` struct where `Workflows` is mandatory. ```go type ForkWorkflowsInput struct { Workflows []ForkWorkflowSpec // Required: The workflows to fork ApplicationVersion string // Optional: Application version for the forked workflows (inherits from originals if empty) QueueName string // Optional: Queue to enqueue the forked workflows on (defaults to the internal queue) QueuePartitionKey string // Optional: Partition key when enqueueing the forked workflows onto a partitioned queue Timeout time.Duration // Optional: Maximum execution time for each forked workflow (default: no timeout) ReplacementChildren map[string]string // Optional: Maps original child workflow IDs to replacement child workflow IDs } type ForkWorkflowSpec struct { OriginalWorkflowID string // Required: The UUID of the original workflow to fork from ForkedWorkflowID string // Optional: Custom workflow ID for the forked workflow (auto-generated if empty) StartStep uint // Optional: Step to start the forked workflow from (default: 0) } ``` The `ApplicationVersion`, `QueueName`, `QueuePartitionKey`, `Timeout`, and `ReplacementChildren` settings apply to every forked workflow in the batch; they have the same meaning as in [`ForkWorkflow`](#forkworkflow). **Example:** ```go handles, err := dbos.ForkWorkflows[any](ctx, dbos.ForkWorkflowsInput{ Workflows: []dbos.ForkWorkflowSpec{ {OriginalWorkflowID: "wf-1", StartStep: 2}, {OriginalWorkflowID: "wf-2"}, }, QueueName: "fork_queue", }) ``` #### SetWorkflowDelay ```go func SetWorkflowDelay(ctx Client, workflowID string, opts ...SetWorkflowDelayOption) error ``` Set or update the delay on a [`DELAYED`](#workflowstatustype) workflow. Provide exactly one of [`WithDelayDuration`](#withdelayduration) (relative) or [`WithDelayUntil`](#withdelayuntil) (absolute). Only affects workflows currently in the `DELAYED` status. **Parameters:** - **ctx**: The DBOS client or context. - **workflowID**: The ID of the workflow whose delay to update. - **opts**: Exactly one of `WithDelayDuration` or `WithDelayUntil`. **Example:** ```go // Shorten the delay to 10 seconds from now err := dbos.SetWorkflowDelay(ctx, workflowID, dbos.WithDelayDuration(10*time.Second)) // Or set an absolute deadline err = dbos.SetWorkflowDelay(ctx, workflowID, dbos.WithDelayUntil(time.Now().Add(time.Hour))) ``` ##### WithDelayDuration ```go func WithDelayDuration(d time.Duration) SetWorkflowDelayOption ``` Set a relative delay measured from now. ##### WithDelayUntil ```go func WithDelayUntil(t time.Time) SetWorkflowDelayOption ``` Set an absolute time until which the workflow should remain delayed. #### SetWorkflowAttributes ```go func SetWorkflowAttributes(ctx Client, workflowID string, attributes map[string]any) error ``` Replace the custom [attributes](./workflows-steps.md#withworkflowattributes) attached to an existing workflow. Pass a `nil` attributes map to clear all attributes. Attributes must be JSON-serializable. Returns an error if the workflow does not exist. **Example:** ```go err := dbos.SetWorkflowAttributes(ctx, "my-workflow-id", map[string]any{"customer": "acme"}) ``` #### DeleteWorkflows ```go func DeleteWorkflows(ctx Client, workflowIDs []string, opts ...DeleteWorkflowOption) error ``` Permanently delete one or more workflows and all their associated data (status, step outputs, events, messages, and streams) from the system database, regardless of their current status, including active (`PENDING`, `ENQUEUED`) workflows. :::warning This operation is irreversible. ::: **Parameters:** - **ctx**: The DBOS client or context. - **workflowIDs**: The IDs of the workflows to delete. - **opts**: Optional configuration, documented below. ##### WithDeleteChildren ```go func WithDeleteChildren() DeleteWorkflowOption ``` Also delete all child workflows, recursively. ```go err := dbos.DeleteWorkflows(ctx, []string{"wf-1", "wf-2"}, dbos.WithDeleteChildren()) ``` #### Workflow Status Some workflow introspection and management methods return a `WorkflowStatus`. This object has the following definition: ```go type WorkflowStatus struct { ID string `json:"workflow_uuid"` // Unique identifier for the workflow Status WorkflowStatusType `json:"status"` // Current execution status Name string `json:"name"` // Function name of the workflow AuthenticatedUser string `json:"authenticated_user"` // User who initiated the workflow (if applicable) AssumedRole string `json:"assumed_role"` // Role assumed during execution (if applicable) AuthenticatedRoles []string `json:"authenticated_roles"` // Roles available to the user (if applicable) Output any `json:"output"` // Workflow output (available after completion) Error error `json:"error"` // Error information (if status is ERROR) ExecutorID string `json:"executor_id"` // ID of the executor running this workflow CreatedAt time.Time `json:"created_at"` // When the workflow was created UpdatedAt time.Time `json:"updated_at"` // When the workflow status was last updated ApplicationVersion string `json:"application_version"` // Version of the application that created this workflow ApplicationID string `json:"application_id"` // Application identifier ApplicationName string `json:"application_name"` // Owning application; empty if the workflow is owned by no application Attempts int `json:"attempts"` // Number of execution attempts QueueName string `json:"queue_name"` // Queue name (if workflow was enqueued) Timeout time.Duration `json:"-"` // Workflow timeout duration; rendered as timeout_ms (integer milliseconds) in JSON Deadline time.Time `json:"deadline"` // Absolute deadline for workflow completion StartedAt time.Time `json:"started_at"` // When the workflow execution actually started CompletedAt time.Time `json:"completed_at"` // When the workflow reached a terminal state (SUCCESS, ERROR, or CANCELLED) ForkedFrom string `json:"forked_from"` // ID of the original workflow if this is a fork WasForkedFrom bool `json:"was_forked_from"` // Whether this workflow has been forked from ParentWorkflowID string `json:"parent_workflow_id"` // ID of the parent workflow if this is a child DeduplicationID string `json:"deduplication_id"` // Queue deduplication identifier (if applicable) Input any `json:"input"` // Input parameters passed to the workflow Priority int `json:"priority"` // Execution priority (lower numbers have higher priority) QueuePartitionKey string `json:"queue_partition_key"` // Queue partition key for partitioned queues ClassName string `json:"class_name"` // Class/namespace name for cross-language dispatch ConfigName *string `json:"config_name"` // Instance/config name for cross-language dispatch Serialization string `json:"serialization"` // Serialization format used for inputs/outputs (e.g., "portable_json") DelayUntil time.Time `json:"delay_until"` // Time before which a DELAYED workflow should not be dequeued Attributes map[string]any `json:"attributes"` // Custom key-value attributes attached to the workflow ScheduleName string `json:"schedule_name"` // Name of the schedule that enqueued this workflow (if any) DebounceDeadline time.Time `json:"debounce_deadline"` // Absolute cap beyond which debounce calls may not extend the delay IsDebounced bool `json:"is_debounced"` // Whether this workflow was created by a debouncer } ``` ##### WorkflowStatusType The `WorkflowStatusType` represents the execution status of a workflow: ```go type WorkflowStatusType string const ( WorkflowStatusPending WorkflowStatusType = "PENDING" // Workflow is running or ready to run WorkflowStatusEnqueued WorkflowStatusType = "ENQUEUED" // Workflow is queued and waiting for execution WorkflowStatusDelayed WorkflowStatusType = "DELAYED" // Workflow is delayed and will transition to ENQUEUED after the delay expires WorkflowStatusSuccess WorkflowStatusType = "SUCCESS" // Workflow completed successfully WorkflowStatusError WorkflowStatusType = "ERROR" // Workflow completed with an error WorkflowStatusCancelled WorkflowStatusType = "CANCELLED" // Workflow was cancelled (manually or due to timeout) WorkflowStatusMaxRecoveryAttemptsExceeded WorkflowStatusType = "MAX_RECOVERY_ATTEMPTS_EXCEEDED" // Workflow exceeded maximum retry attempts ) ``` ### Workflow Schedules DBOS lets you schedule workflows to run on a cron expression. Schedules are stored in the database and can be created, paused, resumed, and deleted at runtime. See the [scheduled workflows tutorial](../tutorials/scheduled-workflows.md) for an overview. Scheduled workflows must accept a [`ScheduledWorkflowInput`](#scheduledworkflowinput) as their input parameter. #### ScheduledWorkflowInput ```go type ScheduledWorkflowInput struct { ScheduledTime time.Time // The cron tick time Context json.RawMessage // The user-defined context attached to the schedule, as raw JSON (nil if none) } ``` The input type of a scheduled workflow function. `Context` carries the value set as [`ScheduleSpec.Context`](#schedulespec) when the schedule was created, as raw JSON; decode it with [`DecodeScheduleContext`](#decodeschedulecontext). #### DecodeScheduleContext ```go func DecodeScheduleContext[T any](input ScheduledWorkflowInput) (T, error) ``` Decode the schedule's user-defined context carried by a `ScheduledWorkflowInput` into `T` — typically the same type that was set as `ScheduleSpec.Context` when the schedule was created. Returns the zero value of `T` if the schedule has no context. ```go type ReportConfig struct { Region string `json:"region"` BatchSize int `json:"batch_size"` } func reportWorkflow(ctx dbos.Context, input dbos.ScheduledWorkflowInput) (any, error) { cfg, err := dbos.DecodeScheduleContext[ReportConfig](input) // ... } ``` #### WorkflowSchedule ```go type WorkflowSchedule struct { ScheduleID string // Unique ID assigned to this schedule revision ScheduleName string // User-supplied unique name WorkflowName string // Fully-qualified or custom name of the workflow WorkflowClassName string // Class/namespace (used for cross-language dispatch) Schedule string // Cron expression Status ScheduleStatus // ACTIVE or PAUSED Context json.RawMessage // User-defined context attached to the schedule, as raw JSON LastFiredAt *time.Time // Last time the schedule fired (nil if never) AutomaticBackfill bool // Whether to backfill missed ticks on application start CronTimezone string // IANA timezone name (empty for UTC) QueueName string // Queue on which scheduled workflows are enqueued ApplicationName string // Owning application; empty if the schedule is owned by no application } ``` ##### ScheduleStatus ```go type ScheduleStatus string const ( ScheduleStatusActive ScheduleStatus = "ACTIVE" // Schedule is firing ScheduleStatusPaused ScheduleStatus = "PAUSED" // Schedule is paused ) ``` #### ScheduleSpec Schedules are described by a `ScheduleSpec`: ```go type ScheduleSpec struct { ScheduleName string // Required: unique name of the schedule Schedule string // Required: cron expression driving the schedule WorkflowName string // Name of the target workflow (required unless Workflow is set) Workflow any // Registered scheduled workflow function (Context only; takes precedence over WorkflowName) WorkflowClassName string // Optional class/namespace name for cross-language dispatch Context any // Optional user-defined context (serialized as JSON) passed to each scheduled invocation; decode with DecodeScheduleContext AutomaticBackfill bool // Backfill missed ticks when the schedule is reloaded after downtime CronTimezone string // Optional IANA timezone used to interpret the cron expression QueueName string // Optional queue to route scheduled invocations to (defaults to the internal queue) ApplicationName string // Optional application that owns the schedule and runs its workflows (defaults to the caller's application) } ``` Field notes: - **Workflow vs. WorkflowName**: from a `Context`, set `Workflow` to a scheduled workflow function already registered via [`RegisterWorkflow`](./workflows-steps.md#registerworkflow). From a [standalone client](./dbos-context.md#newclient) (or to target a workflow owned by another process or language), set `WorkflowName` instead. If both are set, `Workflow` wins. - **WorkflowClassName**: set when the target workflow is owned by a runtime that dispatches by class name (e.g. a Python class-based workflow). - **Context**: an arbitrary value serialized as JSON and passed to each scheduled invocation as [`ScheduledWorkflowInput.Context`](#scheduledworkflowinput); decode it in the workflow with [`DecodeScheduleContext`](#decodeschedulecontext). - **AutomaticBackfill**: backfill missed ticks whenever the schedule is reloaded after downtime, or when a paused schedule is resumed. Missed ticks are computed with the schedule's **current** cron expression, over the window from the last fire to now. If you change a schedule's cron expression (e.g. with [`ApplySchedules`](#applyschedules)) while it is not running, the backfill generates one execution per tick of the *new* expression across that entire window—including times the old expression would never have matched. - **CronTimezone**: an [IANA timezone](https://en.wikipedia.org/wiki/List_of_tz_database_time_zones) name (e.g. `"America/New_York"`) in which to interpret the cron expression. Defaults to UTC. - **QueueName**: route each scheduled invocation to the named [queue](./queues.md) instead of the default internal queue. - **ApplicationName**: the application that owns this schedule and runs its workflows. Defaults to the creating context's own application (for a [standalone client](./dbos-context.md#newclient), its `AppName`). Always set an application name either here or in the client's configuration if multiple applications [share a system database](../../explanations/sharing-a-system-database.md). The workflow function must conform to: ```go type ScheduledWorkflowFunc func(ctx Context, input ScheduledWorkflowInput) (any, error) ``` #### CreateSchedule ```go func CreateSchedule(ctx Client, spec ScheduleSpec) error ``` Create a new schedule. Fails if a schedule with the same name already exists. The reconciler loop picks the new schedule up on its next tick and installs it in the cron scheduler. **Parameters:** - **ctx**: The DBOS client or context. - **spec**: A [`ScheduleSpec`](#schedulespec) describing the schedule. **Example:** ```go // From a Context, with a registered workflow function: err := dbos.CreateSchedule(ctx, dbos.ScheduleSpec{ ScheduleName: "my-schedule", Workflow: myPeriodicTask, Schedule: "0 */5 * * * *", Context: "my context", AutomaticBackfill: true, }) // From a standalone client, by workflow name: err = dbos.CreateSchedule(client, dbos.ScheduleSpec{ ScheduleName: "my-schedule", WorkflowName: "myPeriodicTask", Schedule: "0 */5 * * * *", }) ``` #### ApplySchedules ```go func ApplySchedules(ctx Client, schedules []ScheduleSpec) error ``` Atomically create or update a list of schedules in a single transaction. Existing schedules are upserted by name: all definition fields (workflow name and class, cron expression, context, timezone, queue, and backfill flag) are replaced with the new entry's values, while the schedule's ID, status, and last-fired time are preserved. Useful for defining a fixed set of static schedules on application start. `ApplySchedules` cannot be called from within a workflow. :::warning Because every definition field is replaced, omitting an optional field on re-apply clears it. In particular, a schedule previously routed to a named queue reverts to the internal queue if the new entry does not set `QueueName`. ::: **Example:** ```go err := dbos.ApplySchedules(ctx, []dbos.ScheduleSpec{ {ScheduleName: "a", Workflow: workflowA, Schedule: "0 */10 * * * *"}, {ScheduleName: "b", Workflow: workflowB, Schedule: "0 0 0 * * *"}, }) ``` #### GetSchedule ```go func GetSchedule(ctx Client, scheduleName string) (WorkflowSchedule, error) ``` Retrieve a [`WorkflowSchedule`](#workflowschedule) by name. If no schedule with that name exists, the returned error matches `dbos.ErrScheduleNotFound`. #### ListSchedules ```go func ListSchedules(ctx Client, opts ...ListSchedulesOption) ([]WorkflowSchedule, error) ``` List schedules, optionally filtered. By default, only schedules owned by the calling context's application (plus schedules owned by no application) are listed; a [standalone client](./dbos-context.md#newclient) with no `AppName` lists every application's schedules. ##### WithScheduleStatuses ```go func WithScheduleStatuses(statuses ...ScheduleStatus) ListSchedulesOption ``` Filter by one or more [`ScheduleStatus`](#schedulestatus) values. ##### WithScheduleWorkflowNames ```go func WithScheduleWorkflowNames(names ...string) ListSchedulesOption ``` Filter by workflow name(s). Use the fully qualified name or the custom name registered via [`WithWorkflowName`](./workflows-steps.md#withworkflowname). ##### WithScheduleNamePrefixes ```go func WithScheduleNamePrefixes(prefixes ...string) ListSchedulesOption ``` Filter by schedule name prefix(es). ##### WithScheduleNames ```go func WithScheduleNames(names ...string) ListSchedulesOption ``` Filter by exact schedule name(s). ##### WithScheduleApplicationNames ```go func WithScheduleApplicationNames(names ...string) ListSchedulesOption ``` List schedules owned by these applications (schedules owned by no application are always included). #### PauseSchedule ```go func PauseSchedule(ctx Client, scheduleName string) error ``` Pause a schedule so it stops firing. The schedule's cron entry is removed on the next reconciler tick. #### ResumeSchedule ```go func ResumeSchedule(ctx Client, scheduleName string) error ``` Resume a paused schedule. If the schedule was created with `AutomaticBackfill: true` (see [`ScheduleSpec`](#schedulespec)), missed ticks during the pause are backfilled. #### DeleteSchedule ```go func DeleteSchedule(ctx Client, scheduleName string) error ``` Delete a schedule. The schedule's cron entry is removed on the next reconciler tick. #### BackfillSchedule ```go func BackfillSchedule(ctx Client, scheduleName string, start, end time.Time) ([]string, error) ``` Backfill missed executions for the range `[start, end]`, returning the IDs of the enqueued workflows. Already-executed ticks are automatically skipped, so it is safe to overlap ranges. Cannot be called from within a workflow. **Example:** ```go ids, err := dbos.BackfillSchedule(ctx, "my-schedule", time.Date(2025, 1, 1, 0, 0, 0, 0, time.UTC), time.Date(2025, 1, 2, 0, 0, 0, 0, time.UTC), ) ``` #### TriggerSchedule ```go func TriggerSchedule[R any](ctx Client, scheduleName string) (WorkflowHandle[R], error) ``` Trigger a schedule to fire immediately and return a [`WorkflowHandle`](./workflows-steps.md#workflowhandle) for the enqueued workflow. The generic `TriggerSchedule` returns a typed handle whose `GetResult` decodes the triggered workflow's output into type `R`. Cannot be called from within a workflow. ### Application Versions DBOS tracks each application version that has launched against the system database. You can use these methods to inspect the registered versions and control which one is treated as latest—for example, to recover workflows onto a specific version after a rollout. Versions are tracked per application: if multiple applications [share a system database](../../explanations/sharing-a-system-database.md), these methods only see versions registered by the calling handle's application (plus versions owned by no application), so one application's deployments do not affect which version its peers consider latest. A [standalone client](./dbos-context.md#newclient) with no `AppName` sees every application's versions. #### VersionInfo ```go type VersionInfo struct { ID string // Internal version ID Name string // Application version name Timestamp int64 // Epoch milliseconds; the most recent timestamp identifies the latest version CreatedAt int64 // Epoch milliseconds at which the version was first registered ApplicationName string // Owning application; empty if the version is owned by no application } ``` #### ListApplicationVersions ```go func ListApplicationVersions(ctx Client) ([]VersionInfo, error) ``` Return every application version registered in the system database, ordered by timestamp (newest first). **Parameters:** - **ctx**: The DBOS client or context. #### GetLatestApplicationVersion ```go func GetLatestApplicationVersion(ctx Client) (VersionInfo, error) ``` Return the application version with the most recent timestamp. If no versions are registered, the returned error matches `dbos.ErrNoApplicationVersions`. **Parameters:** - **ctx**: The DBOS client or context. #### SetLatestApplicationVersion ```go func SetLatestApplicationVersion(ctx Client, versionName string) error ``` Mark the named application version as latest by updating its timestamp to the current time. Promoting a version registered by a different application returns an error. **Parameters:** - **ctx**: The DBOS client or context. - **versionName**: The name of the registered application version to mark as latest. ### Application Rename #### RenameApplication ```go func RenameApplication(ctx Client, input RenameApplicationInput) (ApplicationRowCounts, error) ``` ```go type RenameApplicationInput struct { OldName string // The application's previous name. Empty moves nothing but the unclaimed rows, so it requires AdoptUnclaimedRows. NewName string // The application that ends up owning the rows. Required. BatchSize int // Completed workflows and steps re-owned per transaction. Zero defaults to DefaultRenameBatchSize (10,000). AdoptUnclaimedRows bool // Also transfer rows no application owns. } type ApplicationRowCounts struct { Queues int64 `json:"queues"` Schedules int64 `json:"schedules"` Versions int64 `json:"versions"` Workflows int64 `json:"workflows"` Steps int64 `json:"steps"` } ``` Every workflow, step, queue, schedule, and application version is owned by the application (identified by its configured [`AppName`](./configuration.md)) that created it. After renaming an application, use this method (or the `dbos rename-application` [CLI command](./cli.md)) to transfer everything owned by the old name to the new name. Returns the number of rows transferred, by table. Queues, schedules, versions, and in-flight workflows are transferred in a single transaction; completed workflows and their steps are then transferred in batches of `BatchSize`. The operation is idempotent: if interrupted, running it again resumes where it left off. Set `AdoptUnclaimedRows` to also transfer rows owned by no application, such as rows created before upgrading to a DBOS version supporting application ownership. :::warning Stop the application being renamed before running this. A running application would race the rename, creating new work under its old name. ::: ### DBOS Variables #### GetWorkflowID ```go func GetWorkflowID(ctx Context) (string, error) ``` Return the ID of the current workflow, if in a workflow. Returns an error if not called from within a workflow context. **Parameters:** - **ctx**: The DBOS context. #### GetStepID ```go func GetStepID(ctx Context) (int, error) ``` Return the current value of the step counter within a workflow (the ID of the most recently started step). Returns an error if not called from within a workflow context. **Parameters:** - **ctx**: The DBOS context. ### Portable Serialization Options and Types These options enable [cross-language interoperability](../../explanations/portable-workflows.md) by using the portable JSON serialization format. #### WithPortableSend ```go func WithPortableSend() SendOption ``` Configure [`Send`](#send) or [`SendBulk`](#sendbulk) to use the portable JSON serializer, enabling cross-language message passing. #### WithPortableSetEvent ```go func WithPortableSetEvent() SetEventOption ``` Configure [`SetEvent`](#setevent) to use the portable JSON serializer, enabling cross-language event consumption. #### WithPortableWriteStream ```go func WithPortableWriteStream() WriteStreamOption ``` Configure [`WriteStream`](#writestream) to use the portable JSON serializer, enabling cross-language stream reading. #### WithReadStreamSnapshot ```go func WithReadStreamSnapshot() ReadStreamOption ``` Configure [`ReadStream`](#readstream) to return as soon as all currently-available values have been drained, instead of blocking until the stream is closed or the workflow becomes inactive. #### WithReadStreamFromOffset ```go func WithReadStreamFromOffset(offset int) ReadStreamOption ``` Configure [`ReadStream`](#readstream) to start reading from the given base offset (zero-indexed). Combined with [`WithReadStreamSnapshot`](#withreadstreamsnapshot), this allows you to poll a stream incrementally. #### PortableWorkflowError ```go type PortableWorkflowError struct { Name string // The error type/class name Message string // Human-readable error message Code any // Optional application-specific error code Data any // Optional structured error details } ``` A structured error type for workflows using portable serialization. Portable workflows automatically serialize errors in this format. ```go return nil, &dbos.PortableWorkflowError{ Name: "ValidationError", Message: "invalid input", Code: 400, } ``` #### PortableWorkflowArgs ```go type PortableWorkflowArgs struct { PositionalArgs []any `json:"positionalArgs"` NamedArgs map[string]any `json:"namedArgs"` } ``` The cross-language envelope for workflow inputs. When passed as the input to [`Enqueue`](#enqueue), portable JSON serialization is used automatically. Further, a portable workflow ran with [`RunWorkflow`](workflows-steps.md#runworkflow) will serialize its input in this format automatically. ```go args := dbos.PortableWorkflowArgs{ PositionalArgs: []any{"order-123", 42}, } handle, err := dbos.Enqueue[any]( client, "queue", "target_workflow", args, ) ``` ### Alerting #### SetAlertHandler ```go func SetAlertHandler(ctx Context, handler AlertHandler) ``` ```go type AlertHandler func(name string, message string, metadata map[string]string) ``` Register a handler to receive [alerts](../../conductor/alerting.md) from Conductor. The handler function is called with three arguments: - **name**: The type of alert rule. One of `WorkflowFailure`, `SlowQueue`, or `UnresponsiveApplication`. - **message**: The alert message. - **metadata**: A map of string key-value pairs with additional alert information. Only one alert handler may be registered per application, and it must be registered before [`Launch`](./dbos-context.md#launch) is called. If no handler is registered, alerts are logged automatically. **Example syntax:** ```go dbos.SetAlertHandler(dbosContext, func(ruleType string, message string, metadata map[string]string) { slog.Warn(fmt.Sprintf("Alert received: %s - %s", ruleType, message)) for key, value := range metadata { slog.Warn(fmt.Sprintf(" %s: %s", key, value)) } }) ``` --- ## Queues Workflow queues allow you to ensure that workflow functions will be run, without starting them immediately. Queues are useful for controlling the number of workflows run in parallel, or the rate at which they are started. Queue configuration is persisted to the system database, so any DBOS process connected to the same system database can register, retrieve, and reconfigure queues. If multiple applications [share a system database](../../explanations/sharing-a-system-database.md), each queue is owned by the application that registers it, and only that application dequeues workflows from it. ### Queue Management #### RegisterQueue ```go func RegisterQueue(ctx Client, name string, options ...QueueOption) (Queue, error) ``` Register a queue and persist its configuration to the system database, returning a [`Queue`](#queue-interface). If a queue with the same name already exists in the database, the [`WithQueueOnConflict`](#withqueueonconflict) option controls whether its configuration is overwritten. Queues may be registered at any time, including after `Launch()`; live workers periodically reload queue configuration, so changes take effect without a restart. You can enqueue a workflow using the [`WithQueue`](./workflows-steps.md#withqueue) parameter of [`RunWorkflow`](./workflows-steps.md#runworkflow). **Parameters:** - **ctx**: The DBOS client or context. - **name**: The name of the queue. Must be unique among all queues in the application. - **options**: Functional options for the queue, documented below. **Example Syntax:** ```go queue, err := dbos.RegisterQueue(ctx, "email-queue", dbos.WithWorkerConcurrency(5), dbos.WithRateLimiter(&dbos.RateLimiter{ Limit: 100, Period: 60 * time.Second, // 100 workflows per minute }), dbos.WithPriorityEnabled(), ) // Enqueue workflows to this queue by passing its handle to WithQueue: handle, err := dbos.RunWorkflow(ctx, SendEmailWorkflow, emailData, dbos.WithQueue(queue)) ``` ##### WithWorkerConcurrency ```go func WithWorkerConcurrency(concurrency int) QueueOption ``` Set the maximum number of workflows from this queue that may run concurrently within a single DBOS process. ##### WithGlobalConcurrency ```go func WithGlobalConcurrency(concurrency int) QueueOption ``` Set the maximum number of workflows from this queue that may run concurrently. Unset by default (no limit). This concurrency limit is global across all DBOS processes using this queue. ##### WithPriorityEnabled ```go func WithPriorityEnabled() QueueOption ``` Enable setting priority for workflows on this queue. ##### WithRateLimiter ```go func WithRateLimiter(limiter *RateLimiter) QueueOption ``` ```go type RateLimiter struct { Limit int // Maximum number of workflows to start within the period Period time.Duration // Time period for the rate limit } ``` A limit on the maximum number of functions which may be started in a given period. ##### WithPartitionConcurrency ```go func WithPartitionConcurrency(concurrency int) QueueOption ``` Set the maximum number of workflows from any one [partition](../tutorials/queue-tutorial.md#partitioning-queues) of this queue that may run concurrently across all DBOS processes. Must be at least 1 and less than or equal to the queue's global concurrency. Setting any partition limit (`WithPartitionConcurrency`, `WithPartitionWorkerConcurrency`, or `WithPartitionRateLimiter`) makes the queue **partitioned**: every workflow enqueued on it must supply a partition key with [`WithQueuePartitionKey`](./workflows-steps.md#withqueuepartitionkey), and the queue dequeues from each partition separately. A partitioned queue enforces its partition limits and its queue-wide limits (`WithGlobalConcurrency`, `WithWorkerConcurrency`, `WithRateLimiter`) at the same time. ##### WithPartitionWorkerConcurrency ```go func WithPartitionWorkerConcurrency(concurrency int) QueueOption ``` Set the maximum number of workflows from any one partition of this queue that may run concurrently within a single DBOS process. Must be at least 1 and less than or equal to the queue's partition concurrency, worker concurrency, and global concurrency. Setting this limit makes the queue partitioned. ##### WithPartitionRateLimiter ```go func WithPartitionRateLimiter(limiter *RateLimiter) QueueOption ``` A limit on the maximum number of workflows which may be started from any one partition of this queue in a given period. The limit is applied to each partition separately. Setting this limit makes the queue partitioned. **Example Syntax:** ```go // Create a partitioned queue with a per-partition concurrency limit of 1 partitionedQueue, err := dbos.RegisterQueue(ctx, "user-tasks", dbos.WithPartitionConcurrency(1), ) // Enqueue workflows with different partition keys // At most one workflow per user can run at once, but workflows from different users can run concurrently handle1, _ := dbos.RunWorkflow(ctx, ProcessTask, task1, dbos.WithQueue(partitionedQueue), dbos.WithQueuePartitionKey("user-123"), ) handle2, _ := dbos.RunWorkflow(ctx, ProcessTask, task2, dbos.WithQueue(partitionedQueue), dbos.WithQueuePartitionKey("user-456"), ) ``` ##### WithPartitionQueue ```go func WithPartitionQueue() QueueOption ``` :::warning Deprecated `WithPartitionQueue` is deprecated. Partition a queue by setting a partition limit ([`WithPartitionConcurrency`](#withpartitionconcurrency), [`WithPartitionWorkerConcurrency`](#withpartitionworkerconcurrency), or [`WithPartitionRateLimiter`](#withpartitionratelimiter)) instead. ::: Enable the legacy partitioned queue mode, under which the queue's global concurrency, worker concurrency, and rate limit each apply to individual partitions instead of the queue as a whole. For example, a queue registered with `WithPartitionQueue()` and `WithGlobalConcurrency(1)` runs at most one workflow from each partition at a time. The equivalent under the partition limits is `WithPartitionConcurrency(1)`, which additionally lets you keep queue-wide limits (see [Combining Queue-Wide and Per-Partition Limits](../tutorials/queue-tutorial.md#combining-queue-wide-and-per-partition-limits)). `WithPartitionQueue` cannot be combined with a partition limit in the same `RegisterQueue` call. A queue registered with it rejects the `Set*` methods for its limits with an error matching `dbos.ErrInvalidOption`: re-register the queue with the partition limits instead. Its `Get*` methods report each limit at the scope it is enforced, so for example `GetPartitionConcurrency` returns the value passed to `WithGlobalConcurrency` and `GetGlobalConcurrency` returns `nil`. ##### WithQueueBasePollingInterval ```go func WithQueueBasePollingInterval(interval time.Duration) QueueOption ``` Set the base polling interval for this queue. This also acts as the minimum (fastest) interval. Polling intervals are subject to base 2 exponential backoff. **Example Syntax:** ```go queue, err := dbos.RegisterQueue(ctx, "email-queue", dbos.WithQueueBasePollingInterval(100*time.Millisecond)) ``` ##### WithQueueApplicationName ```go func WithQueueApplicationName(name string) QueueOption ``` Set the application that owns the queue and dequeues workflows from it. Defaults to the registering context's own application (for a [standalone client](./dbos-context.md#newclient), its `AppName`). Registering a queue already owned by a different application returns an error. ##### WithQueueOnConflict ```go func WithQueueOnConflict(policy QueueConflictResolution) QueueOption type QueueConflictResolution string const ( QueueConflictUpdateIfLatestVersion QueueConflictResolution = "update_if_latest_version" QueueConflictAlwaysUpdate QueueConflictResolution = "always_update" QueueConflictNeverUpdate QueueConflictResolution = "never_update" ) ``` Set how `RegisterQueue` behaves when a queue with the same name already exists in the system database: - **QueueConflictUpdateIfLatestVersion** (default): overwrite the existing configuration only if the running application is the latest registered application version. This prevents older versions in a rolling deploy from overwriting a newer configuration. - **QueueConflictAlwaysUpdate**: always overwrite the existing configuration. - **QueueConflictNeverUpdate**: leave the existing configuration unchanged. The returned queue reflects the persisted configuration, not the supplied options. #### RetrieveQueue ```go func RetrieveQueue(ctx Client, name string) (Queue, error) ``` Retrieve a queue by name from the system database. If no queue with that name has been registered, returns an error matching `dbos.ErrQueueNotFound`: ```go q, err := dbos.RetrieveQueue(ctx, name) if errors.Is(err, dbos.ErrQueueNotFound) { /* absent */ } ``` **Example Syntax:** ```go queue, err := dbos.RetrieveQueue(ctx, "email-queue") if err != nil { return err } fmt.Println("Priority enabled:", queue.GetPriorityEnabled()) ``` #### ListQueues ```go func ListQueues(ctx Client, opts ...ListQueuesOption) ([]Queue, error) ``` Return queues registered in the system database. By default, only queues owned by the calling context's application (plus queues owned by no application) are returned; a [standalone client](./dbos-context.md#newclient) with no `AppName` returns every application's queues. ##### WithListQueuesApplicationNames ```go func WithListQueuesApplicationNames(names ...string) ListQueuesOption ``` List queues owned by these applications instead (queues owned by no application are always included). #### DeleteQueue ```go func DeleteQueue(ctx Client, name string) error ``` Delete a queue from the system database. No-op if no queue with that name exists. :::warning Workflows already enqueued on a deleted queue can no longer be dequeued, executed, or recovered. However, if a queue with the same name is later registered, it will dequeue the leftover workflows. Do not rely on this: stale workflows unexpectedly resuming on a future queue is rarely the intended behavior. Instead, cancel or drain pending workflows on the queue before deleting it. ::: ### Queue Interface A `Queue` is returned from [`RegisterQueue`](#registerqueue), [`RetrieveQueue`](#retrievequeue), and [`ListQueues`](#listqueues). Its `Get*` methods reflect the queue's configuration as of the most recent read from the database; the `Set*` methods update the configuration in the database. `GetPartitionQueue` reports whether the queue is [partitioned](../tutorials/queue-tutorial.md#partitioning-queues), whether by a partition limit or by the deprecated [`WithPartitionQueue`](#withpartitionqueue) option. Unlike the other properties, ownership cannot be reconfigured: there is no `SetApplicationName`. Ownership is only transferred by [`RenameApplication`](./methods.md#renameapplication). ```go type Queue interface { GetName() string GetGlobalConcurrency() *int GetWorkerConcurrency() *int GetRateLimit() *RateLimiter GetPartitionConcurrency() *int GetPartitionWorkerConcurrency() *int GetPartitionRateLimit() *RateLimiter GetPriorityEnabled() bool GetPartitionQueue() bool GetPollingInterval() time.Duration GetApplicationName() string SetGlobalConcurrency(ctx Client, value *int) error SetWorkerConcurrency(ctx Client, value *int) error SetRateLimit(ctx Client, value *RateLimiter) error SetPartitionConcurrency(ctx Client, value *int) error SetPartitionWorkerConcurrency(ctx Client, value *int) error SetPartitionRateLimit(ctx Client, value *RateLimiter) error SetPriorityEnabled(ctx Client, value bool) error SetPartitionQueue(ctx Client, value bool) error SetPollingInterval(ctx Client, value time.Duration) error } ``` #### Reconfiguring Queues Because queue configuration lives in the system database, you can change a queue's configuration at runtime without redeploying or restarting your workers. Workers pick up the new configuration on their next polling iteration. For the concurrency and rate limit setters, pass `nil` to clear the limit. Each change is validated against the queue's latest persisted configuration: a concurrency limit must be greater than or equal to its partition counterpart, a worker limit must be less than or equal to its global counterpart, and partition limits must be at least 1. Setting any partition limit (`SetPartitionConcurrency`, `SetPartitionWorkerConcurrency`, or `SetPartitionRateLimit`) makes the queue [partitioned](../tutorials/queue-tutorial.md#partitioning-queues); clearing all of them makes it unpartitioned again. `SetPartitionQueue` toggles the deprecated [`WithPartitionQueue`](#withpartitionqueue) mode, and cannot be used on a queue partitioned by its partition limits. :::warning Take care when partitioning a queue at runtime: workflows already enqueued on it have no partition key and will not be dequeued until the queue is unpartitioned. ::: ```go queue, err := dbos.RetrieveQueue(ctx, "email-queue") if err != nil { return err // dbos.ErrQueueNotFound if the queue does not exist } concurrency := 50 if err := queue.SetGlobalConcurrency(ctx, &concurrency); err != nil { return err } if err := queue.SetRateLimit(ctx, &dbos.RateLimiter{Limit: 500, Period: 60 * time.Second}); err != nil { return err } ``` :::warning If your application calls [`RegisterQueue`](#registerqueue) on startup, the next process to start can overwrite settings you applied at runtime via `Set*` methods. Either update the `RegisterQueue` call to match the new configuration, or pass `WithQueueOnConflict(dbos.QueueConflictNeverUpdate)` to preserve the runtime changes. ::: ### ListenQueues ```go func ListenQueues(ctx Context, names ...string) ``` Configure which queues the current DBOS process should listen to for workflow execution. By default, all registered queues are listened to. When `ListenQueues` is called, only the specified queues (and the internal DBOS queue) will be processed by the queue runner. This allows multiple DBOS processes to share the same queues but listen to different subsets. A queue is identified by name, so a queue can be listened to even before it exists in the database; names are resolved against the database on each polling iteration. ```go dbos.RegisterQueue(ctx, "queue-1") dbos.RegisterQueue(ctx, "queue-2") // Only listen to queue-1 and queue-2. dbos.ListenQueues(ctx, "queue-1", "queue-2") ``` Each call to `ListenQueues` replaces the whole listen set; passing an empty set listens to every queue. The listen set may be changed at any time, including after `Launch()`. Use `ListenedQueues(ctx)` to retrieve the current set. --- ### Debouncer A debouncer delays workflow execution until a configurable delay has elapsed since the last invocation. Each subsequent call pushes back the start time by the delay amount. This is useful when you want to coalesce rapid successive triggers (e.g., text field edits, sensor data) into a single workflow execution. See the [debouncing tutorial](../tutorials/workflow-tutorial.md#debouncing) for usage examples. #### NewDebouncer ```go func NewDebouncer[R any, P any](ctx Context, workflow Workflow[P, R], opts ...DebouncerOption) (*Debouncer[R, P], error) ``` Create a new debouncer for the specified workflow. The workflow must be registered before creating the debouncer. Debouncers can be created at any time, including after `Launch()`. Multiple debouncers can be created for the same workflow. Both type parameters are inferred from the workflow function. **Parameters:** - **ctx**: The Context. - **workflow**: The workflow function to debounce (must be registered). - **opts**: Optional configuration, documented below. #### WithDebouncerTimeout ```go func WithDebouncerTimeout(timeout time.Duration) DebouncerOption ``` Set the maximum time before starting the workflow, measured from the first debounce call for a given key. If the timeout is zero (the default), there is no maximum time limit and calling the workflow can be pushed back indefinitely. #### WithDebouncerQueue ```go func WithDebouncerQueue(queueName string) DebouncerOption ``` Run the debounced workflow on the named queue instead of the DBOS internal queue. Debounce keys are scoped to the queue. The queue is fixed per debouncer and must be registered (see [`RegisterQueue`](#registerqueue)); `Debounce` calls cannot override it. `NewDebouncer` validates at creation time that the queue is registered; `NewDebouncerClient` does not. #### WithDebouncerInstance ```go func WithDebouncerInstance(instance ConfiguredInstance) DebouncerOption ``` Target the workflow registration bound to the given configured instance (see [`WithInstance`](./workflows-steps.md#withinstance)). Required when the debounced workflow is a method of a configured instance. ```go debouncer, err := dbos.NewDebouncer(ctx, slack.Send, dbos.WithDebouncerInstance(slack)) ``` #### Debouncer.Debounce ```go func (d *Debouncer[R, P]) Debounce(ctx Context, key string, delay time.Duration, input P, opts ...WorkflowOption) (WorkflowHandle[R], error) ``` Debounce a workflow invocation. If no debouncer is active for the given key, one is started with the specified delay. If a debouncer is already active for the key, the delay is pushed back and the input is updated. When the delay expires _or_ the debouncer preconfigured timeout is reached, the target workflow is executed with the most recent input. **Parameters:** - **ctx**: The Context. - **key**: A unique key to group debounce calls. Calls with the same key are debounced together. - **delay**: Time by which to delay workflow execution from this call. - **input**: Input parameters to pass to the workflow. - **opts**: Optional workflow options (e.g., `WithWorkflowID`). Options the debounce owns or cannot support (`WithQueue`, `WithDeduplicationID`, `WithDelay`, `WithPriority`, `WithQueuePartitionKey`, `WithDeduplicationPolicy`) are rejected with an error matching `dbos.ErrInvalidOption`. **Returns:** - A [WorkflowHandle](./workflows-steps.md#workflowhandle) for the target workflow. You can also create a debouncer from outside a DBOS application using a [`DebouncerClient`](#newdebouncerclient). #### NewDebouncerClient ```go func NewDebouncerClient[R any, P any](workflowName string, client Client, opts ...DebouncerOption) *DebouncerClient[R, P] ``` `R` is the workflow's result type and `P` its input type; neither can be inferred, so both must be named explicitly. Create a new debouncer for use from outside a DBOS application. Similar to [`NewDebouncer`](#newdebouncer) but uses a [standalone client](./dbos-context.md#newclient) instead of a Context and takes a workflow name string instead of a function reference. **Parameters:** - **workflowName**: The name of the workflow to debounce. - **client**: The DBOS client to use for operations. - **opts**: Optional configuration — the same `DebouncerOption`s as [`NewDebouncer`](#newdebouncer), plus the client-specific options below. #### WithDebouncerClassName ```go func WithDebouncerClassName(className string) DebouncerOption ``` Set the class/namespace name recorded for the debounced workflow. Use with `NewDebouncerClient` when the target workflow is registered under a class name — for example by another language's runtime, which may resolve dequeued workflows by class name. #### WithDebouncerConfigName ```go func WithDebouncerConfigName(configName string) DebouncerOption ``` Target the workflow registration bound to the configured instance with the given config name (see [`WithInstance`](./workflows-steps.md#withinstance)). Required when the debounced workflow is a method of a configured instance. Use with `NewDebouncerClient`, where the instance object itself is not available (from a Context, use [`WithDebouncerInstance`](#withdebouncerinstance) instead). ```go dc := dbos.NewDebouncerClient[string, string]("Send", client, dbos.WithDebouncerConfigName("slack")) ``` #### DebouncerClient.Debounce ```go func (dc *DebouncerClient[R, P]) Debounce(key string, delay time.Duration, input P, opts ...WorkflowOption) (WorkflowHandle[R], error) ``` Debounce a workflow invocation from outside a DBOS application. Behaves the same as [`Debouncer.Debounce`](#debouncerdebounce) but does not require a Context. **Parameters:** - **key**: A unique key to group debounce calls. - **delay**: Time by which to delay workflow execution. - **input**: Input parameters to pass to the workflow. - **opts**: Optional workflow options, with the same restrictions as [`Debouncer.Debounce`](#debouncerdebounce). --- ## Workflows & Steps #### RegisterWorkflow ```go func RegisterWorkflow[P any, R any](ctx Context, fn Workflow[P, R], opts ...WorkflowRegistrationOption) ``` Register a function as a DBOS workflow. All workflows must be registered before the context is launched. Workflow functions must be compatible with the following signature: ```go type Workflow[P any, R any] func(ctx Context, input P) (R, error) ``` Returned errors are persisted with [gob](https://pkg.go.dev/encoding/gob), preserving their concrete type when read back (e.g., from a workflow handle in another process). An error that cannot be gob-encoded—including those created by `errors.New` or `fmt.Errorf`, whose fields are unexported—is stored as its message string only, so `errors.Is` and `errors.As` will not match it after a database round-trip. To preserve a custom error type, give it exported fields and register it with [`gob.Register`](https://pkg.go.dev/encoding/gob#Register). **Parameters:** - **ctx**: The Context. - **fn**: The workflow function to register. - **opts**: Functional options for workflow registration, documented below. ##### WithMaxRecoveryAttempts ```go func WithMaxRecoveryAttempts(maxRetries int) WorkflowRegistrationOption ``` Configure the maximum number of times execution of a workflow may be attempted. If `WithMaxRecoveryAttempts(n)` is set, the workflow may be attempted at most `n + 1` times (one initial execution plus `n` retries). If this limit is exceeded, its status is set to `MAX_RECOVERY_ATTEMPTS_EXCEEDED` and it will no longer be recovered automatically. This acts as a [dead letter queue](https://en.wikipedia.org/wiki/Dead_letter_queue), preventing a buggy workflow that crashes its application from doing so infinitely. Use [`ResumeWorkflow`](./methods.md#resumeworkflow) to manually resume a workflow that has exceeded its limit after fixing the underlying issue. ```go // Register a workflow that can be attempted at most 4 times (1 initial + 3 retries) dbos.RegisterWorkflow(dbosContext, myWorkflow, dbos.WithMaxRecoveryAttempts(3)) ``` :::info Workflow Attempts The `Attempts` field in [`WorkflowStatus`](./methods.md#workflow-status) tracks how many times a workflow has been executed: `1` on first execution, `0` if enqueued but not yet dequeued, and incremented by `1` on each recovery or dequeue. The attempt count is incremented one last time before a workflow is placed in the DLQ—for example, a workflow with max retries of 1 that has been moved to the DLQ will show 3 attempts. ::: ##### WithWorkflowName ```go func WithWorkflowName(name string) WorkflowRegistrationOption ``` Register a workflow with a custom name. If not provided, the name of the workflow function is used. ##### WithInstance ```go func WithInstance(instance ConfiguredInstance) WorkflowRegistrationOption ``` Register a workflow method bound to a specific configured instance. Method values bound to different receivers (e.g. `a.Run` and `b.Run`) share a function name, so each instance's method must be registered under a per-instance key, derived from the instance's config name. The instance must implement the `ConfiguredInstance` interface: ```go type ConfiguredInstance interface { ConfigName() string } ``` `ConfigName` must return a stable, unique name for the instance: it is durably recorded so recovery runs the workflow on the correct instance. Instances must be registered with the same config name on every process start, before `Launch()`. Run a workflow registered with `WithInstance` using the matching [`WithRunInstance`](#withruninstance) option. **Example syntax:** ```go type Messenger struct { name string } func (m *Messenger) ConfigName() string { return m.name } func (m *Messenger) Send(ctx dbos.Context, message string) (string, error) { // Workflow implementation using m... return "sent", nil } slack := &Messenger{name: "slack"} email := &Messenger{name: "email"} dbos.RegisterWorkflow(ctx, slack.Send, dbos.WithInstance(slack)) dbos.RegisterWorkflow(ctx, email.Send, dbos.WithInstance(email)) ``` #### RunWorkflow ```go func RunWorkflow[P any, R any](ctx Context, fn Workflow[P, R], input P, opts ...WorkflowOption) (WorkflowHandle[R], error) ``` Execute a workflow function. The workflow may execute immediately or be enqueued for later execution based on options. Returns a [WorkflowHandle](#workflowhandle) that can be used to check the workflow's status or wait for its completion and retrieve its results. **Parameters:** - **ctx**: The Context. - **fn**: The workflow function to execute. - **input** The input to the workflow function. - **opts**: Functional options for workflow execution, documented below. **Example Syntax**: ```go func workflow(ctx dbos.Context, input string) (string, error) { return "success", err } func example(input string) error { handle, err := dbos.RunWorkflow(dbosContext, workflow, input) if err != nil { return err } result, err := handle.GetResult() if err != nil { return err } fmt.Println("Workflow result:", result) return nil } ``` ##### WithWorkflowID ```go func WithWorkflowID(id string) WorkflowOption ``` Run the workflow with a custom workflow ID. If not specified, a UUID workflow ID is generated. ##### WithRunInstance ```go func WithRunInstance(instance ConfiguredInstance) WorkflowOption ``` Run a workflow method registered with [`WithInstance`](#withinstance). ```go handle, err := dbos.RunWorkflow(ctx, slack.Send, input, dbos.WithRunInstance(slack)) ``` ##### WithQueue ```go func WithQueue(queue Queue) WorkflowOption ``` Enqueue the workflow to the given queue instead of executing it immediately. Queued workflows will be dequeued and executed according to the queue's configuration. The queue must be a non-nil [`Queue`](./queues.md#queue-interface) handle returned by [`RegisterQueue`](./queues.md#registerqueue), [`RetrieveQueue`](./queues.md#retrievequeue), or [`ListQueues`](./queues.md#listqueues); passing `nil` makes the enclosing `RunWorkflow` call return an error. To enqueue by name instead (for example, from a standalone client), use [`Enqueue`](./methods.md#enqueue). ##### WithDeduplicationID ```go func WithDeduplicationID(id string) WorkflowOption ``` Set a deduplication ID for this workflow. Should be used alongside `WithQueue`. At any given time, only one workflow with a specific deduplication ID can be enqueued in a given queue. If a workflow with a deduplication ID is currently enqueued or actively executing (status `ENQUEUED` or `PENDING`), subsequent workflow enqueue attempt with the same deduplication ID in the same queue will raise an exception. This behavior can be changed with [`WithDeduplicationPolicy`](#withdeduplicationpolicy). ##### WithDeduplicationPolicy ```go func WithDeduplicationPolicy(policy DeduplicationPolicy) WorkflowOption ``` Set how a colliding deduplication ID is handled for a queued workflow. Must be used alongside `WithQueue` and `WithDeduplicationID`. ```go type DeduplicationPolicy int const ( // DeduplicationPolicyReject (default) returns a ErrorCodeQueueDeduplicated error if another workflow // already holds the deduplication ID on the queue. DeduplicationPolicyReject DeduplicationPolicy = iota // DeduplicationPolicyReturnExisting returns a handle to the existing workflow instead of an error. DeduplicationPolicyReturnExisting ) ``` ```go handle, err := dbos.RunWorkflow(ctx, taskWorkflow, task, dbos.WithQueue(queue), dbos.WithDeduplicationID("user_12345"), dbos.WithDeduplicationPolicy(dbos.DeduplicationPolicyReturnExisting), ) ``` ##### WithPriority ```go func WithPriority(priority uint) WorkflowOption ``` Set a queue priority for the workflow. Should be used alongside `WithQueue`. Workflows with the same priority are dequeued in **FIFO (first in, first out)** order. Priority values can range from `1` to `2,147,483,647`, where **a low number indicates a higher priority**. Workflows without assigned priorities have the highest priority and are dequeued before workflows with assigned priorities. ##### WithQueuePartitionKey ```go func WithQueuePartitionKey(partitionKey string) WorkflowOption ``` Set a queue partition key for the workflow. Use if and only if the queue is [partitioned](../tutorials/queue-tutorial.md#partitioning-queues) (registered with at least one partition limit, such as [`WithPartitionConcurrency`](./queues.md#withpartitionconcurrency)). A partitioned queue applies its partition limits to each partition separately, while its global concurrency, worker concurrency, and rate limit still apply across all partitions. **Example Syntax:** ```go // Create a partitioned queue: at most one workflow per partition runs at once partitionedQueue, err := dbos.RegisterQueue(ctx, "user-tasks", dbos.WithPartitionConcurrency(1), ) // Enqueue workflows with partition keys // At most one task per user runs at once, but tasks from different users run concurrently handle, err := dbos.RunWorkflow(ctx, ProcessUserTask, taskData, dbos.WithQueue(partitionedQueue), dbos.WithQueuePartitionKey(userID), ) ``` :::info - Partition keys are required when enqueueing to a partitioned queue. - Partition keys cannot be used with non-partitioned queues. - Partition keys and deduplication IDs cannot be used together. ::: ##### WithDelay ```go func WithDelay(delay time.Duration) WorkflowOption ``` Delay execution of a queued workflow by the specified duration. Must be used together with [`WithQueue`](#withqueue). The workflow is initially placed in `DELAYED` status and does not execute. After the delay expires, it transitions to `ENQUEUED` status and may be dequeued and executed. This is useful for scheduling a workflow to run at a future time. You can dynamically update or shorten the delay of a `DELAYED` workflow with [`SetWorkflowDelay`](./methods.md#setworkflowdelay). ```go remindersQueue, err := dbos.RegisterQueue(ctx, "reminders") // Run the reminder workflow one hour from now. handle, err := dbos.RunWorkflow(ctx, sendReminder, userID, dbos.WithQueue(remindersQueue), dbos.WithDelay(1 * time.Hour), ) ``` ##### WithPortableWorkflow ```go func WithPortableWorkflow() WorkflowOption ``` Mark the workflow to use the [portable JSON serialization format](../../explanations/portable-workflows.md) for cross-language interoperability. When set, workflow inputs, outputs, and errors are serialized using portable JSON so they can be read by applications in other languages. A DBOS Go portable workflow inputs will be type-asserted into the workflow input type. ```go handle, err := dbos.RunWorkflow(dbosContext, processOrder, "order-123", dbos.WithPortableWorkflow(), ) ``` ##### WithApplicationVersion ```go func WithApplicationVersion(version string) WorkflowOption ``` Set the application version for this workflow, overriding the version in Context. ##### WithAuthenticatedUser ```go func WithAuthenticatedUser(user string) WorkflowOption ``` Associate the workflow execution with a user name. Useful to define workflow identity. Child workflows automatically inherit their parent's authentication information (authenticated user, assumed role, and authenticated roles) unless explicitly overridden. ##### WithAssumedRole ```go func WithAssumedRole(role string) WorkflowOption ``` Set the assumed role recorded on the workflow. ##### WithAuthenticatedRoles ```go func WithAuthenticatedRoles(roles ...string) WorkflowOption ``` Set the authenticated roles recorded on the workflow. ##### WithWorkflowAttributes ```go func WithWorkflowAttributes(attributes map[string]any) WorkflowOption ``` Attach custom key-value attributes to the workflow. Attributes are recorded in the [workflow status](./methods.md#workflow-status) at creation, must be JSON-serializable, and are not inherited by child workflows. On Postgres they are stored as GIN-indexed JSONB and can be searched with [`WithFilterAttributes`](./methods.md#withfilterattributes). Attributes can later be replaced with [`SetWorkflowAttributes`](./methods.md#setworkflowattributes). ```go handle, err := dbos.RunWorkflow(ctx, processOrder, order, dbos.WithWorkflowAttributes(map[string]any{"customer": "acme", "region": "us-east"}), ) ``` #### RunAsStep ```go func RunAsStep[R any](ctx Context, fn Step[R], opts ...StepOption) (R, error) ``` Execute a function as a step in a durable workflow. **Parameters:** - **ctx**: The Context. - **fn**: The step to execute, typically wrapped in an anonymous function. Syntax shown below. - **opts**: Functional options for step execution, documented below. **Example Syntax:** Any Go function can be a step as long as it outputs one [json-encodable](https://pkg.go.dev/encoding/json) value and an error. To pass inputs into a function being called as a step, wrap it in an anonymous function as shown below: ```go func step(ctx context.Context, input string) (string, error) { output := ... return output } func workflow(ctx dbos.Context, input string) (string, error) { output, err := dbos.RunAsStep( ctx, func(stepCtx context.Context) (string, error) { return step(stepCtx, input) } ) } ``` ##### WithStepName ```go func WithStepName(name string) StepOption ``` Set a custom name for a step. ##### WithStepMaxRetries ```go func WithStepMaxRetries(maxRetries int) StepOption ``` Set the maximum number of times this step is automatically retried on failure. A value of 0 (the default) indicates no retries. ##### WithStepMaxInterval ```go func WithStepMaxInterval(interval time.Duration) StepOption ``` WithStepMaxInterval sets the maximum delay between retries. Default value is 5s. ##### WithStepBackoffFactor ```go func WithStepBackoffFactor(factor float64) StepOption ``` WithStepBackoffFactor sets the exponential backoff multiplier between retries. Default value is 2.0. ##### WithStepBaseInterval ```go func WithStepBaseInterval(interval time.Duration) StepOption ``` WithStepBaseInterval sets the initial delay between retries. Default value is 100ms. ##### WithStepRetryPredicate ```go func WithStepRetryPredicate(predicate func(error) bool) StepOption ``` Set a predicate deciding whether a step error is retried. When set, a failed step is retried only if the predicate returns true for the error; otherwise the error is returned immediately, regardless of the remaining retry budget. #### Calling DBOS operations from steps A step body is a strict scope: DBOS operations that write a checkpoint cannot run inside one. Calling any of the following from inside a step returns an `ErrorCodeStepExecution` error: [`RunWorkflow`](#runworkflow) (spawning a child workflow), [`RunAsTransaction`](./datasources.md#runastransaction), [`Go`](#go), [`Enqueue`](./methods.md#enqueue), [`Send`](./methods.md#send), [`Recv`](./methods.md#recv), [`GetEvent`](./methods.md#getevent), [`Sleep`](./methods.md#sleep), [`CloseStream`](./methods.md#closestream), [`GetResult`](#workflowhandlegetresult) on a workflow handle, [`Patch`](#patch), [`DeprecatePatch`](#deprecatepatch), [`Debounce`](./queues.md#debouncerdebounce), workflow-management writes ([`CancelWorkflow(s)`](./methods.md#cancelworkflow), [`ResumeWorkflow(s)`](./methods.md#resumeworkflow), [`ForkWorkflow(s)`](./methods.md#forkworkflow), [`DeleteWorkflows`](./methods.md#deleteworkflows), [`SetWorkflowAttributes`](./methods.md#setworkflowattributes), [`SetWorkflowDelay`](./methods.md#setworkflowdelay)), and schedule writes ([`CreateSchedule`](./methods.md#createschedule), [`PauseSchedule`](./methods.md#pauseschedule), [`ResumeSchedule`](./methods.md#resumeschedule), [`DeleteSchedule`](./methods.md#deleteschedule)). Allowed from inside a step: - Calling another step function — it runs inline as part of the enclosing step, without its own checkpoint. - [`SetEvent`](./methods.md#setevent) and [`WriteStream`](./methods.md#writestream) — at-least-once, attributed to the enclosing step. - Read and list operations ([`ListWorkflows`](./methods.md#listworkflows), [`GetWorkflowSteps`](./methods.md#getworkflowsteps), [`RetrieveWorkflow`](./methods.md#retrieveworkflow), [`ReadStream`](./methods.md#readstream), [aggregates](./methods.md#getworkflowaggregates), [schedule](./methods.md#getschedule) and [queue](./queues.md#listqueues) reads) — they run within the enclosing step's durability scope, without their own checkpoint. #### Concurrent steps ##### Go ```go func Go[R any](ctx Context, fn Step[R], opts ...StepOption) (<-chan StepOutcome[R], error) ``` Launch a step asynchronously and return a channel that will receive the result when the step completes. This is a durable alternative to Go's native goroutines that checkpoints results to the database for deterministic replay. Can only be called from within a workflow (not from inside a step). **Parameters:** - **ctx**: The Context. - **fn**: The step function to execute asynchronously. - **opts**: Functional options for step execution (same options as [`RunAsStep`](#runasstep)). **Returns:** - A receive-only channel of `StepOutcome[R]` that will receive exactly one value when the step completes, then close. - An error if the step could not be launched (e.g., if called outside a workflow). ```go type StepOutcome[R any] struct { Result R Err error } ``` **Example Syntax:** ```go func workflow(ctx dbos.Context, _ string) (string, error) { // Launch step asynchronously resultChan, err := dbos.Go(ctx, func(ctx context.Context) (string, error) { return performWork(ctx) }) if err != nil { return "", err } // Do other work... // Wait for result outcome := <-resultChan if outcome.Err != nil { return "", outcome.Err } return outcome.Result, nil } ``` ##### Select ```go func Select[R any](ctx Context, channels []<-chan StepOutcome[R]) (R, error) ``` Wait for and return the first result from multiple channels obtained from [`Go`](#go). This is a durable alternative to Go's native `select` statement that checkpoints the selected result for deterministic replay. Can only be called from within a workflow. All channels must be of the same type `R`. **Parameters:** - **ctx**: The Context. - **channels**: A slice of receive-only channels from [`Go`](#go) calls. **Returns:** - The result value from the first channel to produce a value. - An error if the selected step returned an error, if the context was cancelled, or if a channel was closed unexpectedly. **Behavior:** - If `channels` is empty, returns the zero value of type `R` with no error. - If the context is cancelled while waiting, returns the context error. - The selected channel index and value are checkpointed, so workflow recovery returns the same result. **Example Syntax:** ```go func workflow(ctx dbos.Context, _ string) (string, error) { ch1, _ := dbos.Go(ctx, func(ctx context.Context) (string, error) { return queryServiceA(ctx) }) ch2, _ := dbos.Go(ctx, func(ctx context.Context) (string, error) { return queryServiceB(ctx) }) // Wait for the first result result, err := dbos.Select(ctx, []<-chan dbos.StepOutcome[string]{ch1, ch2}) if err != nil { return "", err } return result, nil } ``` #### WorkflowHandle ```go type WorkflowHandle[R any] interface { GetResult(opts ...GetResultOption) (R, error) GetStatus() (WorkflowStatus, error) GetWorkflowID() string } ``` WorkflowHandle provides methods to interact with a running or completed workflow. The type parameter `R` represents the expected return type of the workflow. Handles can be used to wait for workflow completion, check status, and retrieve results. ##### WorkflowHandle.GetResult ```go WorkflowHandle.GetResult(opts ...GetResultOption) (R, error) ``` Wait for the workflow to complete and return its result. ###### WithHandleTimeout ```go func WithHandleTimeout(timeout time.Duration) GetResultOption ``` Specify a timeout for obtaining a workflow result. On expiry, the error matches `context.DeadlineExceeded`. ###### WithHandlePollingInterval ```go func WithHandlePollingInterval(interval time.Duration) GetResultOption ``` Set the polling interval for checking workflow completion status in the database. Only positive interval values will be considered. ##### WorkflowHandle.GetStatus ```go WorkflowHandle.GetStatus() (WorkflowStatus, error) ``` Retrieve the WorkflowStatus of the workflow. ##### WorkflowHandle.GetWorkflowID ```go WorkflowHandle.GetWorkflowID() string ``` Retrieve the ID of the workflow. #### Patching ##### Patch ```go func Patch(ctx Context, patchName string) (bool, error) ``` Insert a patch marker at the current point in workflow history, returning `true` if it was successfully inserted and `false` if there is already a checkpoint present at this point in history indicating that the workflow should run unpatched. Used to safely upgrade workflow code; see the [patching tutorial](../tutorials/upgrading-workflows.md#patching) for more detail. **Parameters:** - **ctx**: The Context. - **patchName**: The name to give the patch marker that will be inserted into workflow history. :::info Patching must be enabled in your configuration by setting `EnablePatching: true`. ::: ##### DeprecatePatch ```go func DeprecatePatch(ctx Context, patchName string) error ``` Safely bypass a patch marker at the current point in workflow history if present. Used to safely deprecate patches; see the [patching tutorial](../tutorials/upgrading-workflows.md#patching) for more detail. **Parameters:** - **ctx**: The Context. - **patchName**: The name of the patch marker to be bypassed. :::info Patching must be enabled in your configuration by setting `EnablePatching: true`. ::: #### Serialization Workflow and step inputs, outputs, and errors are checkpointed in the system database. By default they are encoded with a JSON serializer; you can supply a custom serializer via `Config.Serializer`: ```go type Serializer[T any] interface { // Name returns the name of the serialization format (e.g., "DBOS_JSON", "DBOS_GOB"). Name() string // Encode serializes a value to a string representation for database storage. Encode(data T) (*string, error) // Decode deserializes a string from the database back into a value. Decode(data *string) (T, error) } ``` A built-in [gob](https://pkg.go.dev/encoding/gob) serializer is available with `dbos.NewGobSerializer()`, which handles arbitrary Go types (you must register your program's concrete types with `gob.Register`). ##### Custom serializers must be total Every checkpoint written to the system database is encoded with the configured serializer — including engine-internal step outputs, not just your workflow inputs and outputs: - the `int64` wake-up deadline recorded by durable [`Sleep`](./methods.md#sleep) (also written by `Recv` and `GetEvent` timeouts), - the empty-string step output recorded by `WriteStream`, - the `ScheduledWorkflowInput` struct recorded for database-backed schedule firings. This keeps every row in the system database in one format, so it can be decoded uniformly by external tooling. A `Serializer[any]` must therefore encode and decode *arbitrary* Go values, and decoding must preserve concrete types: decoded step results are type-asserted, so an `int64` must round-trip as `int64`, not `float64`. See the [`protobuf-serializer` demo app](https://github.com/dbos-inc/dbos-demo-apps/tree/main/golang/protobuf-serializer) for a detailed example. Additionally, every implementation must honor this contract: 1. `Decode` can be called with a nil `*string`: some checkpoints record an error and never write an output, so the stored value is SQL `NULL`. `Decode` must tolerate nil input (typically returning the zero value). 2. The nil round-trip must be lossless: `Decode(Encode(nil-value))` must yield that nil value back. 3. `Encode` must not return a nil `*string`; to represent nil data, return a pointer to a sentinel string of the serializer's choosing (the built-in gob serializer stores `__DBOS_NIL`; the portable JSON serializer stores `null`). The engine never interprets the sentinel: nil detection always goes through the serializer that wrote the row, so no string is reserved across formats. One deliberate exception: `Recv` and `GetEvent` checkpoint the *sender's* encoded payload verbatim, under the sender's recorded format. The receiver's serializer is never asked to re-encode a message or event it did not produce. #### Errors Errors produced by DBOS APIs are of type `*dbos.Error`: ```go type Error struct { Message string // Human-readable error message Code ErrorCode // Error type code for programmatic handling // Optional context fields — only set when relevant to the error WorkflowID string // Associated workflow identifier DestinationID string // Target workflow identifier (for communication errors) StepName string // Step function name (for step errors) QueueName string // Queue name (for queue-related errors) DeduplicationID string // Deduplication identifier StepID int // Step sequence number ExpectedName string // Expected function name (for determinism errors) RecordedName string // Actually recorded function name (for determinism errors) MaxRetries int // Maximum retry limit (for retry-related errors) } ``` `Error` implements the standard error interface, formatting messages as `DBOS Error : `. ##### Matching errors Prefer matching with `errors.Is` against the package sentinels. Matching is by error code — a sentinel matches any DBOS error carrying the same code, regardless of its other fields: ```go handle, err := dbos.RunWorkflow(ctx, wf, input, dbos.WithWorkflowID(id)) if errors.Is(err, dbos.ErrConflictingWorkflowID) { ... } ``` To read the structured fields, use `errors.As`: ```go var dbosErr *dbos.Error if errors.As(err, &dbosErr) { logger.Error("workflow failed", "workflow_id", dbosErr.WorkflowID, "code", dbosErr.Code) } ``` DBOS errors caused by context cancellation or an expired deadline also wrap the standard library cause, so `errors.Is(err, context.Canceled)` and `errors.Is(err, context.DeadlineExceeded)` match as well. These causes survive serialization: they still match after the error is read back from the system database, for example when awaiting a workflow from another process. Sentinel matching also survives the [portable JSON format](../../explanations/portable-workflows.md): a DBOS error recorded portably keeps its symbolic code, so `errors.Is` against the sentinels still holds when reading it from another process or language runtime. ##### Error codes :::warning Never swallow a DBOS error inside a workflow without knowing which kind you have. The rule is: **handle an error only if DBOS already checkpointed its outcome.** A checkpointed outcome is replayed verbatim, so a branch you take on it is deterministic. An error that signals this execution has lost ownership of the workflow, or has diverged from its recorded history, must be returned from the workflow function. ::: The **In a workflow** column classifies each code: - **Return** — propagate it out of your workflow function. Either DBOS's own handler needs the error to resolve the situation, or the workflow's recorded history can no longer be trusted. - **Handle** — you may catch it and continue. Take extra care where the Meaning column notes that the failure left no checkpoint. - **Check the cause** — a wrapper code; classify by the error it wraps. - **—** — a configuration, registration, or management error that does not arise from workflow execution. Classify with `errors.Is`, not the `Code` field of the outermost error: DBOS wraps some causes, so a failed child start surfaces as `ErrorCodeWorkflowExecution` while still matching `ErrQueueDeduplicated` or `ErrDeadLetterQueue`. | Code | Sentinel | In a workflow | Meaning | |---|---|---|---| | `ErrorCodeConflictingID` | `ErrConflictingWorkflowID` | Return | A concurrent execution of the same workflow recorded a conflicting step checkpoint, or an operation reused a workflow ID already in use. DBOS handles this error at the workflow level: returning it parks your execution until the winning one settles, and it adopts that outcome. Swallowing it makes the two executions race step by step. See [Concurrent Executions](../../explanations/concurrent-executions.md). | | `ErrorCodeInitialization` | — | — | The DBOS context could not be initialized (invalid configuration, system database unreachable, or migrations failed). | | `ErrorCodeNonExistentWorkflow` | `ErrNonExistentWorkflow` | Handle | The referenced workflow does not exist (e.g., `RetrieveWorkflow` or a management method with an unknown ID). | | `ErrorCodeUnexpectedWorkflow` | `ErrUnexpectedWorkflow` | Return | A workflow ID was reused by a different workflow function or on a different queue, indicating non-determinism or conflicting ID reuse. Continuing would write your checkpoints into another workflow's history. | | `ErrorCodeWorkflowCancelled` | `ErrWorkflowCancelled` | Return | This workflow was cancelled during execution. Once cancelled, nothing more is checkpointed: steps are interrupted without recording an outcome, so anything you run after swallowing this error is a non-durable side effect that re-executes on resume. Match it with `errors.Is` (step wrappers may enclose it). When the cancellation came from a cancelled context or expired durable timeout, the wrapped cause also matches `context.Canceled` / `context.DeadlineExceeded`; an API `CancelWorkflow` carries no stdlib cause. See [workflow cancellation](../tutorials/workflow-management.md#cancelling-workflows). | | `ErrorCodeUnexpectedStep` | — | Return | During replay, the step executing at a position differs from the step recorded there: the workflow function is non-deterministic or its code changed. See [Upgrading Workflow Code](../tutorials/upgrading-workflows.md). | | `ErrorCodeAwaitedWorkflowCancelled` | `ErrAwaitedWorkflowCancelled` | Handle | A workflow you were awaiting (via a handle, `GetEvent`, or a child workflow) was cancelled. Unlike `ErrorCodeWorkflowCancelled`, this is the *awaiter's* error: it is checkpointed as the await's outcome, so the caller may handle it and continue. | | `ErrorCodeConflictingRegistration` | — | — | A workflow, queue, or schedule was registered under a name that is already registered (or reserved). | | `ErrorCodeWorkflowUnexpectedType` | — | Return | A recorded input or output could not be decoded into the requested type parameter (e.g., `RetrieveWorkflow[R]` with the wrong `R`). The recorded history does not match the code reading it. | | `ErrorCodeWorkflowExecution` | — | Check the cause | General workflow execution failure, and the wrapper DBOS uses for failures while starting a child workflow. Classify by its wrapped cause with `errors.Is`, not by this code. | | `ErrorCodeStepExecution` | — | Handle | General step execution failure. The step's error is checkpointed and returned verbatim on replay. | | `ErrorCodeDeadLetterQueue` | `ErrDeadLetterQueue` | Handle | The workflow exceeded its maximum recovery attempts (`WithMaxRecoveryAttempts`) and was dead-lettered. Awaiting a dead-lettered workflow checkpoints the error as the await's outcome and is safe to handle; *starting* one does not checkpoint, so treat that case like a rejected enqueue. | | `ErrorCodeMaxStepRetriesExceeded` | `ErrMaxStepRetriesExceeded` | Handle | A step exhausted its configured retries. The error is checkpointed as the step's outcome and wraps the joined errors of all attempts, so `errors.Is`/`errors.As` can reach the underlying failures. | | `ErrorCodeQueueDeduplicated` | `ErrQueueDeduplicated` | Handle | An enqueue was rejected because another workflow with the same deduplication ID is already pending on the queue. The rejected enqueue is rolled back, so no checkpoint records it. | | `ErrorCodePatchingNotEnabled` | — | — | `Patch` or `DeprecatePatch` was called but `Config.EnablePatching` is false. | | `ErrorCodeTimeout` | `ErrTimeout` | Handle | A DBOS wait timed out (e.g., `Recv`, `GetEvent`, or an in-memory handle's `GetResult`). Inside a workflow, `Recv` and `GetEvent` timeouts are checkpointed as the operation's outcome. Deadline-driven timeouts also match `context.DeadlineExceeded`; for handle waits, prefer matching `context.DeadlineExceeded`, which covers all handle flavors. | | `ErrorCodeNoApplicationVersions` | `ErrNoApplicationVersions` | — | An operation required a registered application version, but none exists in the system database. | | `ErrorCodeQueueNotFound` | `ErrQueueNotFound` | Handle | The referenced queue does not exist (e.g., `RetrieveQueue`). | | `ErrorCodeScheduleNotFound` | `ErrScheduleNotFound` | Handle | The referenced schedule does not exist (e.g., `GetSchedule`). | | `ErrorCodeInvalidOption` | `ErrInvalidOption` | — | Invalid or inconsistent options were passed to a DBOS API. | --- ## Queues & Concurrency You can use queues to run many workflows at once with managed concurrency. Queues provide _flow control_, letting you manage how many workflows run at once or how often workflows are started. To create a queue, register it with [`RegisterQueue`](../reference/queues#registerqueue): ```go queue, err := dbos.RegisterQueue(dbosContext, "example_queue") ``` `RegisterQueue` persists the queue's configuration to the system database. It can be called at any time, including after [`Launch`](../reference/dbos-context.md#launch), and the queue's configuration can be [changed at runtime](#reconfiguring-queues). If multiple applications [share a system database](../../explanations/sharing-a-system-database.md), each queue is owned by the application that registers it, and only that application dequeues workflows from it. `RegisterQueue` returns a `Queue` handle. You can then enqueue any workflow by passing the handle to [`WithQueue`](../reference/workflows-steps#withqueue) when calling `RunWorkflow`. You can fetch a queue handle with [`RetrieveQueue`](../reference/queues.md#retrievequeue). Enqueuing a function submits it for execution and returns a [handle](../reference/workflows-steps#workflowhandle) to it. Queued tasks are started in first-in, first-out (FIFO) order. ```go func processTask(ctx dbos.Context, task string) (string, error) { // Process the task... return fmt.Sprintf("Processed: %s", task), nil } func example(dbosContext dbos.Context, queue dbos.Queue) error { // Enqueue a workflow task := "example_task" handle, err := dbos.RunWorkflow(dbosContext, processTask, task, dbos.WithQueue(queue)) if err != nil { return err } // Get the result result, err := handle.GetResult() if err != nil { return err } fmt.Println("Task result:", result) return nil } ``` #### Queue Example Here's an example of a workflow using a queue to process tasks concurrently: ```go func taskWorkflow(ctx dbos.Context, task string) (string, error) { // Process the task... return fmt.Sprintf("Processed: %s", task), nil } func queueWorkflow(ctx dbos.Context, queueName string) ([]string, error) { // Look up the queue handle by name queue, err := dbos.RetrieveQueue(ctx, queueName) if err != nil { return nil, err } // Enqueue each task so all tasks are processed concurrently tasks := []string{"task1", "task2", "task3", "task4", "task5"} var handles []dbos.WorkflowHandle[string] for _, task := range tasks { handle, err := dbos.RunWorkflow(ctx, taskWorkflow, task, dbos.WithQueue(queue)) if err != nil { return nil, fmt.Errorf("failed to enqueue task %s: %w", task, err) } handles = append(handles, handle) } // Wait for each task to complete and retrieve its result var results []string for i, handle := range handles { result, err := handle.GetResult() if err != nil { return nil, fmt.Errorf("task %d failed: %w", i, err) } results = append(results, result) } return results, nil } func example(dbosContext dbos.Context) error { handle, err := dbos.RunWorkflow(dbosContext, queueWorkflow, "example_queue") if err != nil { return err } results, err := handle.GetResult() if err != nil { return err } for _, result := range results { fmt.Println(result) } return nil } ``` Sometimes, you may wish to receive the result of each task as soon as it's ready instead of waiting for all tasks to complete. You can do this using [`Send` and `Recv`](./workflow-communication.md#workflow-messaging-and-notifications). Each enqueued workflow sends a message to the main workflow when it's done processing its task. The main workflow awaits those messages, retrieving the result of each task as soon as the task completes. ```go const TaskCompleteTopic = "task_complete" // The queue handle returned by RegisterQueue, e.g. stored in a package-level variable var queue dbos.Queue type TaskInput struct { ParentWorkflowID string TaskID int Task string } func processTask(ctx dbos.Context, input TaskInput) (string, error) { result := ... // Process the task // Notify the main workflow this task is complete err := dbos.Send(ctx, input.ParentWorkflowID, input.TaskID, TaskCompleteTopic) if err != nil { return "", fmt.Errorf("failed to send completion notification: %w", err) } return result, nil } func processTasks(ctx dbos.Context, tasks []string) ([]string, error) { parentWorkflowID, err := dbos.GetWorkflowID(ctx) if err != nil { return nil, fmt.Errorf("failed to get workflow ID: %w", err) } var handles []dbos.WorkflowHandle[string] for i, task := range tasks { handle, err := dbos.RunWorkflow(ctx, processTask, TaskInput{ParentWorkflowID: parentWorkflowID, TaskID: i, Task: task}, dbos.WithQueue(queue), ) if err != nil { return nil, fmt.Errorf("failed to enqueue task %d: %w", i, err) } handles = append(handles, handle) } var results []string for len(results) < len(tasks) { // Wait for a notification that a task is complete completedTaskID, err := dbos.Recv[int](ctx, TaskCompleteTopic, 5*time.Minute) if err != nil { return nil, fmt.Errorf("timeout waiting for task completion: %w", err) } // Retrieve result of the completed task completedTaskHandle := handles[completedTaskID] result, err := completedTaskHandle.GetResult() if err != nil { return nil, fmt.Errorf("task %d failed: %w", completedTaskID, err) } fmt.Printf("Task %d completed. Result: %s\n", completedTaskID, result) results = append(results, result) } return results, nil } ``` #### Enqueueing from Another Application Often, you want to enqueue a workflow from outside your DBOS application. For example, let's say you have an API server and a data processing service. You're using DBOS to build a durable data pipeline in the data processing service. When the API server receives a request, it should enqueue the data pipeline for execution on the data processing service. You can use the [DBOS Client](../reference/dbos-context.md#newclient) to enqueue workflows from outside your DBOS application by connecting directly to your DBOS application's system database. Since the DBOS Client is designed to be used from outside your DBOS application, workflow and queue metadata must be specified explicitly. For example, this code enqueues the `dataPipeline` workflow on the `pipelineQueue` queue with a `ProcessInput` argument: ```go type ProcessInput struct { TaskID string Data string } type ProcessOutput struct { Result string Status string } config := dbos.ClientConfig{ DatabaseURL: os.Getenv("DBOS_SYSTEM_DATABASE_URL"), // The name of the application that runs the data pipeline AppName: "data-processing-service", } client, err := dbos.NewClient(context.Background(), config) if err != nil { log.Fatal(err) } defer dbos.Shutdown(client, 5*time.Second) handle, err := dbos.Enqueue[ProcessOutput]( client, "pipelineQueue", "dataPipeline", ProcessInput{TaskID: "task-123", Data: "data"}, ) if err != nil { log.Fatal(err) } ``` If your application tables live in the same database as the DBOS system schema, you can enqueue a workflow **atomically** with your own writes by passing your open transaction with [`WithEnqueueTransaction`](../reference/methods.md#withenqueuetransaction). The workflow is enqueued only when you commit; if you roll back, it never existed: ```go tx, err := pool.Begin(ctx) if err != nil { log.Fatal(err) } defer tx.Rollback(ctx) _, err = tx.Exec(ctx, "INSERT INTO tasks (id, data) VALUES ($1, $2)", "task-123", "data") if err != nil { log.Fatal(err) } handle, err := dbos.Enqueue[ProcessOutput]( client, "pipelineQueue", "dataPipeline", ProcessInput{TaskID: "task-123", Data: "data"}, dbos.WithEnqueueTransaction(tx), ) if err != nil { log.Fatal(err) } if err := tx.Commit(ctx); err != nil { log.Fatal(err) } ``` [`Send`](../reference/methods.md#send) accepts a transaction the same way, with [`WithSendTransaction`](../reference/methods.md#withsendtransaction). #### Managing Concurrency You can control how many workflows from a queue run simultaneously by configuring concurrency limits. This helps prevent resource exhaustion when workflows consume significant memory or processing power. ##### Worker Concurrency Worker concurrency sets the maximum number of workflows from a queue that can run concurrently on a single DBOS process. This is particularly useful for resource-intensive workflows to avoid exhausting the resources of any process. For example, this queue has a worker concurrency of 5, so each process will run at most 5 workflows from this queue simultaneously: ```go queue, err := dbos.RegisterQueue(dbosContext, "example_queue", dbos.WithWorkerConcurrency(5)) ``` ##### Global Concurrency Global concurrency limits the total number of workflows from a queue that can run concurrently across all DBOS processes in your application. For example, this queue will have a maximum of 10 workflows running simultaneously across your entire application. :::warning Worker concurrency limits are recommended for most use cases. Take care when using a global concurrency limit as any `PENDING` workflow on the queue counts toward the limit, including workflows from previous application versions ::: ```go queue, err := dbos.RegisterQueue(dbosContext, "example_queue", dbos.WithGlobalConcurrency(10)) ``` #### Rate Limiting You can set _rate limits_ for a queue, limiting the number of functions that it can start in a given period. Rate limits are global across all DBOS processes using this queue. For example, this queue has a limit of 100 workflows with a period of 60 seconds, so it may not start more than 100 workflows in 60 seconds: ```go queue, err := dbos.RegisterQueue(dbosContext, "example_queue", dbos.WithRateLimiter(&dbos.RateLimiter{ Limit: 100, Period: 60 * time.Second, })) ``` Rate limits are especially useful when working with a rate-limited API, such as many LLM APIs. #### Reconfiguring Queues Because queue configuration is persisted to the system database, you can change a queue's configuration at runtime without redeploying or restarting your workers. Workers pick up the new configuration on their next polling iteration. Use [`RetrieveQueue`](../reference/queues.md#retrievequeue) to fetch a queue, then call its [`Set*`](../reference/queues.md#reconfiguring-queues) methods. ```go queue, err := dbos.RetrieveQueue(dbosContext, "example_queue") if err != nil { return err } concurrency := 50 err = queue.SetGlobalConcurrency(dbosContext, &concurrency) ``` :::warning If your application calls [`RegisterQueue`](../reference/queues.md#registerqueue) on startup, the next process to start can overwrite settings you applied at runtime via `Set*` methods. Either update the `RegisterQueue` call to match the new configuration, or pass `WithQueueOnConflict(dbos.QueueConflictNeverUpdate)` to preserve the runtime changes. ::: #### Deduplication You can set a deduplication ID for an enqueued workflow using [`WithDeduplicationID`](../reference/workflows-steps#withdeduplicationid) when calling `RunWorkflow`. At any given time, only one workflow with a specific deduplication ID can be enqueued in the specified queue. If a workflow with a deduplication ID is currently enqueued or actively executing (status `ENQUEUED` or `PENDING`), subsequent workflow enqueue attempts with the same deduplication ID in the same queue will return an error. Alternatively, use [`WithDeduplicationPolicy(dbos.DeduplicationPolicyReturnExisting)`](../reference/workflows-steps#withdeduplicationpolicy) to instead return a handle to the existing workflow holding the deduplication ID. For example, this is useful if you only want to have one workflow active at a time per user—set the deduplication ID to the user's ID. **Example syntax:** ```go func taskWorkflow(ctx dbos.Context, task string) (string, error) { // Process the task... return "completed", nil } func example(dbosContext dbos.Context, queue dbos.Queue) error { task := "example_task" deduplicationID := "user_12345" // Use user ID for deduplication handle, err := dbos.RunWorkflow( dbosContext, taskWorkflow, task, dbos.WithQueue(queue), dbos.WithDeduplicationID(deduplicationID)) if err != nil { // Handle deduplication error or other failures return fmt.Errorf("failed to enqueue workflow: %w", err) } result, err := handle.GetResult() if err != nil { return fmt.Errorf("workflow failed: %w", err) } fmt.Printf("Workflow completed: %s\n", result) return nil } ``` #### Priority You can set a priority for an enqueued workflow using [`WithPriority`](../reference/workflows-steps#withpriority) when calling `RunWorkflow`. Workflows with the same priority are dequeued in **FIFO (first in, first out)** order. Priority values can range from `1` to `2,147,483,647`, where **a low number indicates a higher priority**. If using priority, you must set [`WithPriorityEnabled`](../reference/queues#withpriorityenabled) on your queue. :::tip Workflows without assigned priorities have the highest priority and are dequeued before workflows with assigned priorities. ::: To use priorities in a queue, you must enable it when creating the queue: ```go queue, err := dbos.RegisterQueue(dbosContext, "example_queue", dbos.WithPriorityEnabled()) ``` **Example syntax:** ```go func taskWorkflow(ctx dbos.Context, task string) (string, error) { // Process the task... return "completed", nil } func example(dbosContext dbos.Context, queue dbos.Queue) error { task := "example_task" priority := uint(10) // Lower number = higher priority handle, err := dbos.RunWorkflow(dbosContext, taskWorkflow, task, dbos.WithQueue(queue), dbos.WithPriority(priority)) if err != nil { return err } result, err := handle.GetResult() if err != nil { return fmt.Errorf("workflow failed: %w", err) } fmt.Printf("Workflow completed: %s\n", result) return nil } ``` #### Partitioning Queues You can **partition** queues to distribute work across dynamically created queue partitions. A queue is partitioned if you register it with any per-partition flow control limit: | Option | Meaning | | --- | --- | | [`WithPartitionConcurrency`](../reference/queues.md#withpartitionconcurrency) | Maximum workflows from any one partition running at once across all processes. | | [`WithPartitionWorkerConcurrency`](../reference/queues.md#withpartitionworkerconcurrency) | Maximum workflows from any one partition running at once on a single process. | | [`WithPartitionRateLimiter`](../reference/queues.md#withpartitionratelimiter) | Maximum workflows that may be started from any one partition in a given period. | When you enqueue a workflow on a partitioned queue, you must supply a queue partition key. Essentially, you can think of each partition as a "subqueue" you dynamically create by enqueueing a workflow with a partition key. For example, suppose you want your users to each be able to run at most one task at a time. You can do this with a queue whose partition concurrency is 1, where the partition key is user ID. **Example Syntax:** ```go // Create a partitioned queue with a per-partition concurrency limit of 1 partitionedQueue, err := dbos.RegisterQueue(dbosContext, "user-tasks", dbos.WithPartitionConcurrency(1), ) type Task struct { TaskID string Data string } func processTask(ctx dbos.Context, task Task) (string, error) { // Process the task... return fmt.Sprintf("Processed task %s", task.TaskID), nil } func onUserTaskSubmission(dbosContext dbos.Context, userID string, task Task) error { // Partition the task queue by user ID. As the queue has a // per-partition concurrency of 1, this means that at most one // task can run at once per user (but tasks from different // users can run concurrently). handle, err := dbos.RunWorkflow(dbosContext, processTask, task, dbos.WithQueue(partitionedQueue), dbos.WithQueuePartitionKey(userID), ) if err != nil { return fmt.Errorf("failed to enqueue task: %w", err) } result, err := handle.GetResult() if err != nil { return fmt.Errorf("task failed: %w", err) } fmt.Printf("Task completed: %s\n", result) return nil } ``` :::info - Partition keys are required when enqueueing to a partitioned queue. - Partition keys cannot be used with non-partitioned queues. - Partition keys and deduplication IDs cannot be used together. ::: :::warning Every enqueue on a partitioned queue must supply a partition key. A workflow enqueued on a partitioned queue without a partition key stays `ENQUEUED` and is not dequeued. This also applies to workflows that were enqueued before a queue was partitioned at runtime through its [`Set*`](../reference/queues.md#reconfiguring-queues) methods. ::: ##### Combining Queue-Wide and Per-Partition Limits A partitioned queue enforces its per-partition limits **and** its queue-wide limits ([`WithGlobalConcurrency`](#global-concurrency), [`WithWorkerConcurrency`](#worker-concurrency), and [`WithRateLimiter`](#rate-limiting)) at the same time. This lets you protect your workers from overload while still fairly distributing work between partitions. For example, this "fair queue" runs at most one task per user, but no more than 10 tasks on any single process: ```go fairQueue, err := dbos.RegisterQueue(dbosContext, "fair-queue", dbos.WithPartitionConcurrency(1), dbos.WithWorkerConcurrency(10), ) ``` Each queue-wide limit has a per-partition counterpart, so you can mix and match them freely: ```go // At most 100 tasks running globally and 25 running per tenant, // at most 10 tasks running per process and 2 per tenant per process, // and at most 1000 tasks started per minute globally and 50 per tenant. tenantQueue, err := dbos.RegisterQueue(dbosContext, "tenant-queue", dbos.WithGlobalConcurrency(100), dbos.WithWorkerConcurrency(10), dbos.WithRateLimiter(&dbos.RateLimiter{Limit: 1000, Period: 60 * time.Second}), dbos.WithPartitionConcurrency(25), dbos.WithPartitionWorkerConcurrency(2), dbos.WithPartitionRateLimiter(&dbos.RateLimiter{Limit: 50, Period: 60 * time.Second}), ) ``` Each per-partition concurrency limit must be less than or equal to its queue-wide counterpart, and the partition worker concurrency must be less than or equal to the partition concurrency. :::note Queues registered with the deprecated [`WithPartitionQueue`](../reference/queues.md#withpartitionqueue) option instead apply their queue-wide limits to each partition, and cannot combine the two. ::: #### Delayed Execution You can delay an enqueued workflow's execution using [`WithDelay`](../reference/workflows-steps#withdelay). The workflow is initially placed in `DELAYED` status and does not execute. After the delay expires, it transitions to `ENQUEUED` status and may be dequeued and executed. This is useful for scheduling workflows to run at a future time. **Example syntax:** ```go remindersQueue, err := dbos.RegisterQueue(dbosContext, "reminders") if err != nil { return err } // Send a reminder in one hour handle, err := dbos.RunWorkflow(dbosContext, sendReminder, userID, dbos.WithQueue(remindersQueue), dbos.WithDelay(1 * time.Hour), ) ``` When [enqueueing from a Client](#enqueueing-from-another-application), use [`WithEnqueueDelay`](../reference/methods.md#enqueue) instead. You can also dynamically update or shorten the delay of a `DELAYED` workflow using [`SetWorkflowDelay`](../reference/methods.md#setworkflowdelay): ```go // Shorten the delay to 10 seconds from now err := dbos.SetWorkflowDelay(ctx, handle.GetWorkflowID(), dbos.WithDelayDuration(10*time.Second)) // Or set an absolute deadline err = dbos.SetWorkflowDelay(ctx, handle.GetWorkflowID(), dbos.WithDelayUntil(time.Now().Add(time.Minute))) ``` #### Listening to Specific Queues By default, every DBOS process listens to all registered queues. You can use [`ListenQueues`](../reference/queues.md#listenqueues) to configure a process to listen to only a subset of queues. This is useful when you have multiple DBOS processes and want each one to handle different types of work. ```go dbos.RegisterQueue(dbosContext, "email-queue") dbos.RegisterQueue(dbosContext, "data-queue") // This process will only dequeue workflows from the email queue dbos.ListenQueues(dbosContext, "email-queue") dbos.Launch(dbosContext) ``` Note that `ListenQueues` only controls what workflows are dequeued, not what workflows can be enqueued, so any process can still enqueue workflows on any queue. Queue names may be added to the listen set at any time, including after `Launch()`. --- ## Scheduling Workflows You can schedule DBOS [workflows](./workflow-tutorial.md) to run on a cron schedule. Schedules are stored in the database and can be created, paused, resumed, and deleted at runtime. Each time a schedule fires, its workflow is executed by exactly one worker process. To schedule a workflow, first define a workflow whose input is a [`ScheduledWorkflowInput`](../reference/methods.md#scheduledworkflowinput). This struct carries the cron tick time (`ScheduledTime`) and a user-defined `Context` value attached to the schedule, which you can decode with [`DecodeScheduleContext`](../reference/methods.md#decodeschedulecontext): ```go func myPeriodicTask(ctx dbos.Context, input dbos.ScheduledWorkflowInput) (any, error) { scheduleCtx, err := dbos.DecodeScheduleContext[string](input) if err != nil { return nil, err } logger.Info("running scheduled task", "scheduled_time", input.ScheduledTime, "context", scheduleCtx) return nil, nil } dbos.RegisterWorkflow(dbosContext, myPeriodicTask) ``` Then, create a schedule for it using [`CreateSchedule`](../reference/methods.md#createschedule) with a [`ScheduleSpec`](../reference/methods.md#schedulespec) containing a [crontab](https://en.wikipedia.org/wiki/Cron) expression: ```go err := dbos.CreateSchedule(dbosContext, dbos.ScheduleSpec{ ScheduleName: "my-task-schedule", // The schedule name is a unique identifier of the schedule Workflow: myPeriodicTask, // A registered workflow function Schedule: "0 */5 * * * *", // Every 5 minutes Context: "my context", // Passed into every iteration of the workflow }) ``` Note that `CreateSchedule` will fail if the schedule already exists. If you're defining a set of static schedules to be created on program start, you can instead use [`ApplySchedules`](../reference/methods.md#applyschedules) to create them atomically, updating them if they already exist: ```go err := dbos.ApplySchedules(dbosContext, []dbos.ScheduleSpec{ { ScheduleName: "schedule-a", Workflow: workflowA, Schedule: "0 */10 * * * *", // Every 10 minutes Context: "context-a", }, { ScheduleName: "schedule-b", Workflow: workflowB, Schedule: "0 0 0 * * *", // Every day at midnight Context: "context-b", }, }) ``` When `ApplySchedules` updates an existing schedule, it replaces the entire definition with the new entry, so any optional field left unset is cleared. For example, if a schedule was routed to a named queue and you re-apply it without setting `QueueName`, it reverts to the internal queue. The schedule's status and last-fired time are preserved. To learn more about crontab syntax, see [this guide](https://docs.gitlab.com/ee/topics/cron/) or [this crontab editor](https://crontab.guru/). DBOS Go uses [robfig/cron](https://pkg.go.dev/github.com/robfig/cron/v3) to parse cron schedules, with seconds as the first field. Valid cron schedules contain exactly 6 items, separated by spaces: ``` ┌────────────── second │ ┌──────────── minute │ │ ┌────────── hour │ │ │ ┌──────── day of month │ │ │ │ ┌────── month │ │ │ │ │ ┌──── day of week │ │ │ │ │ │ │ │ │ │ │ │ * * * * * * ``` Cron expressions are evaluated in UTC by default. Set the `CronTimezone` field of [`ScheduleSpec`](../reference/methods.md#schedulespec) to an [IANA timezone name](https://en.wikipedia.org/wiki/List_of_tz_database_time_zones) (e.g. `"America/New_York"`) to evaluate the expression in a different timezone. You can dynamically create many schedules for the same workflow. For example, if you want to perform certain actions periodically for each of your customers, you can create one schedule per customer, using customer ID as context so each workflow knows which customer to act on: ```go func customerWorkflow(ctx dbos.Context, input dbos.ScheduledWorkflowInput) (any, error) { customerID, err := dbos.DecodeScheduleContext[string](input) if err != nil { return nil, err } // ... return nil, nil } dbos.RegisterWorkflow(dbosContext, customerWorkflow) func onCustomerRegistration(ctx dbos.Context, customerID string) error { return dbos.CreateSchedule(ctx, dbos.ScheduleSpec{ ScheduleName: fmt.Sprintf("customer-%s-sync", customerID), Workflow: customerWorkflow, Schedule: "0 0 * * * *", // Every hour Context: customerID, }) } ``` The `Context` field on `ScheduleSpec` is typed as `any` and is serialized as JSON when the schedule is stored. Inside the workflow, recover the original value with [`DecodeScheduleContext`](../reference/methods.md#decodeschedulecontext). #### Managing Schedules You can pause, resume, and delete schedules at runtime: ```go // Pause a schedule so it stops firing err := dbos.PauseSchedule(dbosContext, "my-task-schedule") // Resume a paused schedule err = dbos.ResumeSchedule(dbosContext, "my-task-schedule") // Delete a schedule err = dbos.DeleteSchedule(dbosContext, "my-task-schedule") ``` You can also list and inspect schedules: ```go // List all active schedules schedules, err := dbos.ListSchedules(dbosContext, dbos.WithScheduleStatuses(dbos.ScheduleStatusActive)) // Get a specific schedule by name schedule, err := dbos.GetSchedule(dbosContext, "my-task-schedule") ``` #### Backfilling and Triggering If a schedule was paused or your application was offline, you can backfill missed executions using [`BackfillSchedule`](../reference/methods.md#backfillschedule). Already-executed times are automatically skipped: ```go start := time.Date(2025, 1, 1, 0, 0, 0, 0, time.UTC) end := time.Date(2025, 1, 2, 0, 0, 0, 0, time.UTC) ids, err := dbos.BackfillSchedule(dbosContext, "my-task-schedule", start, end) ``` Alternatively, set `AutomaticBackfill: true` on the [`ScheduleSpec`](../reference/methods.md#schedulespec) when creating a schedule so that missed executions are automatically backfilled whenever your application starts or a paused schedule is resumed. Backfills (manual or automatic) compute missed executions using the schedule's **current** cron expression. If you update a schedule's cron expression and then backfill, the backfill generates one execution per tick of the new expression over the requested window—including times the old expression would never have matched. For example, changing a daily schedule to an hourly one and then backfilling yesterday enqueues 24 executions, not 1. You can also immediately trigger a schedule using [`TriggerSchedule`](../reference/methods.md#triggerschedule): ```go handle, err := dbos.TriggerSchedule[any](dbosContext, "my-task-schedule") ``` #### Scheduling to Queues By default, scheduled workflows are enqueued on an internal queue. You can instead enqueue them on a declared [queue](./queue-tutorial.md) to manage their concurrency or rate limits. Set the `QueueName` field of [`ScheduleSpec`](../reference/methods.md#schedulespec) when creating the schedule: ```go dbos.RegisterQueue(dbosContext, "scheduled_queue", dbos.WithGlobalConcurrency(1)) err := dbos.CreateSchedule(dbosContext, dbos.ScheduleSpec{ ScheduleName: "my-task-schedule", Workflow: myPeriodicTask, Schedule: "0 */5 * * * *", QueueName: "scheduled_queue", }) ``` This ensures that scheduled workflow executions respect the queue's flow control settings. #### Managing Schedules from Another Application You can manage schedules from outside your DBOS application using a [standalone client](../reference/dbos-context.md#newclient). Because workflows are not registered with a client, set the `WorkflowName` field (a string) instead of the `Workflow` function reference: ```go client, err := dbos.NewClient(context.Background(), dbos.ClientConfig{ DatabaseURL: os.Getenv("DBOS_SYSTEM_DATABASE_URL"), // The name of the application that owns and runs the schedule AppName: "my-app", }) err = dbos.CreateSchedule(client, dbos.ScheduleSpec{ ScheduleName: "my-task-schedule", WorkflowName: "myPeriodicTask", Schedule: "0 */5 * * * *", Context: "my context", }) ``` #### How Scheduling Works Under the hood, DBOS constructs an [idempotency key](./workflow-tutorial.md#workflow-ids-and-idempotency) for each scheduled workflow execution. The key is the concatenation of `sched-`, the schedule name, and the scheduled time (RFC3339), ensuring each scheduled invocation occurs exactly once even when multiple application instances share the same schedule. When a schedule fires, its workflow is enqueued by name rather than invoked directly, so the process hosting the schedule does not need to have the workflow registered. Name resolution happens at dequeue time on a worker that has the function, letting any process connected to the system database drive schedules for workflows owned by other processes or languages. Scheduled workflows always run against their owning application's latest registered version, so a stale executor does not pick them up after a new deploy. You can list the workflows started by a schedule by passing [`WithFilterScheduleName`](../reference/methods.md#withfilterschedulename) to [`ListWorkflows`](../reference/methods.md#listworkflows). For the full API reference, see [Workflow Schedules](../reference/methods.md#workflow-schedules). --- ## Steps When using DBOS workflows, you should call any function that performs complex operations or accesses external APIs or services as a _step_. If a workflow is interrupted, upon restart it automatically resumes execution from the **last completed step**. Steps execute **at least once**: if a process crashes after a step's side effects but before the step is checkpointed, the step re-executes on recovery. For exactly-once database writes, use a [datasource transaction](./transaction-tutorial.md) instead. You can use [`RunAsStep`](../reference/workflows-steps#runasstep) to call a function as a step. For a function to be used as a step, it should return a serializable ([json-encodable](https://pkg.go.dev/encoding/json)) value and an error and have this signature: ```go type Step[R any] func(ctx context.Context) (R, error) ``` Here's a simple example: ```go func generateRandomNumber(ctx context.Context) (int, error) { return rand.Int(), nil } func workflowFunction(ctx dbos.Context, n int) (int, error) { randomNumber, err := dbos.RunAsStep( ctx, generateRandomNumber, dbos.WithStepName("generateRandomNumber"), ) if err != nil { return 0, err } return randomNumber, nil } ``` You can pass arguments into a step by wrapping it in an anonymous function, like this: ```go func generateRandomNumber(ctx context.Context, n int) (int, error) { return rand.IntN(n), nil } func workflowFunction(ctx dbos.Context, n int) (int, error) { randomNumber, err := dbos.RunAsStep( ctx, func(stepCtx context.Context) (int, error) { return generateRandomNumber(stepCtx, n) }, dbos.WithStepName("generateRandomNumber"), ) if err != nil { return 0, err } return randomNumber, nil } ``` You should make a function a step if you're using it in a DBOS workflow and it performs a [**nondeterministic**](../tutorials/workflow-tutorial.md#determinism) operation. A nondeterministic operation is one that may return different outputs given the same inputs. Common nondeterministic operations include: - Accessing an external API or service, like serving a file from [AWS S3](https://aws.amazon.com/s3/), calling an external API like [Stripe](https://stripe.com/), or accessing an external data store like [Elasticsearch](https://www.elastic.co/elasticsearch/). - Accessing files on disk. - Generating a random number. - Getting the current time. You **cannot** call, start, or enqueue workflows from within steps. You also cannot call DBOS methods like [`Send`](../reference/methods.md#send), [`Recv`](../reference/methods.md#recv), or [`RunAsTransaction`](../reference/datasources.md#runastransaction) from within steps — they return an error (see the [full list](../reference/workflows-steps.md#calling-dbos-operations-from-steps)). These operations should be performed from workflow functions. You can call one step from another step, but the called step becomes part of the calling step's execution rather than functioning as a separate step. [`SetEvent`](../reference/methods.md#setevent), [`WriteStream`](../reference/methods.md#writestream), and read operations like [`ListWorkflows`](../reference/methods.md#listworkflows) are allowed from steps. #### Configurable Retries You can optionally configure a step to automatically retry any error a set number of times with exponential backoff. This is useful for automatically handling transient failures, like making requests to unreliable APIs. Retries are configurable through step options that can be passed to [`RunAsStep`](../reference/workflows-steps.md#runasstep). Available retry configuration options include: - [`WithStepName`](../reference/workflows-steps#withstepname) - Custom name for the step (default to the [Go runtime reflection value](https://pkg.go.dev/runtime#FuncForPC)) - [`WithStepMaxRetries`](../reference/workflows-steps#withstepmaxretries) - Maximum number of times this step is automatically retried on failure (default 0) - [`WithStepMaxInterval`](../reference/workflows-steps#withstepmaxinterval) - Maximum delay between retries (default 5s) - [`WithStepBackoffFactor`](../reference/workflows-steps#withstepbackofffactor) - Exponential backoff multiplier between retries (default 2.0) - [`WithStepBaseInterval`](../reference/workflows-steps#withstepbaseinterval) - Initial delay between retries (default 100ms) - [`WithStepRetryPredicate`](../reference/workflows-steps#withstepretrypredicate) - Predicate deciding whether a step error is retried; errors it rejects are returned immediately regardless of the remaining retry budget For example, let's configure this step to retry failures (such as if the site to be fetched is temporarily down) up to 10 times: ```go func fetchStep(ctx context.Context, url string) (string, error) { resp, err := http.Get(url) if err != nil { return "", err } defer resp.Body.Close() body, err := io.ReadAll(resp.Body) if err != nil { return "", err } return string(body), nil } func fetchWorkflow(ctx dbos.Context, inputURL string) (string, error) { return dbos.RunAsStep( ctx, func(stepCtx context.Context) (string, error) { return fetchStep(stepCtx, inputURL) }, dbos.WithStepName("fetchFunction"), dbos.WithStepMaxRetries(10), dbos.WithStepMaxInterval(30*time.Second), dbos.WithStepBackoffFactor(2.0), dbos.WithStepBaseInterval(500*time.Millisecond), ) } ``` If a step exhausts all retry attempts, it returns an error to the calling workflow. #### Step Timeouts A step receives a `context.Context` like any other Go function, so you can apply a timeout or deadline to it using the standard library and react to cancellation inside the step by selecting on `ctx.Done()`. ```go func waitStep(ctx context.Context) (string, error) { select { case <-time.After(10 * time.Second): return "done", nil case <-ctx.Done(): return "", ctx.Err() } } func exampleWorkflow(ctx dbos.Context, _ string) (string, error) { result, err := dbos.RunAsStep( ctx, func(stepCtx context.Context) (string, error) { stepCtx, cancel := context.WithTimeout(stepCtx, 2*time.Second) defer cancel() return waitStep(stepCtx) }, dbos.WithStepName("waitStep"), ) if err != nil { // The workflow decides what to do: retry, fall back, or return the error. return "", err } return result, nil } ``` A few important things to keep in mind: - **Timing out a step does not cancel the workflow.** When the step returns with an error (e.g. `context.DeadlineExceeded`), the workflow continues to run and is free to handle that error—retry, fall back to another step, or return. To formally transition a workflow into the `CANCELLED` terminal status, use a workflow-level timeout instead. See [Workflow Timeouts](./workflow-tutorial.md#workflow-timeouts). - **A step can inherit the workflow's cancellable context.** If you derive the step's context from a cancellable workflow's `Context`, then when the workflow's timeout fires the workflow will become `CANCELLED`, but the currently executing step will **not** be preempted—it keeps running and can still record its outcome (success or error) to the database when it returns. The workflow will not be able to enter the next step: the next call to `RunAsStep` will fail because the workflow is already cancelled. - **If you don't want that behavior**, Handle the resulting cancellation just as you would in any normal Go program—by selecting on `ctx.Done()` in long-running loops or by passing the context through to cancellation-aware APIs. --- ## Testing & Mocking [`Context`](../reference/dbos-context) is a fully mockable interface, which you manually mock, or can generate mocks using tools like [mockery](https://github.com/vektra/mockery).
Sample .mockery.yml v3 configuration You can use this configuration file to generate mocks by running `mockery`: ```yaml all: false dir: './mocks' filename: '{{.InterfaceName}}_mock.go' force-file-write: true formatter: goimports include-auto-generated: false log-level: info structname: 'Mock{{.InterfaceName}}' pkgname: 'mocks' recursive: false require-template-schema-exists: true template: testify template-schema: '{{.Template}}.schema.json' packages: github.com/dbos-inc/dbos-transact-golang/dbos: interfaces: Context: WorkflowHandle: ```
Here is an example workflow which: - Calls a step - Spawns a child workflow - Calls a workflow management operation, [ListWorkflows](../reference/methods#listworkflows) ```go func step(ctx context.Context) (int, error) { return 1, nil } func childWorkflow(ctx dbos.Context, i int) (int, error) { return i + 1, nil } func workflow(ctx dbos.Context, i int) ([]dbos.WorkflowStatus, error) { // Test RunAsStep _, err := dbos.RunAsStep(ctx, step) if err != nil { return nil, err } // Child wf ch, err := dbos.RunWorkflow(ctx, childWorkflow, i) if err != nil { return nil, err } _, err = ch.GetResult() if err != nil { return nil, err } return dbos.ListWorkflows(ctx) } ``` Here is how you can test this workflow, assuming mocks generated with [mockery](https://github.com/vektra/mockery). The idea is that you can mock any package-level DBOS method, because it has a mirror on the `Context` interface. ```go // test file package dbos_test import ( "context" "fmt" "testing" "mocks" // Replace with the location of your generated mocks "github.com/dbos-inc/dbos-transact-golang/dbos" "github.com/stretchr/testify/mock" ) func step(ctx context.Context) (int, error) { return 1, nil } func childWorkflow(ctx dbos.Context, i int) (int, error) { return i + 1, nil } func workflow(ctx dbos.Context, i int) ([]dbos.WorkflowStatus, error) { // Test RunAsStep _, err := dbos.RunAsStep(ctx, step) if err != nil { return nil, err } // Child wf childHandle, err := dbos.RunWorkflow(ctx, childWorkflow, i) if err != nil { return nil, err } _, err = childHandle.GetResult() if err != nil { return nil, err } return dbos.ListWorkflows(ctx) } func TestMocks(t *testing.T) { // Create a mock Context mockCtx := mocks.NewMockContext(t) // Step mockCtx.On("RunAsStep", mockCtx, mock.Anything, mock.Anything).Return(1, nil) // Child workflow mockChildHandle := mocks.NewMockWorkflowHandle[any](t) // mock WorkflowHandle mockCtx.On("RunWorkflow", mockCtx, mock.Anything, 2, mock.Anything).Return(mockChildHandle, nil).Once() mockChildHandle.On("GetResult").Return(1, nil) // Workflow management mockCtx.On("ListWorkflows", mockCtx).Return([]dbos.WorkflowStatus{}, nil) workflow(mockCtx, 2) mockCtx.AssertExpectations(t) mockChildHandle.AssertExpectations(t) } ``` --- ## Transactions & Datasources A _datasource_ is a handle to a database **you** own, over which DBOS can run durable transactions. When you run a transaction through a datasource inside a workflow, your application writes and the DBOS durability record commit atomically in your database, so the transaction executes **exactly once** even if your program crashes and the workflow is recovered. This is a stronger guarantee than a [step](./step-tutorial.md) provides: a step that writes to a database may re-execute (and re-commit) if the process crashes after the write but before the step is checkpointed. ### Creating a Datasource Create a datasource with [`NewDataSource`](../reference/datasources.md#newdatasource), passing a connection pool to your database: a `*pgxpool.Pool` for Postgres or CockroachDB, or a `*sql.DB` for SQLite. ```go import "github.com/jackc/pgx/v5/pgxpool" pool, err := pgxpool.New(context.Background(), os.Getenv("APP_DATABASE_URL")) if err != nil { log.Fatal(err) } ds, err := dbos.NewDataSource(dbosContext, pool, dbos.WithDataSourceName("app")) if err != nil { log.Fatal(err) } ``` `NewDataSource` may be called at any time, before or after `Launch()`. It provisions a `transaction_completion` durability table in your database (in the `dbos` schema by default, configurable with [`WithDataSourceSchema`](../reference/datasources.md#withdatasourceschema)) unless the table already exists. If you prefer to manage DDL yourself, pre-create the table in your own migrations and connect with a role that only needs `SELECT, INSERT` on it. If you pass the same engine handle you gave DBOS as its system database (via [`Config.SystemDBPool` or `Config.SQLiteSystemDB`](../reference/configuration.md)), no `transaction_completion` table is created or managed at all: your application writes and the DBOS checkpoint commit together in a single transaction. See [Sharing the System Database Engine](../reference/datasources.md#sharing-the-system-database-engine). ### Running Transactions Inside a workflow, run a transaction with [`RunAsTransaction`](../reference/datasources.md#runastransaction). Your function receives a [`Tx`](../reference/datasources.md#the-tx-interface) on which to run queries; DBOS commits the transaction if the function returns successfully and rolls it back if it returns an error. ```go func checkoutWorkflow(ctx dbos.Context, item string) (int64, error) { // This transaction runs exactly once, even across crashes and recovery. orderID, err := dbos.RunAsTransaction(ctx, ds, func(txCtx context.Context, tx dbos.Tx) (int64, error) { var id int64 err := tx.QueryRow(txCtx, "INSERT INTO orders(item) VALUES ($1) RETURNING id", item).Scan(&id) return id, err }) if err != nil { return 0, err } // Transactions and steps share the workflow's step counter and can be interleaved. _, err = dbos.RunAsStep(ctx, func(stepCtx context.Context) (string, error) { return sendConfirmation(stepCtx, orderID) }) return orderID, err } ``` `RunAsTransaction` must be called from within a workflow. It accepts the same options as [`RunAsStep`](../reference/workflows-steps.md#runasstep) (`WithStepName`, `WithStepMaxRetries`, and so on). Serialization and deadlock conflicts are retried automatically with a fresh transaction; errors returned by your function follow your step retry policy. :::tip If your application tables live in the same database that hosts the DBOS system schema (for example, the [shared-engine setup](../reference/datasources.md#sharing-the-system-database-engine)), you can atomically write application data and enqueue a workflow in one transaction by calling the [`dbos.enqueue_workflow`](../../explanations/system-tables.md#dbosenqueue_workflow) PL/pgSQL function from within it. From outside a workflow (for example, from a [standalone client](../reference/dbos-context.md#newclient)), pass your own open transaction to `Enqueue` or `Send` with [`WithEnqueueTransaction`](../reference/methods.md#withenqueuetransaction) or [`WithSendTransaction`](../reference/methods.md#withsendtransaction) instead. ::: ### Guarantees - A `RunAsTransaction` called at the top level of a workflow is **exactly-once**: the application writes and the durability record commit atomically, and recovery replays the recorded output instead of re-running the function. - Nesting a `RunAsTransaction` inside a `RunAsStep` or inside another `RunAsTransaction` is rejected with an error: a nested transaction could not be checkpointed, so it would silently lose its exactly-once guarantee. Run transactions from workflow code. See the [datasources reference](../reference/datasources.md) for details on durability, recovery, and permissions. --- ## Upgrading Workflow Code A challenge encountered when operating long-running durable workflows in production is **how to deploy breaking changes without disrupting in-progress workflows.** A breaking change to a workflow is one that changes which steps are run, or the order in which the steps are run. If a breaking change was made to a workflow and that workflow is replayed by the recovery system, the checkpoints created by the previous version of the code may not match the steps called by the workflow in the new version of the code, causing recovery to fail. DBOS supports two strategies for safely upgrading workflow code: **patching** and **versioning**. ### Patching In patching, the result of a call to [`dbos.Patch()`](../reference/workflows-steps.md#patch) is used to conditionally execute the new code. `dbos.Patch()` returns `true` for new calls (those executing after the breaking change) and `false` for old calls (those that executed before the breaking change). Therefore, if `dbos.Patch()` returns `true`, the workflow should follow the new code path, otherwise it must follow the prior codepath. To use patching, you must enable it in the configuration: ```go dbosCtx, err := dbos.NewContext(context.Background(), dbos.Config{ DatabaseURL: os.Getenv("DBOS_DATABASE_URL"), AppName: "my-app", EnablePatching: true, }) ``` For example, let's say our original workflow is: ```go func workflow(ctx dbos.Context, input string) (string, error) { _, err := dbos.RunAsStep(ctx, func(stepCtx context.Context) (string, error) { return foo(stepCtx) }) if err != nil { return "", err } _, err = dbos.RunAsStep(ctx, func(stepCtx context.Context) (string, error) { return bar(stepCtx) }) if err != nil { return "", err } return "success", nil } ``` We want to replace the call to `foo()` with a call to `baz()`. This is a breaking change because it changes what steps run. We can make this breaking change safely using a patch: ```go func workflow(ctx dbos.Context, input string) (string, error) { patched, err := dbos.Patch(ctx, "use-baz") if err != nil { return "", err } if patched { _, err = dbos.RunAsStep(ctx, func(stepCtx context.Context) (string, error) { return baz(stepCtx) }) } else { _, err = dbos.RunAsStep(ctx, func(stepCtx context.Context) (string, error) { return foo(stepCtx) }) } if err != nil { return "", err } _, err = dbos.RunAsStep(ctx, func(stepCtx context.Context) (string, error) { return bar(stepCtx) }) if err != nil { return "", err } return "success", nil } ``` Now, new workflows will run `baz()`, while old workflows will reexecute `foo()`. Examples of workflows taking the old code: - Recovered workflows that executed up to or beyond the patch point - Workflows forked after the patch point Examples of workflows taking the new code: - Entirely new workflows - Recovered workflows that executed before the patch point - Workflows forked before the patch point #### Deprecating and Removing Patches Patches add complexity and runtime overhead; fortunately they don't need to stay in your code forever. Once all workflows that started before you deployed the patch are complete, you can safely remove patches from your code. :::tip You can use the [list workflows APIs](./workflow-management.md#listing-workflows) to see what workflows are still active. ::: First, you must deprecate the patch with [`dbos.DeprecatePatch()`](../reference/workflows-steps.md#deprecatepatch). `DeprecatePatch` must be used for a transition period prior to fully removing the patch, as it allows coexistence with any ongoing workflows that used `Patch()`. For example, here's how to deprecate the patch above: ```go func workflow(ctx dbos.Context, input string) (string, error) { err := dbos.DeprecatePatch(ctx, "use-baz") if err != nil { return "", err } _, err = dbos.RunAsStep(ctx, func(stepCtx context.Context) (string, error) { return baz(stepCtx) }) if err != nil { return "", err } _, err = dbos.RunAsStep(ctx, func(stepCtx context.Context) (string, error) { return bar(stepCtx) }) if err != nil { return "", err } return "success", nil } ``` Then, when all workflows that started before you deprecated the patch are complete, you can remove the patch entirely: ```go func workflow(ctx dbos.Context, input string) (string, error) { _, err := dbos.RunAsStep(ctx, func(stepCtx context.Context) (string, error) { return baz(stepCtx) }) if err != nil { return "", err } _, err = dbos.RunAsStep(ctx, func(stepCtx context.Context) (string, error) { return bar(stepCtx) }) if err != nil { return "", err } return "success", nil } ``` If any mistakes happen during the process (a breaking change is not patched, or a patch is deprecated or removed prematurely), the workflow will return an error pointing to the step where the problem occurred. #### How Patching Works Under the hood, when you call `dbos.Patch()` from a workflow, it attempts to insert a "patch marker" at its current point in your workflow history (this is a new row in the `operation_outputs` table in the DBOS [system database](../../explanations/system-tables.md)). If it successfully inserts the patch marker or if the patch marker is already present, then the workflow should take the patch codepath. If there is already a record present in this point in your workflow history and it is not a patch marker, then the workflow must be old (it already continued past this point with old code), and `dbos.Patch()` returns `false`. When you deprecate a patch with `dbos.DeprecatePatch()`, new workflows no longer insert patch markers into their workflow history. However, if a workflow contains the patch marker in its history, it continues past that patch marker, safely ignoring it. Once all workflows with patch markers are complete, the patch may be safely removed. ### Versioning When using versioning, DBOS **versions** applications and workflows, and only continues workflow execution with the same application version that started the workflow. All workflows are tagged with the application version on which they started. By default, application version is automatically computed from a hash of workflow source code. However, you can set your own version through configuration. ```go dbosCtx, err := dbos.NewContext(context.Background(), dbos.Config{ DatabaseURL: os.Getenv("DBOS_DATABASE_URL"), AppName: "my-app", ApplicationVersion: "1.0.0", }) ``` When DBOS tries to recover workflows, it only recovers workflows whose version matches the current application version. This prevents recovery of workflows that depend on different code. Enqueued workflows with no application version (for example, enqueued from a [DBOS Client](../reference/dbos-context.md#newclient) without specifying a version) are only dequeued by processes running their owning application's **latest** registered version. This ensures version-less work is picked up by the newest code during a rolling deploy. When using versioning, we recommend **blue-green** code upgrades: - When deploying a new version of your code, launch new processes running your new code version, but retain some processes running your old code version. - Direct new traffic to your new processes while your old processes "drain" and complete all workflows of the old code version. - Then, once all workflows of the old version are complete (you can use [`ListWorkflows`](../reference/methods.md#listworkflows) to check), you can retire the old code version. --- ## Communicating with Workflows DBOS provides a few different ways to communicate with your workflows. You can: - [Send messages to workflows](#workflow-messaging-and-notifications) - [Publish events from workflows for clients to read](#workflow-events) - [Stream values from workflows to clients](#workflow-streaming) ### Workflow Messaging and Notifications You can send messages to a specific workflow. This is useful for signaling a workflow or sending notifications to it while it's running. ##### Send ```go func Send[P any](ctx Client, destinationID string, message P, topic string, opts ...SendOption) error ``` You can call `Send()` to send a message to a workflow. Messages can optionally be associated with a topic and are queued on the receiver per topic. Pass [`WithIdempotencyKey`](../reference/methods.md#withidempotencykey) to make a retried `Send` deliver at most once: retrying with the same key (for example, after a crash or network failure) inserts the message only once. To send many messages at once, possibly to different workflows, use [`SendBulk`](../reference/methods.md#sendbulk). The batch is sent in a single transaction: either every message is delivered or none is. ```go err := dbos.SendBulk(dbosContext, []dbos.SendMessage{ {DestinationID: orderWorkflowID, Message: "confirmed", Topic: "orders"}, {DestinationID: inventoryWorkflowID, Message: order, Topic: "reserve", IdempotencyKey: "reserve-42"}, }) ``` ##### Recv ```go func Recv[R any](ctx Context, topic string, timeout time.Duration) (R, error) ``` Workflows can call `Recv()` to receive messages sent to them, optionally for a particular topic. Each call to `Recv()` waits for and consumes the next message to arrive in the queue for the specified topic, returning an error if the wait times out. If the topic is not specified, this method only receives messages sent without a topic. ##### Messages Example Messages are especially useful for sending notifications to a workflow. For example, in an e-commerce application, the checkout workflow, after redirecting customers to a secure payments service, must wait for a notification from that service that the payment has finished processing. To wait for this notification, the payments workflow uses `Recv()`, executing failure-handling code if the notification doesn't arrive in time: ```go const PaymentStatusTopic = "payment_status" func checkoutWorkflow(ctx dbos.Context, orderData OrderData) (string, error) { // Process initial checkout steps... // Wait for payment notification with a 5-minute timeout notification, err := dbos.Recv[PaymentNotification](ctx, PaymentStatusTopic, 5*time.Minute) if err != nil { ... // Handle timeout or other errors } // Handle the notification if notification.Status == "completed" { ... // Handle the notification. } else { ... // Handle a failure } } ``` A webhook waits for the payment processor to send the notification, then uses `Send()` to forward it to the workflow: ```go func paymentWebhookHandler(w http.ResponseWriter, r *http.Request) { // Parse the notification from the payment processor notification := ... // Retrieve the workflow ID from notification metadata workflowID := ... // Send the notification to the waiting workflow err := dbos.Send(dbosContext, workflowID, notification, PaymentStatusTopic) if err != nil { http.Error(w, "Failed to send notification", http.StatusInternalServerError) return } } ``` ##### Reliability Guarantees All messages are persisted to the database, so if `Send` completes successfully, the destination workflow is guaranteed to be able to `Recv` it. If you're sending a message from a workflow, DBOS guarantees exactly-once delivery. ### Workflow Events Workflows can publish _events_, which are key-value pairs associated with the workflow. They are useful for publishing information about the status of a workflow or to send a result to clients while the workflow is running. ##### SetEvent ```go func SetEvent[P any](ctx Context, key string, message P, opts ...SetEventOption) error ``` Any workflow can call [`SetEvent`](../reference/methods.md#setevent) to publish a key-value pair, or update its value if it has already been published. If called from a step, the write is at-least-once (the step may be retried on failure). ##### GetEvent ```go func GetEvent[R any](ctx Client, targetWorkflowID, key string, timeout time.Duration) (R, error) ``` You can call [`GetEvent`](../reference/methods.md#getevent) to retrieve the value published by a particular workflow ID for a particular key. If the event does not yet exist, this call waits for it to be published, returning an error if the wait times out. ##### Events Example Events are especially useful for writing interactive workflows that communicate information to their caller. For example, in an e-commerce application, the checkout workflow, after validating an order, directs the customer to a secure payments service to handle credit card processing. To communicate the payments URL to the customer, it uses events. The checkout workflow emits the payments URL using `SetEvent()`: ```go const PaymentURLKey = "payment_url" func checkoutWorkflow(ctx dbos.Context, orderData OrderData) (string, error) { // Process order validation... paymentsURL := ... err := dbos.SetEvent(ctx, PaymentURLKey, paymentsURL) if err != nil { return "", fmt.Errorf("failed to set payment URL event: %w", err) } // Continue with checkout process... } ``` The HTTP handler that originally started the workflow uses `GetEvent()` to await this URL, then redirects the customer to it: ```go func webCheckoutHandler(dbosContext dbos.Context, w http.ResponseWriter, r *http.Request) { orderData := parseOrderData(r) // Parse order from request handle, err := dbos.RunWorkflow(dbosContext, checkoutWorkflow, orderData) if err != nil { http.Error(w, "Failed to start checkout", http.StatusInternalServerError) return } // Wait up to 30 seconds for the payment URL event url, err := dbos.GetEvent[string](dbosContext, handle.GetWorkflowID(), PaymentURLKey, 30*time.Second) if err != nil { // Handle a timeout } // Redirect the customer } ``` ##### Reliability Guarantees All events are persisted to the database, so the latest version of an event is always retrievable. Additionally, if `GetEvent` is called in a workflow, the retrieved value is persisted in the database so workflow recovery can use that value, even if the event is later updated. ### Workflow Streaming Workflows can stream data in real time to clients. This is useful for streaming results from a long-running workflow or LLM call, or for monitoring and progress reporting. ##### Writing to Streams ```go func WriteStream[P any](ctx Context, key string, value P, opts ...WriteStreamOption) error ``` You can write values to a stream from a workflow or its steps using [`WriteStream`](../reference/methods.md#writestream). A workflow may have any number of streams, each identified by a unique key. When you are done writing to a stream, you should close it with [`CloseStream`](../reference/methods.md#closestream). Otherwise, streams are automatically closed when the workflow terminates. ```go func CloseStream(ctx Context, key string) error ``` DBOS streams are immutable and append-only. Writes to a stream from a workflow happen exactly-once. Writes to a stream from a step happen at-least-once; if a step fails and is retried, it may write to the stream multiple times. Readers will see all values written to the stream from all tries of the step in the order in which they were written. **Example syntax:** ```go func producerWorkflow(ctx dbos.Context, _ string) (string, error) { err := dbos.WriteStream(ctx, "progress", "step 1 complete") if err != nil { return "", err } err = dbos.WriteStream(ctx, "progress", "step 2 complete") if err != nil { return "", err } err = dbos.CloseStream(ctx, "progress") if err != nil { return "", err } return "done", nil } ``` ##### Reading from Streams ```go func ReadStream[R any](ctx Client, workflowID string, key string, opts ...ReadStreamOption) ([]R, bool, error) ``` You can read values from a stream from anywhere using [`ReadStream`](../reference/methods.md#readstream). This function reads all values from a stream identified by a workflow ID and key. It blocks until the stream is closed or the workflow becomes inactive (status is not `PENDING` or `ENQUEUED`). It returns the values, whether the stream is closed, and any error. To read without blocking, pass [`WithReadStreamSnapshot`](../reference/methods.md#withreadstreamsnapshot), which returns as soon as all currently-available values have been drained. Pair it with [`WithReadStreamFromOffset`](../reference/methods.md#withreadstreamfromoffset) to set the offset to start reading from, so you can poll a stream incrementally: ```go // Read whatever is available right now, starting from offset 0, without blocking. values, closed, err := dbos.ReadStream[string](ctx, workflowID, "progress", dbos.WithReadStreamSnapshot()) ``` You can also read from a stream asynchronously using [`ReadStreamAsync`](../reference/methods.md#readstreamasync), which returns a channel: ```go func ReadStreamAsync[R any](ctx Client, workflowID string, key string) (<-chan StreamValue[R], error) ``` ```go type StreamValue[R any] struct { Value R // The stream value (zero value if error/closed) Err error // Error if one occurred (nil otherwise) Closed bool // Whether the stream is closed } ``` You can also read from a stream from outside a DBOS application with a [DBOS Client](../reference/methods.md#streams). **Example syntax:** ```go // Blocking read: wait for all values, then process values, closed, err := dbos.ReadStream[string](ctx, workflowID, "progress") if err != nil { return err } for _, value := range values { fmt.Printf("Received: %s\n", value) } // Async read: process values as they arrive ch, err := dbos.ReadStreamAsync[string](ctx, workflowID, "progress") if err != nil { return err } for streamValue := range ch { if streamValue.Err != nil { return streamValue.Err } if streamValue.Closed { break } fmt.Printf("Received: %s\n", streamValue.Value) } ``` ##### Stream Reliability Guarantees All stream values are persisted to the database, so streams are durable and survive restarts. If `WriteStream` is called from a workflow, the write is exactly-once. If called from a step, the write is at-least-once (the step may be retried on failure). --- ## Workflow Management(Tutorials) You can view and manage your durable workflow executions via the [DBOS Console](../../conductor/workflow-management.md) or programmatically. ### Listing Workflows You can list your application's workflows programmatically via [`ListWorkflows`](../reference/methods#listworkflows). You can also view a searchable and expandable list of your application's workflows from its page on the [DBOS Console](../../conductor/workflow-management.md). ### Visualizing Workflow Execution You can also visualize a workflow's execution as a trace timeline (showing the workflow, its steps, and its child workflows and their steps) from its page on the [DBOS Console](../../conductor/workflow-management.md). For example, here is the trace of a workflow that processes multiple tasks concurrently by enqueueing child workflows: ### Tagging Workflows with Attributes You can attach custom key-value attributes to a workflow with [`WithWorkflowAttributes`](../reference/workflows-steps.md#withworkflowattributes) when starting or enqueueing it: ```go handle, err := dbos.RunWorkflow(ctx, processOrder, order, dbos.WithWorkflowAttributes(map[string]any{"customer": "acme", "region": "us-east"}), ) ``` Attributes are recorded in the workflow's status and can be used to find workflows with [`ListWorkflows`](../reference/methods.md#listworkflows) and the [`WithFilterAttributes`](../reference/methods.md#withfilterattributes) filter (Postgres only): ```go workflows, err := dbos.ListWorkflows(ctx, dbos.WithFilterAttributes(map[string]any{"customer": "acme"}), ) ``` You can replace a workflow's attributes at any time with [`SetWorkflowAttributes`](../reference/methods.md#setworkflowattributes). ### Cancelling Workflows You can cancel the execution of a workflow from the web UI or programmatically via [`CancelWorkflow`](../reference/methods#cancelworkflow). To cancel many workflows at once, use [`CancelWorkflows`](../reference/methods#cancelworkflows), which cancels them in a single database round-trip. Pass [`WithCancelChildren`](../reference/methods#withcancelchildren) to also cancel all the workflow's children, recursively. If the workflow is enqueued, cancelling removes it from the queue. If the workflow is currently executing, cancelling sets its status to `CANCELLED`; the execution is not interrupted mid-step and its output is checkpointed, but the execution stops at the start of its **next durable operation** (step, child workflow, sleep, `Send`/`Recv`, and so on). That operation returns an error matching [`dbos.ErrWorkflowCancelled`](../reference/workflows-steps.md#error-codes). Do not ignore that error to continue execution. Ignoring a cancellation error cannot turn the workflow back into a success. #### Cancellation via context or timeout A workflow is also cancelled when its [durable timeout](./workflow-tutorial.md#workflow-timeouts) expires, when the context it was started from is cancelled, or on shutdown. You can use this to cancel a workflow directly: start it under a cancellable context obtained with [`WithCancel`](../reference/dbos-context.md#withcancel), then call the cancel function. ```go dbosCtx, cancel := dbos.WithCancel(ctx) handle, err := dbos.RunWorkflow(dbosCtx, myWorkflow, input) // ... later: cancel() ``` This form of cancellation enables **cooperative cancellation**: an executing step receives the cancellation through its `context.Context` and can select on `ctx.Done()` to return early instead of running to completion. ```go func longRunningStep(ctx context.Context) (string, error) { for { select { case <-ctx.Done(): return "", ctx.Err() // Return early; this step is not checkpointed and re-executes on resume default: if done, result := doSomeWork(); done { return result, nil } } } } ``` A step interrupted this way returns an error matching `dbos.ErrWorkflowCancelled` that also wraps the standard-library cause — `errors.Is(err, context.Canceled)` or `errors.Is(err, context.DeadlineExceeded)` matches too. The interrupted step is deliberately **not** checkpointed, so if the workflow is later resumed, that step re-executes. (An API `CancelWorkflow`, by contrast, does not cancel the running execution's `Context`, and its cancellation errors carry no standard-library cause.) :::note Cancelling the context durably cancels the workflow: DBOS immediately marks it `CANCELLED` in the database, exactly as [`CancelWorkflow`](#cancelling-workflows) would. The cancellation is then enforced at the step boundary: a step that does not watch `ctx.Done()` keeps running, but when it returns, its result — success or error — is discarded rather than checkpointed, and the step call reports the cancellation instead. Handle cooperative cancellation in long-running steps and return early: any work done after the context is cancelled only computes a result DBOS will throw away. ::: A durable [`Sleep`](../reference/methods.md#sleep) wakes immediately when the workflow's context is cancelled. An API `CancelWorkflow` does not wake an in-progress sleep: the workflow sleeps out the remaining time and stops at its next durable operation. #### Awaiting a cancelled workflow Waiting on a cancelled workflow's handle returns an error matching [`dbos.ErrAwaitedWorkflowCancelled`](../reference/workflows-steps.md#error-codes) — a distinct code, so an awaiting workflow can tell "the workflow I awaited was cancelled" apart from "I was cancelled" and may choose to handle it and continue. When the awaiter is itself a workflow, this outcome is checkpointed like any other child error, so replay is deterministic: resuming the cancelled workflow later does not change what the awaiter observed. ### Resuming Workflows You can resume a workflow from its last completed step from the web UI or programmatically via [`ResumeWorkflow`](../reference/methods#resumeworkflow). You can use this to resume workflows that are cancelled or that have exceeded their maximum recovery attempts. You can also use this to start an enqueued workflow immediately, bypassing its queue. Resuming restarts the workflow function from the beginning, but every checkpointed step replays its recorded result instead of re-executing, so only unfinished work runs again (including a step that was interrupted by cancellation, which is never checkpointed). Resuming also resets the workflow's recovery-attempt counter and clears any durable timeout the workflow carried before it was resumed. ### Forking Workflows You can start a new execution of a workflow by **forking** it from a specific step. When you fork a workflow, DBOS generates a new workflow with a new workflow ID, copies to that workflow the original workflow's inputs and all its steps up to the selected step, then begins executing the new workflow from the selected step. Forking a workflow is useful for recovering from outages in downstream services (by forking from the step that failed after the outage is resolved) or for "patching" workflows that failed due to a bug in a previous application version (by forking from the bugged step to an application version on which the bug is fixed). You can fork a workflow programmatically using [`ForkWorkflow`](../reference/methods#forkworkflow). The forked workflow can be given its own timeout, and a parent whose children were also forked can be pointed at the new children with `ReplacementChildren`. You can also fork a workflow from a step from the web UI by clicking on that step in the workflow's trace timeline: --- ## Workflows Workflows provide **durable execution** so you can write programs that are **resilient to any failure**. Workflows are comprised of [steps](./step-tutorial.md), which wrap ordinary Go functions. If a workflow is interrupted for any reason (e.g., an executor restarts or crashes), when your program restarts the workflow automatically resumes execution from the last completed step. To write a workflow, register a Go function with [`RegisterWorkflow`](../reference/workflows-steps.md#registerworkflow). Workflow registration must happen before launching the DBOS context with `dbos.Launch()` The function's signature must match: ```go type Workflow[P any, R any] func(ctx Context, input P) (R, error) ``` In other words, a workflow must take in a DBOS context and one other input of any serializable ([json-encodable](https://pkg.go.dev/encoding/json)) type and must return one output of any serializable type and error. For example: ```go func stepOne(ctx context.Context) (string, error) { fmt.Println("Step one completed") return "success", nil } func stepTwo(ctx context.Context) (string, error) { fmt.Println("Step two completed") return "success", nil } func workflow(ctx dbos.Context, _ string) (string, error) { _, err := dbos.RunAsStep(ctx, stepOne) if err != nil { return "failure", err } _, err = dbos.RunAsStep(ctx, stepTwo) if err != nil { return "failure", err } return "success", err } func main() { ... // Create the DBOS context dbos.RegisterWorkflow(dbosContext, workflow) ... // Launch DBOS after registering all workflows } ``` Call workflows with [`RunWorkflow`](../reference/workflows-steps.md#runworkflow). This starts the workflow in the background and returns a [workflow handle](../reference/workflows-steps.md#workflowhandle) from which you can access information about the workflow or wait for it to complete and return its result. Here's an example: ```go func runWorkflowExample(dbosContext dbos.Context, input string) error { handle, err := dbos.RunWorkflow(dbosContext, workflow, input) if err != nil { return err } result, err := handle.GetResult() if err != nil { return err } fmt.Println("Workflow result:", result) return nil } ``` ### Workflow IDs and Idempotency Every time you execute a workflow, that execution is assigned a unique ID, by default a [UUID](https://en.wikipedia.org/wiki/Universally_unique_identifier). You can access this ID through [`GetWorkflowID`](../reference/methods.md#getworkflowid), or from the handle's [`GetWorkflowID`](../reference/workflows-steps.md#workflowhandlegetworkflowid) method. Workflow IDs are useful for communicating with workflows and developing interactive workflows. You can set the workflow ID of a workflow using [`WithWorkflowID`](../reference/workflows-steps.md#withworkflowid) when calling `RunWorkflow`. Workflow IDs must be **globally unique** for your application. An assigned workflow ID acts as an idempotency key: if a workflow is called multiple times with the same ID, it executes only once. This is useful if your operations have side effects like making a payment or sending an email. For example: ```go func exampleWorkflow(ctx dbos.Context, input string) (string, error) { workflowID, err := dbos.GetWorkflowID(ctx) if err != nil { return "", err } fmt.Printf("Running workflow with ID: %s\n", workflowID) // ... return "success", nil } func example(dbosContext dbos.Context, input string) error { myID := "unique-workflow-id-123" handle, err := dbos.RunWorkflow(dbosContext, exampleWorkflow, input, dbos.WithWorkflowID(myID)) if err != nil { log.Fatal(err) } result, err := handle.GetResult() if err != nil { log.Fatal(err) } fmt.Println("Result:", result) return nil } ``` ### Determinism Workflows are in most respects normal Go functions. They can have loops, branches, conditionals, and so on. However, a workflow function must be **deterministic**: if called multiple times with the same inputs, it should invoke the same steps with the same inputs in the same order (given the same return values from those steps). If you need to perform a non-deterministic operation like accessing the database, calling a third-party API, generating a random number, or getting the local time, you shouldn't do it directly in a workflow function. Instead, you should perform all non-deterministic operations in [steps](./step-tutorial.md). :::warning Go's goroutine scheduler and `select` operation are non-deterministic. You should use them only inside steps, or use the durable [`Go`](#concurrent-steps) and [`Select`](#selecting-the-first-result) functions instead. Go's map iteration order is also random: each iteration over the same map may visit keys in a different order. Don't call steps or start workflows while ranging over a map—on recovery, the iteration order changes and operations replay in a different order, breaking determinism. Sort the keys first and iterate over the sorted slice instead: ```go keys := slices.Sorted(maps.Keys(m)) for _, k := range keys { handle, err := dbos.RunWorkflow(ctx, processItem, m[k]) // ... } ``` ::: For example, **don't do this**: ```go func exampleWorkflow(ctx dbos.Context, input string) (string, error) { randomChoice := rand.Intn(2) if randomChoice == 0 { return dbos.RunAsStep(ctx, stepOne) } else { return dbos.RunAsStep(ctx, stepTwo) } } ``` Instead, do this: ```go func generateChoice(ctx context.Context) (int, error) { return rand.Intn(2), nil } func exampleWorkflow(ctx dbos.Context, input string) (string, error) { randomChoice, err := dbos.RunAsStep(ctx, generateChoice) if err != nil { return "", err } if randomChoice == 0 { return dbos.RunAsStep(ctx, stepOne) } else { return dbos.RunAsStep(ctx, stepTwo) } } ``` ### Workflow Timeouts You can set a timeout for a workflow using its input [`Context`](../reference/dbos-context.md). Use [`WithTimeout`](../reference/dbos-context#withtimeout) to obtain a cancellable `Context`, as you would with a normal [`context.Context`](https://pkg.go.dev/context#Context). When the timeout expires, the workflow and all its children are cancelled. Cancelling a workflow sets its status to CANCELLED and preempts its execution at the beginning of its next step; the workflow's `Context` is also cancelled, so an executing step can react to `ctx.Done()`. DBOS calls then return an error matching `dbos.ErrWorkflowCancelled` whose cause matches `context.DeadlineExceeded` — see [cancellation behavior](./workflow-management.md#cancelling-workflows) for details. You can detach a child workflow by passing it an uncancellable context, which you can obtain with [`WithoutCancel`](../reference/dbos-context#withoutcancel). Timeouts are **start-to-completion**: if a workflow is [enqueued](./queue-tutorial.md), the timeout does not begin until the workflow is dequeued and starts execution. Also, timeouts are durable: they are stored in the database and persist across restarts, so workflows can have very long timeouts. ```go func exampleWorkflow(ctx dbos.Context, input string) (string, error) {} timeoutCtx, cancelFunc := dbos.WithTimeout(dbosCtx, 12*time.Hour) handle, err := dbos.RunWorkflow(timeoutCtx, exampleWorkflow, "wait-for-cancel") ``` You can also manually cancel the workflow by calling its `cancel` function (or calling [CancelWorkflow](./workflow-management.md#cancelling-workflows)). ### Durable Sleep You can use [`Sleep`](../reference/methods#sleep) to put your workflow to sleep for any period of time. This sleep is **durable**—DBOS saves the wakeup time in the database so that even if the workflow is interrupted and restarted multiple times while sleeping, it still wakes up on schedule. Sleeping is useful for scheduling a workflow to run in the future (even days, weeks, or months from now). For example: ```go func runTask(ctx context.Context, task string) (string, error) { // Execute the task... return "task completed", nil } func exampleWorkflow(ctx dbos.Context, input struct { TimeToSleep time.Duration Task string }) (string, error) { // Sleep for the specified duration _, err := dbos.Sleep(ctx, input.TimeToSleep) if err != nil { return "", err } // Execute the task after sleeping result, err := dbos.RunAsStep( ctx, func(stepCtx context.Context) (string, error) { return runTask(stepCtx, input.Task) }, ) if err != nil { return "", err } return result, nil } ``` ### Scheduled Workflows You can schedule workflows to run on a cron expression. Schedules are stored in the database and can be created, paused, resumed, and deleted at runtime. See the [scheduled workflows tutorial](./scheduled-workflows.md) for details. ### Debouncing **Debouncing** delays a workflow's execution until some time has passed since it was last called. This is useful when rapid successive triggers should be coalesced into a single workflow execution. For example, if a user is editing a text field, you may want to start a processing workflow only after the user stops typing. To debounce a workflow, define the workflow and queue, then create a [`Debouncer`](../reference/queues.md#debouncer) for it: ```go func processInput(ctx dbos.Context, input string) (string, error) { fmt.Printf("Processing input: %s\n", input) return "processed", nil } func main() { dbosContext, _ := dbos.NewContext(context.Background(), dbos.Config{ AppName: "debounce-example", ApplicationVersion: "0.1.0", DatabaseURL: os.Getenv("DBOS_SYSTEM_DATABASE_URL"), }) dbos.RegisterWorkflow(dbosContext, processInput) // Create a debouncer with a maximum timeout of 30 seconds debouncer, err := dbos.NewDebouncer(dbosContext, processInput, dbos.WithDebouncerTimeout(30*time.Second)) if err != nil { log.Fatal(err) } dbos.Launch(dbosContext) defer dbos.Shutdown(dbosContext, 5*time.Second) // Each call to Debounce pushes back the workflow start time by the delay. // The workflow runs with the most recent input once the delay expires. handle, err := debouncer.Debounce(dbosContext, "user-123", 5*time.Second, "first input") if err != nil { log.Fatal(err) } // If this call arrives within 5 seconds, the delay resets and the input updates handle, err = debouncer.Debounce(dbosContext, "user-123", 5*time.Second, "updated input") if err != nil { log.Fatal(err) } result, err := handle.GetResult() fmt.Println("Result:", result) // Processed with "updated input" } ``` Key behaviors: - Each call to `Debounce` with the same key pushes back the workflow start time by the specified delay. - When the delay expires without another call, the workflow executes with the most recent input. - The optional [`WithDebouncerTimeout`](../reference/queues.md#withdebouncertimeout) caps the maximum wait time from the first call. If the timeout is zero (the default), the delay can be pushed back indefinitely. - Different keys debounce independently, so you can debounce per-user, per-tenant, or per-resource. - You can create multiple debouncers per workflow, with different timeouts. - Debouncers can be created at any time, before or after `Launch()`. - To debounce a workflow method of a configured instance (registered with [`WithInstance`](../reference/workflows-steps.md#withinstance)), pass the instance with [`WithDebouncerInstance`](../reference/queues.md#withdebouncerinstance): `dbos.NewDebouncer(ctx, slack.Send, dbos.WithDebouncerInstance(slack))`. From a [`DebouncerClient`](../reference/queues.md#newdebouncerclient), pass the instance's config name with [`WithDebouncerConfigName`](../reference/queues.md#withdebouncerconfigname) instead. #### Debouncing from an External Application You can also debounce workflows from outside your DBOS application using a [`DebouncerClient`](../reference/queues.md#newdebouncerclient): ```go config := dbos.ClientConfig{ DatabaseURL: os.Getenv("DBOS_SYSTEM_DATABASE_URL"), } client, err := dbos.NewClient(context.Background(), config) if err != nil { log.Fatal(err) } defer dbos.Shutdown(client, 5*time.Second) dc := dbos.NewDebouncerClient[string, string]("processInput", client, dbos.WithDebouncerTimeout(30*time.Second)) handle, err := dc.Debounce("user-123", 5*time.Second, "some input") if err != nil { log.Fatal(err) } ``` ### Concurrent Steps Golang offers two building blocks to execute work concurrently: `go` and `select`. `go` starts a new goroutine and `select` allows to poll from a list of channels. Unfortunately these primitive are non-deterministic: the Golang scheduler does not guarantee the order in which goroutines are scheduled, nor does it guarantee that the same channel will, out of a set of ready channels, will be selected, when the same code runs multiple time. This is a challenge for durable execution frameworks that require code to be deterministic. To make these building blocks available to your workflows, DBOS provides durable [`Go`](../reference/workflows-steps.md#go) and [`Select`](../reference/workflows-steps.md#select) functions to run multiple steps concurrently within a workflow while preserving durability guarantees. - **`Go`** launches a step asynchronously and returns a channel for retrieving the result later. - **`Select`** waits for the first result from multiple concurrent steps. #### Running Concurrent Steps with Go Use [`Go`](../reference/workflows-steps.md#go) to launch a step that runs in the background while your workflow continues. The function returns immediately with a channel that will receive the step's result when it completes. ```go func workflow(ctx dbos.Context, _ string) (string, error) { // Launch a step asynchronously resultChan, err := dbos.Go(ctx, func(ctx context.Context) (string, error) { // Perform some work... return "step completed", nil }) if err != nil { return "", err } // Do other work while the step runs... // Wait for the result outcome := <-resultChan if outcome.Err != nil { return "", outcome.Err } return outcome.Result, nil } ``` You can launch multiple steps concurrently: ```go func workflow(ctx dbos.Context, urls []string) ([]string, error) { // Launch multiple steps concurrently var channels []<-chan dbos.StepOutcome[string] for _, url := range urls { url := url // Capture loop variable ch, err := dbos.Go(ctx, func(ctx context.Context) (string, error) { return fetchURL(ctx, url) }) if err != nil { return nil, err } channels = append(channels, ch) } // Collect all results var results []string for _, ch := range channels { outcome := <-ch if outcome.Err != nil { return nil, outcome.Err } results = append(results, outcome.Result) } return results, nil } ``` #### Selecting the First Result Use [`Select`](../reference/workflows-steps.md#select) to wait for the first result from multiple concurrent steps. This is useful for racing multiple operations or implementing timeout patterns. ```go func workflow(ctx dbos.Context, _ string) (string, error) { // Launch two concurrent steps ch1, err := dbos.Go(ctx, func(ctx context.Context) (string, error) { // Query primary database return queryPrimaryDB(ctx) }) if err != nil { return "", err } ch2, err := dbos.Go(ctx, func(ctx context.Context) (string, error) { // Query replica database return queryReplicaDB(ctx) }) if err != nil { return "", err } // Wait for the first result result, err := dbos.Select(ctx, []<-chan dbos.StepOutcome[string]{ch1, ch2}) if err != nil { return "", err } return result, nil } ``` #### Determinism and Recovery `Go` and `Select` maintain workflow determinism by checkpointing: - Each `Go` call is assigned a deterministic step ID, ensuring steps execute in the same order during recovery. - `Select` checkpoints which channel was selected and its value, so replays return the same result regardless of actual execution timing. This means you can safely use `Go` and `Select` for concurrent operations without worrying about non-deterministic behavior during workflow recovery. ### Workflow Guarantees Workflows provide the following reliability guarantees. These guarantees assume that the application and database may crash and go offline at any point in time, but are always restarted and return online. 1. Workflows always run to completion. If a DBOS process is interrupted while executing a workflow and restarts, it resumes the workflow from the last completed step. 2. [Steps](./step-tutorial.md) are tried _at least once_ but are never re-executed after they complete. If a failure occurs inside a step, the step may be retried, but once a step has completed (returned a value or thrown an exception to the calling workflow), it will never be re-executed. If an exception is thrown from a workflow, the workflow **terminates**—DBOS records the exception, sets the workflow status to `ERROR`, and **does not recover the workflow**. This is because uncaught exceptions are assumed to be nonrecoverable. If your workflow performs operations that may transiently fail (for example, sending HTTP requests to unreliable services), those should be performed in [steps with configured retries](./step-tutorial.md#configurable-retries). DBOS provides [tooling](./workflow-management.md) to help you identify failed workflows and examine the specific uncaught exceptions. ### Workflow Versioning and Recovery DBOS **versions** applications and workflows. All workflows are tagged with the application version on which they started. By default, application version is automatically computed from a hash of workflow source code. However, you can set your own version through configuration. ```go dbosContext, err := dbos.NewContext(context.Background(), dbos.Config{ AppName: "dbos-app", DatabaseURL: os.Getenv("DBOS_SYSTEM_DATABASE_URL"), ApplicationVersion: "1.0.0", }) ``` When DBOS tries to recover workflows, it only recovers workflows whose version matches the current application version. This prevents unsafe recovery of workflows that depend on different code. When using versioning, we recommend **blue-green** code upgrades. When deploying a new version of your code, launch new processes running your new code version, but retain some processes running your old code version. Direct new traffic to your new processes while your old processes "drain" and complete all workflows of the old code version. Then, once all workflows of the old version are complete (you can use [`ListWorkflows`](../reference/methods.md#listworkflows) to check), you can retire the old code version. --- ## Upgrading ### Upgrading to v1.0.0 This guide covers migrating an application from the last v0.x release to v1.0.0: ```shell go get github.com/dbos-inc/dbos-transact-golang@latest ``` v1 is a breaking release. The sections below cover every breaking change; [§10](#10-suggested-migration-order-per-app) suggests an order to apply them in. --- #### 1. Core renames: `DBOSContext` → `Context` The central interface was renamed. This is the highest-volume change and is purely mechanical: | v0 | v1 | |---|---| | `dbos.DBOSContext` | `dbos.Context` | | `dbos.NewDBOSContext(ctx, cfg)` | `dbos.NewContext(ctx, cfg)` | | `Config.SqliteSystemDB` | `Config.SQLiteSystemDB` | Workflow and step signatures change accordingly: ```go // v0 func myWorkflow(ctx dbos.DBOSContext, input string) (string, error) // v1 func myWorkflow(ctx dbos.Context, input string) (string, error) ``` `Shutdown` now returns an error (non-nil if the timeout expired before all resources stopped). Use the package-level functions for launch and shutdown — like the other interface methods, `Shutdown` now takes a leading `Client` receiver-argument, so method-form calls must be migrated: ```go // v0 ctx.Launch() ctx.Shutdown(30 * time.Second) // v1 err := dbos.Launch(ctx) err := dbos.Shutdown(ctx, 30*time.Second) // accepts Client or Context ``` Behavioral (no code change): `Launch()` now recovers this executor's PENDING workflows *before* starting the queue runner and scheduler. Behavioral (no code change): a failed `Launch()` is now terminal — the context is torn down (system database closed) and cannot be relaunched. Code that retried `Launch` on the same context after a transient failure must create a fresh context with `NewContext` for each attempt. #### 2. Client redesign — `Client` is now a subset of `Context` `Client` is no longer a separate wrapper type with its own method shapes. It is a formal sub-interface of `Context` (both embed `context.Context`), owning every operation that doesn't require `Launch()`: enqueue, send/get-event/read-stream, workflow management, queue management, schedule management, app-version management, `Shutdown`. Every `Context` **is** a `Client`, so a launched `Context` can be passed anywhere a `Client` is accepted. - `dbos.NewClient(ctx, dbos.ClientConfig{...})` still exists and returns a `Client`. `ClientConfig` gained `SystemDBStartupTimeout` and renamed `SqliteSystemDB` → `SQLiteSystemDB`. - Client interface methods now take a leading `Client` receiver-argument like `Context` methods always did (`c.ListWorkflows(c, ...)`). In practice, always use the package-level generic functions instead. ##### The `Client*` generic helpers are gone The duplicated client-side generics were removed. The single package-level generic functions now accept a `Client` (and therefore also a `Context`): | v0 (client) | v1 (unified) | |---|---| | `dbos.ClientRetrieveWorkflow[R](client, id)` | `dbos.RetrieveWorkflow[R](client, id)` | | `dbos.ClientGetEvent[R](client, wfid, key, timeout)` | `dbos.GetEvent[R](client, wfid, key, timeout)` | | `dbos.ClientReadStream[R](client, wfid, key)` | `dbos.ReadStream[R](client, wfid, key)` | | `dbos.ClientReadStreamAsync[R](client, wfid, key)` | `dbos.ReadStreamAsync[R](client, wfid, key)` | | `dbos.ClientResumeWorkflow[R](client, id)` | `dbos.ResumeWorkflow[R](client, id)` | | `dbos.ClientResumeWorkflows[R](client, ids)` | `dbos.ResumeWorkflows[R](client, ids)` | | `dbos.ClientForkWorkflow[R](client, input)` | `dbos.ForkWorkflow[R](client, input)` | | `dbos.ClientForkWorkflows[R](client, input)` | `dbos.ForkWorkflows[R](client, input)` | | `dbos.ClientTriggerSchedule[R](client, name)` | `dbos.TriggerSchedule[R](client, name)` | | `client.Enqueue(queue, wf, input)` (method) | `dbos.Enqueue[R](client, queue, wf, input)` | ##### `Enqueue` type parameters reordered `Enqueue[P, R]` became `Enqueue[R, P]` so the input type is inferred and you name only the result type: ```go // v0 handle, err := dbos.Enqueue[string, int](client, "queue", "MyWorkflow", "input") // v1 handle, err := dbos.Enqueue[int](client, "queue", "MyWorkflow", "input") ``` Same reorder applies to both debouncer types: `Debouncer[P, R]` / `NewDebouncer[P, R]` → `Debouncer[R, P]` / `NewDebouncer[R, P]`, and `DebouncerClient[P, R]` / `NewDebouncerClient[P, R]` → `DebouncerClient[R, P]` / `NewDebouncerClient[R, P]`. `NewDebouncer`'s type parameters are still inferred from the workflow function, but it now also returns an error (see §7), and explicitly-typed variables must swap: `var d *dbos.Debouncer[string, int]` → `*dbos.Debouncer[int, string]`. `NewDebouncerClient` cannot infer its parameters — every call site must swap them by hand; watch for cases where input and result types are mutually assignable (e.g. both `string`), which compile unchanged with swapped meaning. ##### Functions that widened from `Context` to `Client` These package-level functions now take a `Client` first argument. Existing call sites passing a `Context` keep compiling — listed for completeness: `Send`, `GetEvent`, `ReadStream`, `ReadStreamAsync`, `RetrieveWorkflow`, `CancelWorkflow(s)`, `ResumeWorkflow(s)`, `ForkWorkflow(s)`, `DeleteWorkflows`, `SetWorkflowDelay`, `ListWorkflows`, `GetWorkflowSteps`, `GetWorkflowAggregates`, `GetStepAggregates`, all queue-management, all schedule-management, all app-version functions. Still `Context`-only (require a launched runtime / workflow scope): `RunWorkflow`, `RunAsStep`, `RunAsTransaction`, `Go`, `Select`, `Recv`, `SetEvent`, `WriteStream`, `CloseStream`, `Sleep`, `Patch`, `DeprecatePatch`, `GetWorkflowID`, `GetStepID`, `RegisterWorkflow`, `ListenQueues`. #### 3. Queues ##### `NewWorkflowQueue` removed → `RegisterQueue` In-memory queues are gone; all queues are database-backed. The public `WorkflowQueue` struct was unexported; the `Queue` interface (`GetName()`, `GetGlobalConcurrency()`, `Set*` methods, …) is the public handle. ```go // v0 queue := dbos.NewWorkflowQueue(ctx, "queue", dbos.WithWorkerConcurrency(5)) // before Launch only // v1 queue, err := dbos.RegisterQueue(ctx, "queue", dbos.WithWorkerConcurrency(5)) // works before or after Launch, and from a Client ``` - `RegisterQueue` returns `(Queue, error)` — handle the error. This typically ripples outward: per-module registration helpers that were `func Register(ctx dbos.Context)` become error-returning, and every host binary that calls them needs the plumbing. Budget for that in apps with per-agent registration functions. - `WithMaxTasksPerIteration` option removed (no replacement). - Field access via struct (`queue.Name`) → interface getters (`queue.GetName()`). - Queue `Set*` methods now take a `Client` (any `Context` works). - Registering the reserved internal queue name is an error. ##### `WithQueue` takes a `Queue`, not a name ```go // v0 handle, err := dbos.RunWorkflow(ctx, task, i, dbos.WithQueue(queue.Name)) // v1 handle, err := dbos.RunWorkflow(ctx, task, i, dbos.WithQueue(queue)) ``` Apps that only kept the queue *name* (a common v0 pattern, since `WithQueue` took a string) must now retain the `Queue` handle returned by `RegisterQueue` — typically in the package var where the name constant used to suffice. If only the name is at hand, fetch the handle first with `dbos.RetrieveQueue(ctx, name)` — which now returns an error matching `dbos.ErrQueueNotFound` when absent, instead of `(nil, nil)`. Update any `q == nil` existence checks: ```go // v1 q, err := dbos.RetrieveQueue(ctx, name) if errors.Is(err, dbos.ErrQueueNotFound) { /* absent */ } ``` ##### Listen set - `ListenQueues(ctx, queues ...WorkflowQueue)` → `ListenQueues(ctx, names ...string)`. Each call **replaces** the whole set; empty set = listen to every queue. - New: `ListenedQueues(ctx) []string`. - `ListRegisteredQueues` (deprecated in v0) removed — use `ListQueues(ctx)`. - `DeleteQueue` on a nonexistent queue is not an error. #### 4. Scheduled workflows ##### `WithSchedule` registration option removed Cron scheduling is no longer a registration-time option; schedules are database-backed rows created **after Launch** (or from a Client). The workflow must accept `dbos.ScheduledWorkflowInput`: ```go // v0 dbos.RegisterWorkflow(ctx, func(ctx dbos.DBOSContext, scheduledTime time.Time) (string, error) { ... }, dbos.WithSchedule("* * * * * *")) // v1 reportWorkflow := func(ctx dbos.Context, input dbos.ScheduledWorkflowInput) (any, error) { // input.ScheduledTime is the cron tick time ... } dbos.RegisterWorkflow(ctx, reportWorkflow) // ... after dbos.Launch(ctx): err := dbos.CreateSchedule(ctx, dbos.ScheduleSpec{ ScheduleName: "report-schedule", Schedule: "* * * * * *", Workflow: reportWorkflow, // or WorkflowName: "..." }) ``` ##### One input struct: `ScheduleSpec` `CreateScheduleRequest` + its functional options (`WithScheduleContext`, `WithAutomaticBackfill`, `WithCronTimezone`, `WithScheduleQueueName`, `WithScheduleWorkflowClassName`), `ApplySchedulesRequest`, and `ClientScheduleInput` all collapse into `ScheduleSpec`: ```go type ScheduleSpec struct { ScheduleName string // required Schedule string // required, cron expression WorkflowName string // required unless Workflow is set Workflow any // registered Go workflow fn (Context only; wins over WorkflowName) WorkflowClassName string Context any // serialized as JSON AutomaticBackfill bool CronTimezone string QueueName string // defaults to the internal queue } ``` - `CreateSchedule(ctx, fn, req, opts...)` → `CreateSchedule(ctx, spec)`; `ApplySchedules` takes `[]ScheduleSpec`. - Field rename: `ApplySchedulesRequest.WorkflowFn` → `ScheduleSpec.Workflow`. Renaming the *type* alone leaves `WorkflowFn:` behind at every schedule literal — rename the field too (included in the §10 sed list). - `ScheduledWorkflowInput.Context` is now `json.RawMessage`; decode with the new `dbos.DecodeScheduleContext[T](input)`. *Creating* schedules is unaffected — `ScheduleSpec.Context` stays `any` and still takes your struct directly; the SDK serializes it. Only code that hand-constructs a `ScheduledWorkflowInput` (test harnesses, manual-trigger paths that invoke a scheduled workflow directly to simulate a tick) must now `json.Marshal` the context into the field. > **Gob-serializer apps: drain scheduled firings before upgrading.** `ScheduledWorkflowInput.Context` changed from `any` to `json.RawMessage` under the same gob registration, and gob cannot decode the old interface-encoded field into the new concrete type. Any schedule firing still ENQUEUED (or otherwise not yet dequeued) at upgrade time in an app using `NewGobSerializer` will fail input decoding and error when it runs — those workflows do not make it through the upgrade. Let pending firings drain (or cancel them) before deploying v1. Apps on the default JSON serializer are unaffected. - `GetSchedule` returns `(WorkflowSchedule, error)` by value (was `*WorkflowSchedule`), with `dbos.ErrScheduleNotFound` when absent. - `TriggerSchedule` is now generic: `TriggerSchedule[R](ctx, name)`. #### 5. Errors `DBOSError`/`DBOSErrorCode` renamed and sentinels introduced: | v0 | v1 | |---|---| | `*dbos.DBOSError` | `*dbos.Error` | | `dbos.DBOSErrorCode` | `dbos.ErrorCode` | | `ConflictingIDError` | `ErrorCodeConflictingID` | | `InitializationError` | `ErrorCodeInitialization` | | `NonExistentWorkflowError` | `ErrorCodeNonExistentWorkflow` | | `ConflictingWorkflowError` | `ErrorCodeUnexpectedWorkflow` (renamed concept) | | `WorkflowCancelled` | `ErrorCodeWorkflowCancelled` | | `UnexpectedStep` | `ErrorCodeUnexpectedStep` | | `AwaitedWorkflowCancelled` | `ErrorCodeAwaitedWorkflowCancelled` | | `ConflictingRegistrationError` | `ErrorCodeConflictingRegistration` | | `WorkflowUnexpectedTypeError` | `ErrorCodeWorkflowUnexpectedType` | | `WorkflowExecutionError` | `ErrorCodeWorkflowExecution` | | `StepExecutionError` | `ErrorCodeStepExecution` | | `DeadLetterQueueError` | `ErrorCodeDeadLetterQueue` | | `MaxStepRetriesExceeded` | `ErrorCodeMaxStepRetriesExceeded` | | `QueueDeduplicated` | `ErrorCodeQueueDeduplicated` | | `PatchingNotEnabled` | `ErrorCodePatchingNotEnabled` | | `TimeoutError` | `ErrorCodeTimeout` | | `NoApplicationVersions` | `ErrorCodeNoApplicationVersions` | | — (new) | `ErrorCodeQueueNotFound`, `ErrorCodeScheduleNotFound`, `ErrorCodeInvalidOption` | Prefer the new sentinels with `errors.Is` over constructing an `&Error{Code: ...}` probe: ```go // v0 if errors.Is(err, &dbos.DBOSError{Code: dbos.QueueDeduplicated}) { ... } // v1 if errors.Is(err, dbos.ErrQueueDeduplicated) { ... } ``` Sentinels: `ErrWorkflowCancelled`, `ErrAwaitedWorkflowCancelled`, `ErrQueueDeduplicated`, `ErrNonExistentWorkflow`, `ErrConflictingWorkflowID`, `ErrUnexpectedWorkflow`, `ErrMaxStepRetriesExceeded`, `ErrDeadLetterQueue`, `ErrTimeout`, `ErrQueueNotFound`, `ErrScheduleNotFound`, `ErrNoApplicationVersions`, `ErrInvalidOption`. Timeout errors from an expired deadline also match `context.DeadlineExceeded`; context-driven cancellations (durable timeout, cancelled context) carry a cause matching `context.Canceled` or `context.DeadlineExceeded`. Every code and sentinel is documented in the [errors reference](./reference/workflows-steps.md#errors). Type assertions on the error struct: `var e *dbos.DBOSError; errors.As(err, &e)` → `var e *dbos.Error; errors.As(err, &e)`. #### 6. Option renames ##### `ListWorkflows` filters: `With*` → `WithFilter*` (several became variadic) | v0 | v1 | |---|---| | `WithWorkflowIDs(ids []string)` | `WithFilterWorkflowIDs(ids ...string)` | | `WithStatus(s []WorkflowStatusType)` | `WithFilterStatus(s ...WorkflowStatusType)` | | `WithStartTime(t)` | `WithFilterCreatedAfter(t)` | | `WithEndTime(t)` | `WithFilterCreatedBefore(t)` | | `WithName(n ...string)` | `WithFilterName(n ...string)` | | `WithAppVersion(v ...string)` | `WithFilterAppVersion(v ...string)` | | `WithUser(u ...string)` | `WithFilterUser(u ...string)` | | `WithLimit(n)` / `WithOffset(n)` | `WithFilterLimit(n)` / `WithFilterOffset(n)` | | `WithSortDesc()` | `WithFilterSortDesc()` | | `WithWorkflowIDPrefix(p ...string)` | `WithFilterWorkflowIDPrefix(p ...string)` | | `WithLoadInput(b)` / `WithLoadOutput(b)` | `WithFilterLoadInput(b)` / `WithFilterLoadOutput(b)` | | `WithQueueName(q ...string)` | `WithFilterQueueName(q ...string)` | | `WithQueuesOnly()` | `WithFilterQueuesOnly()` | | `WithExecutorIDs(ids []string)` | `WithFilterExecutorIDs(ids ...string)` | | `WithForkedFrom(f ...string)` | `WithFilterForkedFrom(f ...string)` | | `WithParentWorkflowID(p ...string)` | `WithFilterParentWorkflowID(p ...string)` | | `WithCompletedAfter/Before(t)` | `WithFilterCompletedAfter/Before(t)` | | `WithDequeuedAfter/Before(t)` | `WithFilterDequeuedAfter/Before(t)` | | `WithWasForkedFrom(b)` | `WithFilterWasForkedFrom(b)` | | `WithHasParent(b)` | `WithFilterHasParent(b)` | Slice-taking filters became variadic — call sites passing a literal slice need `...` or unwrapping: `WithStatus([]dbos.WorkflowStatusType{dbos.WorkflowStatusSuccess})` → `WithFilterStatus(dbos.WorkflowStatusSuccess)`. ##### Step retry options | v0 | v1 | |---|---| | `WithBackoffFactor(f)` | `WithStepBackoffFactor(f)` | | `WithBaseInterval(d)` | `WithStepBaseInterval(d)` | | `WithMaxInterval(d)` | `WithStepMaxInterval(d)` | | `WithRetryPredicate(fn)` | `WithStepRetryPredicate(fn)` | | — (new) | `WithTxIsolation(level IsoLevel)` | ##### Other option/function renames | v0 | v1 | |---|---| | `WithMaxRetries(n)` (workflow registration) | `WithMaxRecoveryAttempts(n)` | | `CancelWorkflowOptions` (type) | `CancelWorkflowOption` | | `WithAuthenticatedRoles(roles []string)` | `WithAuthenticatedRoles(roles ...string)` | | `WithEnqueueAuthenticatedRoles(roles []string)` | `WithEnqueueAuthenticatedRoles(roles ...string)` | | `UpdateWorkflowAttributes(ctx, id, attrs)` | `SetWorkflowAttributes(ctx, id, attrs)` | | `WithReadStreamSnapshot(fromOffset int)` | `WithReadStreamSnapshot()` + `WithReadStreamFromOffset(offset int)` | #### 7. Miscellaneous API changes - `Go[R](ctx, fn)` now returns a receive-only `<-chan StepOutcome[R]` (was `chan`). Only breaks call sites that stored it in an explicitly-typed bidirectional variable. - `GobSerializer` struct unexported — construct via `dbos.NewGobSerializer()` (unchanged). Persisted v0 gob error payloads remain decodable. - `ListRegisteredWorkflows(ctx)` returns `[]WorkflowRegistryEntry` — no error, no options; `WithScheduledOnly` / `ListRegisteredWorkflowsOption` removed. - `GetLatestApplicationVersion` returns `(VersionInfo, error)` by value (was `*VersionInfo`). `nil` checks → `errors.Is(err, dbos.ErrNoApplicationVersions)`. - `ClearRegistries` removed from the public API (was testing-only). - New package-level accessors: `dbos.GetApplicationVersion(ctx)`, `dbos.GetExecutorID(ctx)`, `dbos.GetApplicationID(ctx)`. - `Config.DatabaseURL` accepts Postgres/CockroachDB URLs or key=value DSNs, and sqlite URLs (`sqlite:/path/to.db`, `sqlite:relative.db`, `sqlite::memory:`). - **SQLite now requires a driver import.** The SQLite driver moved out of the core package into `dbos/driver/sqlite`, registered database/sql-style. Apps using a sqlite URL or `Config.SQLiteSystemDB` must add `import _ "github.com/dbos-inc/dbos-transact-golang/dbos/driver/sqlite"` (one blank import, anywhere in the binary); without it, startup fails with an error naming this import. Postgres-only apps need no change and no longer compile or link `modernc.org/sqlite`. - **One additive system-database schema change.** The first v1 startup applies an additive migration that adds two defaulted columns to `workflow_status`. Existing workflows (including long-DELAYED ones), queues, and schedule rows carry over as-is, and v0.x and v1 executors can share a system database during a rolling upgrade: old executors ignore the new columns and skip migration versions they don't know. - **Debouncer changes.** `NewDebouncer` now returns `(*Debouncer[R, P], error)` — call sites gain error handling (type parameters are still inferred). Debouncers can be created at any time, including after `Launch()` (v0 required creation before launch). New options: `WithDebouncerQueue(name)` runs the debounced workflow on a named registered queue instead of the DBOS internal queue, and `WithDebouncerClassName(name)` (client-side) targets workflows registered under a class name. `Debounce` now rejects options a debounce owns or cannot support (`WithQueue`, `WithDeduplicationID`, `WithDelay`, `WithPriority`, `WithQueuePartitionKey`, `WithDeduplicationPolicy`) with an error matching `dbos.ErrInvalidOption`. Pending debounced workflows are now visible to workflow management: they appear as `DELAYED` on their queue with `IsDebounced` set and can be listed with `dbos.WithFilterIsDebounced(true)`. The debounce timeout is captured by the call that enqueues the workflow — later calls coalescing on the same pending workflow extend the delay up to that fixed deadline but never move it — and a deadline on the `Debounce` context is recorded as the debounced workflow's execution timeout (the clock starts at dequeue), never as an absolute deadline. - **Drain debouncers before upgrading.** A debounce still pending when you deploy v1 will not fire under the new version. Before upgrading, stop calling `Debounce` and let in-flight debounces fire (this takes at most the debounce delay or timeout); cancel any stragglers left after the upgrade with `CancelWorkflow`. ##### Custom serializers must handle non-user values Every checkpoint written to the system database is encoded with the configured custom serializer — including engine-internal step outputs: the `int64` deadline recorded by `DBOS.sleep` (written by `Sleep` and, new in v1, by `Recv`/`GetEvent` timeouts), `DBOS.writeStream`'s empty-string step output, and the `ScheduledWorkflowInput` struct for DB-backed schedule firings. This keeps every row in one format when inspecting the database from another program. A `Serializer[any]` must therefore be total: it must encode and decode arbitrary Go values, and decoding must preserve concrete types (decoded step results are type-asserted, so an `int64` must round-trip as `int64`, not `float64`). Strictly-typed serializers need a mapping for bare Go scalars — see the [`protobuf-serializer` demo app](https://github.com/dbos-inc/dbos-demo-apps/tree/main/golang/protobuf-serializer), which packs `proto.Message` values into `anypb.Any` directly and wraps scalars in the protobuf well-known wrapper types (`wrapperspb.Int64Value`, `StringValue`, …), unwrapping them back to native scalars on decode. This keeps every stored row decodable by standard proto tooling in any language. See the [serialization reference](./reference/workflows-steps.md#serialization) for the full serializer contract. One deliberate exception: `Recv` and `GetEvent` checkpoint the *sender's* encoded payload verbatim under the sender's recorded format — the receiver's serializer is never asked to re-encode a message or event it didn't produce. #### 8. DBOS calls are rejected inside step bodies v0 let most DBOS operations run inside a step body and silently dropped their durability guarantees: `CloseStream`, workflow-management writes, and `RunAsTransaction` invoked from inside a step executed in a real database transaction but recorded **no checkpoint**, so recovery could repeat them. v1 rejects them instead with an `ErrorCodeStepExecution` error (`cannot call within a step`). Newly rejected inside a step (v0 already rejected `Send`, `Recv`, `GetEvent`, `Sleep`, `Patch`, `DeprecatePatch`, and spawning a child workflow with `RunWorkflow`): - `RunAsTransaction` - `Enqueue` - `Go` - `handle.GetResult()` — awaiting another workflow's result from inside a step - `CloseStream` - Workflow-management writes: `CancelWorkflow(s)`, `ResumeWorkflow(s)`, `ForkWorkflow(s)`, `DeleteWorkflows`, `SetWorkflowAttributes`, `SetWorkflowDelay` - Schedule writes: `CreateSchedule`, `PauseSchedule`, `ResumeSchedule`, `DeleteSchedule` - `Debounce` Still allowed inside a step: - [`SetEvent`](./tutorials/workflow-communication.md#workflow-events) and [`WriteStream`](./tutorials/workflow-communication.md#workflow-streaming) — at-least-once, attributed to the enclosing step (a retried step may write again). - Read/list operations — `ListWorkflows`, `GetWorkflowSteps`, `RetrieveWorkflow`, `ReadStream`, `GetWorkflowAggregates`, `GetStepAggregates`, `GetSchedule`, `ListSchedules`, queue reads, app-version reads. Called from a step they run within the enclosing step's durability scope, without their own checkpoint; called from workflow code they are checkpointed as steps, as before. - Nested [`RunAsStep`](./tutorials/step-tutorial.md) — the inner function runs inline as part of the enclosing step, without its own checkpoint (unchanged from v0). To migrate, hoist rejected calls out of step bodies into the surrounding workflow code. If a step needs another workflow's result, return from the step and await the handle in the workflow. #### 9. Mocks / tests Mocks of the old `DBOSContext` must be regenerated from `dbos.Context` (which now embeds `Client`). With mockery, point the config at `Context` — see the [Testing & Mocking tutorial](./tutorials/testing.md) for a sample `.mockery.yml`. Assertions on error types/codes need the §5 renames. #### 10. Suggested migration order per app 1. `go get github.com/dbos-inc/dbos-transact-golang@latest`; `go build ./...` to enumerate breakage. 2. Mechanical renames (safe to sed, word-boundary matched): `DBOSContext`→`Context`, `NewDBOSContext`→`NewContext`, `DBOSError`→`Error`, `DBOSErrorCode`→`ErrorCode`, `SqliteSystemDB`→`SQLiteSystemDB`, `UpdateWorkflowAttributes`→`SetWorkflowAttributes`, `CancelWorkflowOptions`→`CancelWorkflowOption`, step-retry `With*`→`WithStep*`, list-filter `With*`→`WithFilter*` (per §6 table — names collide with unrelated options, so match exact identifiers), error-code constants per §5, `Client`→`` per §2, `WorkflowFn:`→`Workflow:` (schedule-spec literals, per §4). 3. Structural fixes by hand: queue registration (`NewWorkflowQueue`→`RegisterQueue` + error handling + `WithQueue(queue)`), scheduled workflows (`WithSchedule`→`CreateSchedule`+`ScheduleSpec`, input signature), `Enqueue` type-param reorder, `RetrieveQueue`/`GetSchedule`/`GetLatestApplicationVersion`/`NewDebouncer` return-shape changes, `Shutdown` error handling, hoisting DBOS calls out of step bodies (§8). 4. `go vet ./...` && `go build ./...` && run tests. --- ## Welcome to DBOS! DBOS is a library for building reliable programs. Add a few annotations to your application to **durably execute** it and make it **resilient to any failure**. #### Get Started import { TbHexagonNumber1, TbHexagonNumber2, TbHexagonNumber3, TbHexagonNumber4 } from "react-icons/tb"; import { FaHackerNews } from "react-icons/fa6";
#### Example Applications import { MdOutlineShoppingCart } from "react-icons/md"; import { PiFileMagnifyingGlassBold } from "react-icons/pi"; import { VscGraphLine } from "react-icons/vsc";
#### Features import { IoIosRocket } from "react-icons/io"; import { BsDatabaseCheck } from "react-icons/bs"; import { SiOpentelemetry, SiApachekafka } from "react-icons/si"; import { RiCalendarScheduleLine, RiRewindStartMiniLine } from "react-icons/ri"; import { PiQueueBold } from "react-icons/pi";
#### Join the Community If you have any questions or feedback about DBOS, you can reach out to DBOS community members and developers on our [Discord server](https://discord.gg/fMwQjeW5zg).
--- ## CockroachDB Here's how to connect your DBOS application running on your computer or cloud environment to your CockroachDB database. #### 1. Set up a Local Application If you haven't already, follow the [quickstart](../quickstart.md) to set up a DBOS application locally. The rest of this guide will assume you have a local application. #### 2. Install the CockroachDB driver Install a CockroachDB-compatible PostgreSQL driver: ```shell pip install psycopg2-binary sqlalchemy-cockroachdb ``` #### 3. Connect to your CockroachDB Database Retrieve your CockroachDB database connection information. Then create a connection string with the following format: ``` cockroachdb://user:password@host:port/database ``` Be sure to specify the `cockroachdb://` driver! Export this as an environment variable: ``` export DBOS_COCKROACHDB_URL="" ``` #### 4. Configure Your DBOS Application Now, configure your DBOS application to connect to CockroachDB as follows: ```python import os from sqlalchemy import create_engine from dbos import DBOS, DBOSConfig database_url = os.environ.get("DBOS_COCKROACHDB_URL") engine = create_engine(database_url) config: DBOSConfig = { "name": "dbos-app", "application_version": "0.1.0", "system_database_url": database_url, # Create a custom SQLAlchemy engine to utilize the CockroachDB drivers "system_database_engine": engine, # CockroachDB does not support LISTEN/NOTIFY "use_listen_notify": False, } DBOS(config=config) DBOS.launch() ``` When you launch your application, it should connect to your CockroachDB database! #### 1. Set up a Local Application If you haven't already, follow the [quickstart](../quickstart.md) to set up a DBOS application locally. The rest of this guide will assume you have a local application. #### 2. Connect to your CockroachDB Database Retrieve your CockroachDB database connection information. CockroachDB is PostgreSQL wire-compatible, so you can use a standard PostgreSQL connection string: ``` postgresql://user:password@host:port/database ``` Export this as an environment variable: ``` export DBOS_COCKROACHDB_URL="" ``` #### 3. Configure Your DBOS Application Now, configure your DBOS application to connect to CockroachDB as follows: ```typescript DBOS.setConfig({ name: 'my-application', applicationVersion: '0.1.0', // Your CockroachDB connection string. systemDatabaseUrl: process.env.DBOS_COCKROACHDB_URL, // CockroachDB does not support LISTEN/NOTIFY useListenNotify: false, }); await DBOS.launch(); ``` When you launch your application, it should connect to your CockroachDB database! #### 1. Set up a Local Application If you haven't already, follow the [quickstart](../quickstart.md) to set up a DBOS application locally. The rest of this guide will assume you have a local application. #### 2. Connect to your CockroachDB Database Retrieve your CockroachDB database connection information. CockroachDB is PostgreSQL wire-compatible, so you can use a standard PostgreSQL connection string: ``` postgresql://user:password@host:port/database ``` Export this as an environment variable: ``` export DBOS_SYSTEM_DATABASE_URL="" ``` #### 3. Configure Your DBOS Application The Go SDK **automatically detects** CockroachDB when it connects to the database. No special configuration or drivers are needed—just provide your CockroachDB connection URL as you would for PostgreSQL: ```go dbosContext, err := dbos.NewDBOSContext(context.Background(), dbos.Config{ AppName: "dbos-app", ApplicationVersion: "0.1.0", DatabaseURL: os.Getenv("DBOS_SYSTEM_DATABASE_URL"), }) if err != nil { log.Fatal(err) } ``` When you launch your application, DBOS will detect CockroachDB and automatically adjust its behavior: - LISTEN/NOTIFY (not supported by CockroachDB) is replaced with polling-based notifications. - Schema migrations are adapted for CockroachDB compatibility. --- ## Django ## DBOS + Django This guide shows you how to add DBOS durable workflows to your existing Django application to make it resilient to any failure. :::info The guide is bootstrapped from the Django [quickstart](https://docs.djangoproject.com/en/5.2/intro/tutorial01/). See its source code on [GitHub](https://github.com/dbos-inc/dbos-demo-apps/tree/main/python/dbos-django-starter). ::: ### Installation and Requirements Install DBOS Python with: ```shell pip install dbos ``` ### Configuring and Launching DBOS In your Django application `AppConfig`, configure and launch DBOS inside the `ready` method: ```python title="polls/apps.py" import os from django.apps import AppConfig from dbos import DBOS, DBOSConfig class PollsConfig(AppConfig): default_auto_field = 'django.db.models.BigAutoField' name = 'polls' def ready(self): dbos_config: DBOSConfig = { "name": "django-app", "application_version": "0.1.0", "system_database_url": os.environ.get("DBOS_SYSTEM_DATABASE_URL"), } DBOS(config=dbos_config) DBOS.launch() return super().ready() ``` :::tip If you're starting your Django app with `python manage.py runserver`, we recommend setting the `--noreload` flag during development to avoid repeatedly relaunching DBOS. ::: ### Making Your Views Reliable Once you've configured DBOS, you can add durable workflows to your application. They're just ordinary Python functions you can call from your services or views! For example, here's a simple `callWorkflow` view that invokes a workflow of two steps: ```python title="polls/views.py" def callWorkflow(request, a, b): return JsonResponse(workflow(a, b)) @DBOS.step() def step_one(a): print("Step one completed!", a) @DBOS.step() def step_two(b): print("Step two completed!", b) @DBOS.workflow() def workflow(a, b): step_one(a) step_two(b) return {"result": "success"} ``` --- ## Google ADK You can use DBOS to add durable execution to an agent built with [Google ADK](https://adk.dev/). With durable execution, you can build reliable agents that preserve progress across transient API failures, application errors, and restarts, while also handling long-running, asynchronous, and human-in-the-loop workflows with production-grade reliability. ### Installation To get started, install DBOS and the [durable Google ADK agents integration](https://github.com/dbos-inc/dbos-google-adk). ``` pip install dbos dbos-google-adk ``` You may also need a [Gemini API key](https://aistudio.google.com/app/api-keys) (or any [supported model](https://adk.dev/agents/models/)). ### Building Reliable Agents To wrap your Google ADK agents for durable execution, [install and configure DBOS](../python/integrating-dbos.md) then follow these three guidelines: 1. Add `DBOSPlugin` to your `Runner`'s list of plugins. 2. Annotate the function calling `runner.run_async` with `@DBOS.workflow` to run your agent as a durably executed workflow. 3. Annotate your agent's tool call functions and guardrail functions with `@DBOS.step` or `@DBOS.workflow()` to mark them as steps or sub-workflows of your durably executed agentic workflow. Here is a simple but complete example of wrapping an agent for durable execution. With < 10 lines of code (highlighted below), you can add DBOS into an existing Google ADK application. ```python title="dbos_agent.py" import asyncio import logging # highlight-start from dbos import DBOS, DBOSConfig from dbos_google_adk import DBOSPlugin # highlight-end from google.adk.agents import LlmAgent from google.adk.runners import Runner from google.adk.sessions import InMemorySessionService from google.genai import types # Decorate tool calls with @DBOS.step() for durable execution # highlight-next-line @DBOS.step() async def get_weather(city: str) -> str: """Get the weather for a city.""" return f"Sunny in {city}" agent = LlmAgent(name="weather", model="gemini-flash-latest", tools=[get_weather]) runner = Runner( app_name="my-agent", agent=agent, # highlight-next-line plugins=[DBOSPlugin()], session_service=InMemorySessionService(), ) # Drive the agent from a DBOS workflow for durable execution # highlight-next-line @DBOS.workflow() async def run_agent(user_id: str, session_id: str, message: str) -> str: new_message = types.Content(role="user", parts=[types.Part.from_text(text=message)]) async for event in runner.run_async( user_id=user_id, session_id=session_id, new_message=new_message ): if event.is_final_response(): return event.content.parts[0].text return "" async def main(): # highlight-start # DBOS checkpoints to SQLite by default. Postgres is recommended for production. config: DBOSConfig = {"name": "my-agent", "application_version": "0.1.0", "system_database_url": "sqlite:///dbostest.sqlite"} DBOS(config=config) DBOS.launch() # highlight-end await runner.session_service.create_session( app_name="my-agent", user_id="u", session_id="s" ) print(await run_agent("u", "s", "How is the weather in San Francisco?")) if __name__ == "__main__": asyncio.run(main()) ``` **Notes:** 1. Workflows and agents must be defined before [`DBOS.launch()`](../python/reference/dbos-class.md#launch). 2. This example uses SQLite for ease of getting started. Postgres is recommended for production. ### Learn More For more details on building agents, see the [Google ADK documentation](https://adk.dev/). For information about durable execution and workflow design, see the [DBOS programming guide](../python/programming-guide). Together, these resources cover everything from getting started with simple agents to designing production-ready, fault-tolerant applications. --- ## Lakebase ## Use DBOS With Lakebase Here's how to connect your DBOS application running on your computer or cloud environment to your Databricks Lakebase Postgres database. #### 1. Set up a Local Application If you haven't already, follow the [quickstart](../quickstart.md) to set up a DBOS application locally. The rest of this guide will assume you have a local application. #### 2. Connect to Lakebase From the "Roles & Databases" tab of your Lakebase console, create a Postgres role that you will use to access your Lakebase database from your DBOS application. Next, click "Connect" on your Lakebase console to retrieve connection information for your Lakebase database. You should see a screen that looks like this: This page shows the connection string for your database. There are a few settings you may wish to alter before retrieving this connection string: 1. Select the Postgres role you created from the dropdown. 2. Make sure you are viewing your "Connection string" from the dropdown. 3. By default, you will use the `databricks_postgres` database. If you want to use a different database, create it from the dropdown. When you are ready, copy the connection string (including the password) from the dashboard and set the `DBOS_SYSTEM_DATABASE_URL` environment variable to it: ``` export DBOS_SYSTEM_DATABASE_URL="" ``` :::info Programmatic OAuth token generation through the Databricks SDK is only available in Python. ::: First, install the Databricks SDK into your application, then follow [this guide](https://docs.databricks.com/aws/en/dev-tools/sdk-python#authentication) to authenticate: ```shell pip install databricks-sdk ``` Then, add this code to your application to configure DBOS to retrieve and use a Databricks OAuth token for authentication with your Lakebase database:
Databricks OAuth Authentication ```python from databricks.sdk import WorkspaceClient import psycopg import sqlalchemy as sa def create_databricks_oauth_engine(database_url: str, lakebase_endpoint: str) -> sa.Engine: """ Create a SQLAlchemy engine that uses Databricks OAuth tokens for authentication. Each new connection will fetch a fresh OAuth token from Databricks. """ workspace_client = WorkspaceClient() def get_connection() -> psycopg.Connection: token = workspace_client.postgres.generate_database_credential( endpoint=lakebase_endpoint ).token return psycopg.connect(database_url, password=token) return sa.create_engine("postgresql+psycopg://", creator=get_connection) if __name__ == "__main__": engine = create_databricks_oauth_engine( os.environ["DBOS_SYSTEM_DATABASE_URL"], os.environ["DATABRICKS_LAKEBASE_ENDPOINT"], ) config: DBOSConfig = { "name": "my-app", "application_version": "0.1.0", "system_database_engine": engine, } DBOS(config=config) DBOS.launch() uvicorn.run(app, host="0.0.0.0", port=8000) ```
Next, click "Connect" on your Lakebase console to retrieve connection information for your Lakebase database. You should see a screen that looks like this: This page shows the connection string for your database. There are a few settings you may wish to alter before retrieving this connection string: 1. Make sure you are viewing your "Connection string" from the dropdown. 2. By default, you will use the `databricks_postgres` database. If you want to use a different database, create it from the dropdown. When you are ready, copy the connection string from the dashboard and set the `DBOS_SYSTEM_DATABASE_URL` environment variable to it: ``` export DBOS_SYSTEM_DATABASE_URL="" ``` Also, set the `DATABRICKS_LAKEBASE_ENDPOINT` environment variable to your Lakebase endpoint: ``` export DATABRICKS_LAKEBASE_ENDPOINT="projects//branches//endpoints/" ```
#### 3. Launch Your Application Now, launch your DBOS application. It should successfully connect to your Lakebase database, printing your masked Lakebase database URL on startup. After connecting your DBOS application to Lakebase, you can use the Lakebase console to view your DBOS system tables. Open the "Tables" tab in the Lakebase console. For your database, select the database to which you connected your DBOS application (default `databricks_postgres`). For your schema, select "dbos". You can now see DBOS durably checkpoint your workflows to your Lakebase database: --- ## LlamaIndex ## LlamaIndex Durable Workflows with DBOS [LlamaIndex Workflows](https://developers.llamaindex.ai/python/llamaagents/workflows/) is a Python framework for orchestrating AI agents by composing steps and events into structured workflows. This guide explains how to build durable LlamaIndex agents using the DBOS runtime, enabling fault-tolerant, persistent AI workflows that can safely recover from crashes or restarts. By integrating DBOS with LlamaIndex Workflows through the `llama-agents-dbos` package, every workflow transition is automatically persisted. This allows long-running AI workflows to resume exactly where they left off, without requiring manual checkpointing or snapshot logic. This is especially useful for: - AI agents with long-running tasks - Multi-step LlamaIndex workflows - LLM pipelines that must survive failures - Production AI systems that require reliability :::info Also check out the integration guide in [the LlamaIndex docs](https://developers.llamaindex.ai/python/llamaagents/workflows/dbos/)! ::: ### Installation To get started, install the [`llama-agents-dbos`](https://github.com/run-llama/workflows-py/tree/main/packages/llama-agents-dbos) package. **pip** ```shell pip install llama-agents-dbos ``` **uv** ```shell uv add llama-agents-dbos ``` ### Quick Start: Standalone Durable Workflow The example below defines a simple workflow that counts from 0 to 20. DBOS persists each transition so the workflow can safely resume after a crash or restart. ```python import asyncio from dbos import DBOS, DBOSConfig from llama_agents.dbos import DBOSRuntime from pydantic import Field from workflows import Context, Workflow, step from workflows.events import Event, StartEvent, StopEvent # 1. Configure DBOS — SQLite by default config: DBOSConfig = { "name": "llamaindex-counter-example", "application_version": "0.1.0", "system_database_url": "sqlite:///counter_example.sqlite", } DBOS(config=config) # 2. Define events and workflow (nothing DBOS-specific here) class Tick(Event): count: int = Field(description="Current count") class CounterResult(StopEvent): final_count: int = Field(description="Final counter value") class CounterWorkflow(Workflow): @step async def start(self, ctx: Context, ev: StartEvent) -> Tick: await ctx.store.set("count", 0) print("[Start] Initializing counter to 0") return Tick(count=0) @step async def increment(self, ctx: Context, ev: Tick) -> Tick | CounterResult: count = ev.count + 1 await ctx.store.set("count", count) print(f"[Tick {count:2d}] count = {count}") if count >= 20: return CounterResult(final_count=count) await asyncio.sleep(0.5) return Tick(count=count) # 3. Create runtime, attach to workflow, and launch runtime = DBOSRuntime() workflow = CounterWorkflow(runtime=runtime) async def main() -> None: await runtime.launch() result = await workflow.run(run_id="counter-run-1") print(f"Result: final_count = {result.final_count}") asyncio.run(main()) ``` If you kill the process mid-run (e.g. Ctrl+C at tick 8), calling `workflow.run(run_id="counter-run-1")` again will resume from tick 8 instead of restarting from zero. **Notes:** 1. Workflows must be defined before `runtime.launch()`. 2. The `run_id` parameter uniquely identifies a workflow run, which is equivalent to DBOS workflow ID. 3. This example uses SQLite for ease of getting started. Postgres is recommended for production. ### Durable Workflow Server `DBOSRuntime` integrates with [`WorkflowServer`](https://developers.llamaindex.ai/python/llamaagents/workflows/deployment/) to serve workflows over HTTP with durable execution out of the box. Pass `runtime.create_workflow_store()` as the persistence backend and `runtime.build_server_runtime()` as the execution engine: ```python import asyncio from dbos import DBOS, DBOSConfig from llama_agents.dbos import DBOSRuntime from llama_agents.server import WorkflowServer from pydantic import Field from workflows import Context, Workflow, step from workflows.events import Event, StartEvent, StopEvent config: DBOSConfig = { "name": "llamaindex-server", "application_version": "0.1.0", "system_database_url": "sqlite:///server_example.sqlite", } DBOS(config=config) class Tick(Event): count: int = Field(description="Current count") class CounterResult(StopEvent): final_count: int = Field(description="Final counter value") class CounterWorkflow(Workflow): """Counts to 5, emitting stream events along the way.""" @step async def start(self, ctx: Context, ev: StartEvent) -> Tick: return Tick(count=0) @step async def tick(self, ctx: Context, ev: Tick) -> Tick | CounterResult: count = ev.count + 1 ctx.write_event_to_stream(Tick(count=count)) print(f" tick {count}") await asyncio.sleep(0.5) if count >= 5: return CounterResult(final_count=count) return Tick(count=count) async def main() -> None: runtime = DBOSRuntime() server = WorkflowServer( workflow_store=runtime.create_workflow_store(), runtime=runtime.build_server_runtime(), ) server.add_workflow("counter", CounterWorkflow(runtime=runtime)) print("Serving on http://localhost:8000") print("Try: curl -X POST http://localhost:8000/workflows/counter/run") await server.start() try: await server.serve(host="0.0.0.0", port=8000) finally: await server.stop() asyncio.run(main()) ``` The LlamaIndex workflow debugger UI at `http://localhost:8000/` works exactly the same as with the default runtime: DBOS is transparent to the LlamaIndex workflow server layer. ### Learn More For more details on building agents and workflows with LlamaIndex, see the [LlamaAgents documentation](https://developers.llamaindex.ai/python/llamaagents/workflows/). For information about durable execution and workflow design, see the [DBOS programming guide](../python/programming-guide). Together, these resources cover everything from getting started with simple agents to designing production-ready, fault-tolerant applications. --- ## Logfire ## Use DBOS With Logfire [Pydantic Logfire](https://logfire.pydantic.dev/docs/) is an observability platform built on OpenTelemetry that makes it easy to monitor your application. This guide shows how to configure your DBOS application to export OpenTelemetry traces and logs to Logfire. ### Installation and Requirements First, set up your Logfire account by following the instructions [here](https://logfire.pydantic.dev/docs/#logfire). Next, generate a [write token](https://logfire.pydantic.dev/docs/how-to-guides/create-write-tokens/) so your application can push data to Logfire. If you're not using the Logfire SDK, follow the instructions in [Configure DBOS OpenTelemetry Export](#configure-dbos-opentelemetry-export). If you're using the Logfire SDK, follow the instructions in [Configure DBOS with Logfire SDK](#configure-dbos-with-logfire-sdk). ### Configure DBOS OpenTelemetry Export The following section assumes you're not using the Logfire SDK. If you're using the Logfire SDK, skip this section and see [Configure DBOS with Logfire SDK](#configure-dbos-with-logfire-sdk). First, set the `OTEL_EXPORTER_OTLP_HEADERS` environment variable to include your token: ```shell export OTEL_EXPORTER_OTLP_HEADERS='Authorization=your-write-token' ``` :::tip If you're deploying your app on DBOS Cloud, make sure to set `OTEL_EXPORTER_OTLP_HEADERS` in your application's [environment variables](../conductor/reference/dbos-cloud/secrets). ::: Then, configure your DBOS application to enable OpenTelemetry traces and export them to Logfire: ```python dbos_config: DBOSConfig = { "name": "my-app", "application_version": "0.1.0", # highlight-start "otlp_traces_endpoints": ["https://logfire-us.pydantic.dev/v1/traces"], "otlp_logs_endpoints": ["https://logfire-us.pydantic.dev/v1/logs"], "enable_otlp": True, # highlight-end } DBOS(config=dbos_config) ``` ```typescript DBOS.setConfig({ name: 'my-app', applicationVersion: '0.1.0', // highlight-start otlpTracesEndpoints: ["https://logfire-us.pydantic.dev/v1/traces"], otlpLogsEndpoints: ["https://logfire-us.pydantic.dev/v1/logs"], enableOTLP: true, // highlight-end }); await DBOS.launch(); ``` :::tip This page shows https://logfire-us.pydantic.dev as the base URL which is for the US region. If you are using the EU region, use https://logfire-eu.pydantic.dev instead. ::: Now start your DBOS application. You should see your logs and traces appear on the Logfire dashboard! ![DBOS Logs and Traces on Logfire](./assets/logfire-screenshot.png) ### Configure DBOS with Logfire SDK If you're using the Logfire SDK (currently only supported in Python), set the `LOGFIRE_TOKEN` environment variable to your write token: ```shell export LOGFIRE_TOKEN='your-write-token' ``` :::tip If you're deploying your app on DBOS Cloud, make sure to set `LOGFIRE_TOKEN` in your application's [environment variables](../conductor/reference/dbos-cloud/secrets). ::: Then, configure your DBOS application to use Logfire and enable OpenTelemetry traces. You don't need to set `otlp_traces_endpoints` or `otlp_logs_endpoints`, because the Logfire SDK automatically configures exporters. ```python logfire.configure(service_name='my-app') # (Optional) If you're using Pydantic AI, instrument the agents logfire.instrument_pydantic_ai() dbos_config: DBOSConfig = { "name": "my-app", "application_version": "0.1.0", # highlight-start "enable_otlp": True, # highlight-end } DBOS(config=dbos_config) ``` Now start your DBOS application. You should see your logs and traces appear on the Logfire dashboard! The following screenshot shows how in a durable Pydantic AI agent built with DBOS, you can correlate Pydantic AI agent traces (e.g., token usage) and DBOS workflow execution in the same view. ![DBOS Agent Logs and Traces on Logfire](./assets/logfire-agent-screenshot.png) For more details on using Logfire instrumentation with your application, see the [Pydantic Logfire documentation](https://logfire.pydantic.dev/docs/). --- ## DBOS MCP Server You can use the [DBOS Model Context Protocol (MCP) server](https://github.com/dbos-inc/dbos-mcp) to augment your LLM or agent with tools that can analyze and manage your DBOS workflows. This enables your LLM or agent to retrieve information on your applications' workflows and steps, for example to help you debug issues in development or production. To use the server, your application should be connected to [Conductor](../conductor/overview.md). You may want to use the MCP server alongside a DBOS prompt or skills ([Python](../python/prompting.md), [TypeScript](../typescript/prompting.md), [Go](../golang/prompting.md), [Java](../java/prompting.md)) so your model has the most up-to-date information on DBOS. ### Setup #### Install `uv` Before using this MCP server, you must install `uv`. For installation instructions, see the [`uv` installation docs](https://docs.astral.sh/uv/getting-started/installation/). #### Setup with Claude Code To use this MCP server with Claude Code, first install it: ```bash claude mcp add dbos-conductor -- uvx dbos-mcp ``` Then start Claude Code and ask it questions about your DBOS apps! Claude will prompt you to log in by clicking the URL it offers and authenticating in the browser. Credentials are stored in `~/.dbos-mcp/credentials`. ### Tools The DBOS MCP server provides the following tools: ##### Application Introspection - `list_applications` - List all applications - `list_executors` - List connected executors for an application - `list_application_versions` - List all versions of an application - `set_latest_application_version` - Set an application's latest version ##### Workflow Introspection - `list_workflows` - List/filter workflows - `get_workflow` - Get workflow details - `list_steps` - Get execution steps for a workflow - `get_workflow_events` - Get the events a workflow published - `get_workflow_notifications` - Get the notifications a workflow received - `get_workflow_aggregates` - Aggregate workflow counts, grouped by status, name, queue, executor, version, or application ##### Workflow Management - `cancel_workflow` - Cancel a running workflow - `resume_workflow` - Resume a pending or failed workflow - `fork_workflow` - Fork a workflow from a specific step - `delete_workflow` - Delete a workflow and its history - `bulk_cancel_workflows` - Cancel multiple workflows at once - `bulk_resume_workflows` - Resume multiple workflows at once - `bulk_delete_workflows` - Delete multiple workflows at once - `fork_from_failure` - Fork multiple failed workflows from the point at which they failed ##### Schedule Management - `list_schedules` - List an application's schedules - `get_schedule` - Get details of a specific schedule - `pause_schedule` - Pause a schedule, stopping it from triggering new workflows - `resume_schedule` - Resume a paused schedule - `trigger_schedule` - Manually trigger a schedule to run its workflow immediately ##### Authentication - `login` - Start login flow (returns URL to login page) - `login_complete` - Complete login after authenticating --- ## Neon ## Use DBOS With Neon Here's how to connect your DBOS application running on your computer or cloud environment to your Neon database. #### 1. Set up a Local Application If you haven't already, follow the [quickstart](../quickstart.md) to set up a DBOS application locally. The rest of this guide will assume you have a local application. #### 2. Connect to Your Neon Database Next, open your [Neon dashboard](https://console.neon.tech), select a project, and click "Connect" to retrieve connection information for your Neon database. You should see a screen that looks like this: This page shows the connection string for your database. There are a few settings you may wish to alter before retrieving this connection string: 1. Make sure you are viewing your "Connection string" from the dropdown. 2. We recommend disabling connection pooling when connecting from DBOS. 3. By default, you will use the `neondb` database. If you want to use a different database, create it from the dropdown. When you are ready, copy the connection string (including the password) from the dashboard and set the `DBOS_SYSTEM_DATABASE_URL` environment variable to it: ``` export DBOS_SYSTEM_DATABASE_URL="" ``` #### 3. Launch Your Application Now, launch your DBOS application. It should successfully connect to your Neon database, printing your masked Neon database URL on startup. After connecting your DBOS application to Neon, you can use the Neon console to view your DBOS system tables. Open the "Tables" tab in the Neon console. For your database, select the database to which you connected your DBOS application (default `neondb`). For your schema, select "dbos". You can now see DBOS durably checkpoint your workflows to your Neon database: --- ## Nest.js ## DBOS + Nest.js This guide shows you how to add DBOS durable workflows to your existing [Nest.js](https://nestjs.com/) application to make it resilient to any failure. :::info This example was bootstrapped with `nest new`. You can see its full code on [GitHub](https://github.com/dbos-inc/dbos-demo-apps/tree/main/typescript/dbos-nestjs-starter). ::: ### Installation and Requirements Install the [open-source DBOS TypeScript library](https://github.com/dbos-inc/dbos-transact-ts) with: ```shell npm install @dbos-inc/dbos-sdk ``` ### Bootstrapping DBOS First, modify your bootstrap function to configure and launch DBOS: ```typescript title="src/main.ts" import { NestFactory } from '@nestjs/core'; import { AppModule } from './app.module'; // highlight-next-line import { DBOS } from "@dbos-inc/dbos-sdk"; async function bootstrap() { const app = await NestFactory.create(AppModule); // highlight-next-line DBOS.setConfig({ // highlight-next-line name: 'dbos-nestjs-starter', // highlight-next-line applicationVersion: '0.1.0', // highlight-next-line systemDatabaseUrl: process.env.DBOS_SYSTEM_DATABASE_URL, // highlight-next-line }); // highlight-next-line await DBOS.launch(); await app.listen(process.env.PORT ?? 3000); } void bootstrap(); ``` ### Add Workflows to Services Next, you can integrate DBOS workflows into your Nest.js services by annotating or registering service methods. To register a service instance method as a workflow, its class must extend [`ConfiguredInstance`](../typescript/tutorials/instantiated-objects.md). By extending `ConfiguredInstance`, you add your workflow methods to a DBOS internal registry so that if DBOS needs to recover your workflows, it can do so using the appropriate instance of your service. Here is an example of a Nest.js service implementing a simple two-step workflow: ```typescript title="src/app.service.ts" // highlight-next-line import { ConfiguredInstance, DBOS } from '@dbos-inc/dbos-sdk'; import { Injectable } from '@nestjs/common'; @Injectable() // highlight-next-line export class AppService extends ConfiguredInstance { constructor(name: string) { super(name); } async stepOne() { console.log('Step one completed!'); return Promise.resolve(); } async stepTwo() { console.log('Step two completed!'); return Promise.resolve(); } // highlight-next-line @DBOS.workflow() async workflow() { await DBOS.runStep(() => this.stepOne(), { name: 'stepOne' }); await DBOS.runStep(() => this.stepTwo(), { name: 'stepTwo' }); return 'Hello World!'; } } ``` ### Configure Service Instantiation You can instantiate classes containing DBOS workflows during dependency injection just like any other Nest.js class. If you create multiple instances of a class containing DBOS workflows, you should give them distinct names (`dbos-service-instance` in this case). ```typescript title="src/app.module.ts" import { Module, Provider } from '@nestjs/common'; import { AppController } from './app.controller'; import { AppService } from './app.service'; export const appProvider: Provider = { provide: AppService, useFactory: () => { const service = new AppService('dbos-service-instance'); return service; }, }; @Module({ imports: [], controllers: [AppController], providers: [appProvider], }) export class AppModule {} ``` --- ## Next.js ## DBOS + Next.js :::info To learn how to run DBOS with Next.js on Vercel, see the [Vercel integration guide](./vercel.md). ::: This guide shows you how to add DBOS durable workflows to a [Next.js](https://nextjs.org/) application. Running DBOS directly inside the Next.js process is difficult because Next.js may evaluate modules containing the DBOS instance or DBOS workflows and queues multiple times in different contexts. Instead, we recommend a two-process architecture: 1. A **Next.js app** that enqueues workflows using the [DBOS client](../typescript/reference/client.md) from server actions. 2. A **DBOS worker** that dequeues and executes workflows. Both processes connect to the same Postgres database. You can see a full working example on [GitHub](https://github.com/dbos-inc/dbos-demo-apps/tree/main/typescript/dbos-nextjs-starter). ### Installation Install DBOS and its OpenTelemetry integration (required by Turbopack): ```shell npm install @dbos-inc/dbos-sdk @dbos-inc/otel ``` ### 1. Create a Worker Create a standalone script that registers your workflows and launches DBOS. When started, it will automatically dequeue and execute workflows. ```ts title="worker/index.ts" import { DBOS } from "@dbos-inc/dbos-sdk"; async function greetingWorkflowFn(name: string) { return await DBOS.runStep( async () => `Hello, ${name}!`, { name: "generateGreeting" } ); } DBOS.registerWorkflow(greetingWorkflowFn); async function main() { DBOS.setConfig({ name: "my-nextjs-app", applicationVersion: "0.1.0", systemDatabaseUrl: process.env.DBOS_SYSTEM_DATABASE_URL, }); await DBOS.launch(); await DBOS.registerQueue("task_queue"); console.log("Worker listening for workflows..."); } main().catch(console.error); ``` ### 2. Enqueue Workflows from Server Actions In your Next.js app, use the [DBOS client](../typescript/reference/client.md) to enqueue workflows and check their status. The client connects directly to the DBOS system database without needing a full DBOS launch. ```ts title="app/actions.ts" "use server"; import { DBOSClient } from "@dbos-inc/dbos-sdk"; async function getClient() { return DBOSClient.create({ systemDatabaseUrl: process.env.DBOS_SYSTEM_DATABASE_URL!, }); } export async function launchWorkflow(name: string) { const client = await getClient(); try { const handle = await client.enqueue( { workflowName: "greetingWorkflowFn", queueName: "task_queue", }, name ); return { workflowID: handle.workflowID }; } finally { await client.destroy(); } } export async function getWorkflowStatus(workflowID: string) { const client = await getClient(); try { const handle = client.retrieveWorkflow(workflowID); const status = await handle.getStatus(); let result: string | null = null; if (status?.status === "SUCCESS") { result = await handle.getResult(); } return { status: status?.status ?? "UNKNOWN", result }; } finally { await client.destroy(); } } ``` Then call these server actions from your page: ```tsx title="app/page.tsx" "use client"; import { useState } from "react"; import { launchWorkflow, getWorkflowStatus } from "./actions"; export default function Home() { const [result, setResult] = useState(null); async function handleClick() { const { workflowID } = await launchWorkflow("World"); // Poll until the workflow completes const interval = setInterval(async () => { const data = await getWorkflowStatus(workflowID); if (data.status === "SUCCESS") { setResult(data.result); clearInterval(interval); } }, 500); } return (
{result &&

{result}

}
); } ``` --- ## OpenAI Agents SDK You can use DBOS to add durable execution to an agent built with the [OpenAI Agents SDK](https://platform.openai.com/docs/guides/agents-sdk). With durable execution, you can build reliable agents that preserve progress across transient API failures, application errors, and restarts, while also handling long-running, asynchronous, and human-in-the-loop workflows with production-grade reliability. ### Installation To get started, install DBOS and the [durable OpenAI agents integration](https://github.com/dbos-inc/dbos-openai-agents). ``` pip install dbos dbos-openai-agents ``` ### Building Reliable Agents To wrap your OpenAI agents for durable execution, [install and configure DBOS](../python/integrating-dbos.md) then follow these three guidelines: 1. Use `DBOSRunner.run` and `DBOSRunner.run_sync` as drop-in replacements for [`Runner.run`](https://openai.github.io/openai-agents-python/ref/run/#agents.run.Runner.run) and [`Runner.run_sync`](https://openai.github.io/openai-agents-python/ref/run/#agents.run.Runner.run_sync). 2. Annotate the function calling `DBOSRunner.run` or `DBOSRunner.run_sync` with `@DBOS.workflow` to run your agent as a durably executed workflow. 3. Annotate your agent's tool call functions and guardrail functions with `@DBOS.step` or `@DBOS.workflow()` to mark them as steps or sub-workflows of your durably executed agentic workflow. Here is a simple but complete example of wrapping an agent for durable execution. With just 10 lines of code (highlighted below), you can add DBOS into an existing OpenAI Agents application. ```python title="dbos_agent.py" import asyncio from agents import Agent, function_tool # highlight-start from dbos import DBOS, DBOSConfig from dbos_openai_agents import DBOSRunner #highlight-end # Decorate tool calls and guardrails with @DBOS.step() for durable execution @function_tool # highlight-next-line @DBOS.step() async def get_weather(city: str) -> str: """Get the weather for a city.""" return f"Sunny in {city}" agent = Agent(name="weather", tools=[get_weather]) # Use DBOSRunner to call your agent from a workflow # highlight-start @DBOS.workflow() async def run_agent(user_input: str) -> str: result = await DBOSRunner.run(agent, user_input) return str(result.final_output) # highlight-end async def main(): # highlight-start config: DBOSConfig = { "name": "my-agent", "application_version": "0.1.0", "system_database_url": 'sqlite:///my_agent.sqlite', } DBOS(config=config) DBOS.launch() # highlight-end output = await run_agent("How is the weather in San Francisco") print(output) if __name__ == "__main__": asyncio.run(main()) ``` **Notes:** 1. Workflows and agents must be defined before [`DBOS.launch()`](../python/reference/dbos-class.md#launch). 2. This example uses SQLite for ease of getting started. Postgres is recommended for production. ### Learn More For more details on building agents, see the [OpenAI Agents SDK documentation](https://platform.openai.com/docs/guides/agents-sdk). For information about durable execution and workflow design, see the [DBOS programming guide](../python/programming-guide). Together, these resources cover everything from getting started with simple agents to designing production-ready, fault-tolerant applications. --- ## Parseable ## Use DBOS With Parseable [Parseable](https://www.parseable.com/) can ingest OpenTelemetry logs and traces from DBOS application processes as well as [Conductor Metrics](../conductor/metrics.md). ![DBOS Logs and Traces on Parseable](./assets/parseable-screenshot.png) For details, see the [Parseable DBOS Guide](https://www.parseable.com/docs/ingest-data/ai-agents/dbos). --- ## Pydantic AI ## Use DBOS With Pydantic AI [Pydantic AI](https://ai.pydantic.dev/) is a Python agent framework for building production-grade applications and workflows powered by Generative AI. :::info Also check out the integration guide in [the Pydantic docs](https://ai.pydantic.dev/durable_execution/dbos)! ::: This guide shows you how to build a durable AI agent with DBOS and Pydantic AI. By combining the two, you can build reliable agents that preserve progress across transient API failures, application errors, and restarts, while also handling long-running, asynchronous, and human-in-the-loop workflows with production-grade reliability. ### Overview Pydantic AI has native support for DBOS agents. The diagram below shows the architecture of an agentic application built with DBOS. DBOS runs fully in-process as a library: functions remain standard Python functions but are checkpointed to a database for durability. ```text Clients (HTTP, RPC, Kafka, etc.) | v +------------------------------------------------------+ | Application Servers | | | | +----------------------------------------------+ | | | Pydantic AI + DBOS Libraries | | | | | | | | [ Workflows (Agent Run Loop) ] | | | | [ Steps (Tool, MCP, Model) ] | | | | [ Queues ] [ Cron Jobs ] [ Messaging ] | | | +----------------------------------------------+ | | | +------------------------------------------------------+ | v +------------------------------------------------------+ | Database | | (Stores workflow and step state, schedules tasks) | +------------------------------------------------------+ ``` Any Pydantic AI agent can be wrapped in a [`DBOSAgent`](https://ai.pydantic.dev/api/durable_exec/#pydantic_ai.durable_exec.dbos.DBOSAgent) to enable durable execution. `DBOSAgent` automatically: * Wraps `Agent.run` and `Agent.run_sync` (the agent's main loop) as DBOS workflows. * Wraps [model requests](https://ai.pydantic.dev/models/overview) and [MCP communication](https://ai.pydantic.dev/mcp/client) as DBOS steps. Custom tool functions and event stream handlers are **not automatically wrapped** by DBOS. You can decide how to integrate them: * Decorate with `@DBOS.step` if the function involves non-determinism or I/O. * Skip the decorator if durability isn't needed (e.g., debug logging), to avoid the extra DB checkpoint write. * If the function needs to enqueue tasks or invoke other DBOS workflows, skip the decorator so it runs inside the agent's main workflow (not as a step). The original agent, model, and MCP server can still be used normally outside the DBOS agent. ### Installation and Requirements Simply install Pydantic AI with the DBOS optional dependency: **pip** ```shell pip install pydantic-ai[dbos] ``` **uv** ```shell uv add pydantic-ai[dbos] ``` Or if you're using the slim package, you can install it with the DBOS optional dependency: **pip** ```shell pip install pydantic-ai-slim[dbos] ``` **uv** ```shell uv add pydantic-ai-slim[dbos] ``` ### Using DBOS Agent Here is a simple but complete example of wrapping an agent for durable execution. With fewer than 10 additional lines (highlighted below), you can add DBOS into an existing Pydantic AI application. ```python title="dbos_agent.py" import asyncio # highlight-next-line from dbos import DBOS, DBOSConfig from pydantic_ai import Agent # highlight-next-line from pydantic_ai.durable_exec.dbos import DBOSAgent # highlight-start dbos_config: DBOSConfig = { 'name': 'pydantic_dbos_agent', 'application_version': '0.1.0', 'system_database_url': 'sqlite:///dbostest.sqlite', } DBOS(config=dbos_config) #highlight-end agent = Agent( 'gpt-5', instructions="You're an expert in geography.", name='geography', ) # highlight-next-line dbos_agent = DBOSAgent(agent) async def main(): # highlight-next-line DBOS.launch() result = await dbos_agent.run('What is the capital of Mexico?') print(result.output) #> Mexico City (Ciudad de México, CDMX) if __name__ == "__main__": asyncio.run(main()) ``` **Notes:** 1. Workflows and `DBOSAgent` must be defined before [`DBOS.launch()`](../python/reference/dbos-class.md#launch) so that recovery can correctly find all workflows. 2. [`DBOSAgent.run()`](https://ai.pydantic.dev/api/durable_exec/#pydantic_ai.durable_exec.dbos.DBOSAgent.run) works like [`Agent.run()`](https://ai.pydantic.dev/api/agent/#pydantic_ai.agent.AbstractAgent.run), but runs as a DBOS workflow and executes model requests, decorated tool calls, and MCP communication as DBOS steps. 3. This example uses SQLite for simplicity. Postgres is recommended for production. 4. Each agent must have a unique `name`, which DBOS uses to identify its workflows. For more details on building agents, see the [Pydantic AI documentation](https://ai.pydantic.dev/durable_execution/dbos). For information about durable execution and workflow design, see the [DBOS programming guide](../python/programming-guide). Together, these resources cover everything from getting started with simple agents to designing production-ready, fault-tolerant applications. --- ## Supabase Edge Functions ## Use DBOS On Supabase Edge Functions You can use DBOS to add durable workflows, background jobs, or AI agents to your Supabase project. We recommend the following architecture: 1. Enqueue workflows for execution from an [Edge Function](https://supabase.com/docs/guides/functions) using the [DBOS client](../typescript/reference/client.md). 2. Create a DBOS worker in a second Edge Function to serverlessly dequeue and execute your workflows. 3. Configure a [pg_cron](https://supabase.com/docs/guides/cron) job that starts your worker whenever there is new work for it to execute. Both functions connect to your project's Postgres database, so DBOS checkpoints your workflows next to the rest of your data. Because an Edge Function is terminated when it exhausts its CPU or wall-clock budget, connect your worker to [Conductor](../conductor/overview.md), which detects the disconnect and recovers interrupted workflows onto the next worker that starts. :::info You can check out a working example of this integration [on GitHub](https://github.com/dbos-inc/dbos-demo-apps/tree/main/typescript/supabase-edge-functions). ::: ### 1. Enqueue Workflows From an Edge Function First, enqueue workflows using the [DBOS client](../typescript/reference/client.md). ```ts title="supabase/functions/enqueue/index.ts" import { DBOSClient } from "npm:@dbos-inc/dbos-sdk@4.27.6"; // Cached per isolate; Supabase reuses isolates across invocations. const client = await DBOSClient.create({ systemDatabaseUrl: Deno.env.get("DBOS_SYSTEM_DATABASE_URL")!, applicationName: "supabase-edge-functions", }); Deno.serve(async (req) => { const task = await req.json(); const handle = await client.enqueue( { queueName: "supabase-queue", workflowName: "processTask" }, task, ); // Optionally, start the worker immediately instead of waiting for it to poll. EdgeRuntime.waitUntil( fetch(`${Deno.env.get("SUPABASE_URL")}/functions/v1/worker`, { method: "POST", headers: { Authorization: `Bearer ${Deno.env.get("SUPABASE_SERVICE_ROLE_KEY")}` }, }), ); return Response.json({ workflowID: handle.workflowID }); }); ``` You can also use the client to list past workflows or retrieve their results. The `applicationName` you pass to the client must match the name your worker is configured with, as queues are scoped by application name. ### 2. Create a Worker Edge Function Next, create a DBOS worker in a second Edge Function. In it, define and register your workflows and steps: ```ts title="supabase/functions/worker/index.ts" import { DBOS } from "npm:@dbos-inc/dbos-sdk@4.27.6"; async function validate(task: Task) { /* ... */ } async function transform(task: Task) { /* ... */ } async function processTaskFn(task: Task) { const validated = await DBOS.runStep(() => validate(task), { name: "validate" }); return await DBOS.runStep(() => transform(validated), { name: "transform" }); } DBOS.registerWorkflow(processTaskFn, { name: "processTask" }); ``` Then configure and launch DBOS, register your queue, and process its workflows as a [background task](https://supabase.com/docs/guides/functions/background-tasks), so the function answers its caller immediately instead of holding the request open until the workflows finish. ```ts title="supabase/functions/worker/index.ts" async function processWorkflows() { DBOS.setConfig({ name: "supabase-edge-functions", systemDatabaseUrl: Deno.env.get("DBOS_SYSTEM_DATABASE_URL")!, applicationVersion: "v1", }); await DBOS.launch({ conductorKey: Deno.env.get("DBOS_CONDUCTOR_KEY") }); try { await DBOS.registerQueue("supabase-queue"); // Run until the queue is empty, then exit. If Supabase terminates the function // first, Conductor recovers the in-flight workflows onto the next worker. while (true) { const workflows = await DBOS.listWorkflows({ status: ["ENQUEUED", "PENDING"], applicationName: "supabase-edge-functions", limit: 1, }); if (workflows.length === 0) break; await new Promise((resolve) => setTimeout(resolve, 1000)); } } finally { await DBOS.shutdown(); } } let processing = false; Deno.serve(() => { // Launch DBOS at most once per Edge Functions isolate if (processing) { return Response.json({ accepted: false, reason: "already-processing" }); } processing = true; // waitUntil keeps the isolate alive until processing is complete EdgeRuntime.waitUntil( processWorkflows() .catch((e) => console.error(e)) .finally(() => { processing = false; }), ); return Response.json({ accepted: true }); }); ``` ### 3. Schedule the Worker With pg_cron Finally, configure a [pg_cron](https://supabase.com/docs/guides/cron) job to start your worker whenever there is new work for it. This function periodically checks your database for active (`ENQUEUED` or `PENDING`) workflows. If there are any, it launches a worker to process them. ```sql title="supabase/sql/02_cron.sql" select cron.schedule( 'dbos-worker-tick', '* * * * *', $$ -- Start a worker by POSTing to it with your service role key... select net.http_post( url := (select decrypted_secret from vault.decrypted_secrets where name = 'project_url') || '/functions/v1/worker', headers := jsonb_build_object( 'Content-Type', 'application/json', 'Authorization', 'Bearer ' || (select decrypted_secret from vault.decrypted_secrets where name = 'service_role_key') ), body := jsonb_build_object('source', 'cron'), timeout_milliseconds := 5000 ) -- ...but only when there are workflows waiting to be executed. where exists ( select 1 from dbos.workflow_status where application_name = 'supabase-edge-functions' and status in ('ENQUEUED', 'PENDING') ); $$ ); ``` You can check out a working example of this integration (with deployment instructions) [on GitHub](https://github.com/dbos-inc/dbos-demo-apps/tree/main/typescript/supabase-edge-functions). --- ## Supabase ## Use DBOS With Supabase :::info To learn more about how DBOS and Supabase are working together, check out [this blog post](https://supabase.com/blog/durable-workflows-in-postgres-dbos)! ::: Here's how to connect your DBOS application running on your computer or cloud environment to your Supabase. #### 1. Set up a Local Application If you haven't already, follow the [quickstart](../quickstart.md) to set up a DBOS application locally. The rest of this guide will assume you have a local application. #### 2. Connect to Your Supabase Database Next, open your Supabase dashboard at [`supabase.com/dashboard`](https://supabase.com/dashboard), select a project, and click "Connect" to retrieve connection information for your Supabase database. You should see a screen that looks like this, showing the connection string for your database: :::tip Make sure your connection method is set to "Direct connection" (if your database supports it) or to "Session pooler". ::: When you are ready, copy the connection string (filling in your Supabase password) from the dashboard and set the `DBOS_SYSTEM_DATABASE_URL` environment variable to it: ``` export DBOS_SYSTEM_DATABASE_URL="" ``` #### 3. Launch Your Application Now, launch your DBOS application. It should successfully connect to your Supabase database, printing your masked Supabase database URL on startup. After connecting your DBOS application to Supabase, you can use the Supabase console to view your DBOS system tables. Open the "Table Editor" tab in the Supabase console. For your schema, select "dbos". You can now see DBOS durably checkpoint your workflows to your Supabase database: --- ## Tiger Data ## Use DBOS With Tiger Data Here's how to connect your DBOS application running on your computer or cloud environment to a Postgres or TimescaleDB database running in Tiger Cloud (the makers of TimescaleDB). #### 1. Set up a Local Application If you haven't already, follow the [quickstart](../quickstart.md) to set up a DBOS application locally. The rest of this guide will assume you have a local application. #### 2. Connect to Your Database Next, open your [dashboard](https://console.cloud.timescale.com/dashboard/services) and select your service. You should see a screen that looks like this: Copy the connection string that appears on the right and set the `DBOS_SYSTEM_DATABASE_URL` environment variable to it: ``` export DBOS_SYSTEM_DATABASE_URL="" ``` For security, the TimescaleDB connection string does not include your password, so also set the `PGPASSWORD` environment variable to your database password: ``` export PGPASSWORD="" ``` #### 3. Launch Your Application Now, launch your DBOS application. It should successfully connect to your database, printing your masked Tiger Cloud database URL on startup. After connecting your DBOS application, you can use the console to view your DBOS system tables. Open the "SQL editor" tab in the Tiger Cloud console. Run the following query: ```sql SELECT * FROM dbos.workflow_status; ``` You should see the durable checkpoints DBOS makes for your workflows: --- ## Vercel AI SDK ## Use DBOS With the Vercel AI SDK You can use DBOS to add [durable execution](../typescript/tutorials/workflow-tutorial.md) to agents built with the [Vercel AI SDK](https://ai-sdk.dev/) through the [`@dbos-inc/vercel-ai`](https://www.npmjs.com/package/@dbos-inc/vercel-ai) package. This package makes AI SDK agents durable, backed by your Postgres database. All you have to do is wrap your model with `durableCalls` and your tools with `durableTools` and run your agents inside a DBOS workflow. Then, this integration automatically checkpoints every action your agents take in Postgres. If your process is interrupted, DBOS replays your agent from its checkpoints so it resumes from where it left off. ```ts import { DBOS } from '@dbos-inc/dbos-sdk'; import { generateText, wrapLanguageModel } from 'ai'; import { openai } from '@ai-sdk/openai'; import { durableCalls } from '@dbos-inc/vercel-ai'; const model = wrapLanguageModel({ model: openai('gpt-5'), middleware: durableCalls({ retriesAllowed: true, maxAttempts: 5 }), }); const researchAgent = DBOS.registerWorkflow( async (question: string) => { const { text } = await generateText({ model, prompt: question, system: 'You are a helpful research assistant.', }); return text; }, { name: 'researchAgent' }, ); DBOS.setConfig({ name: 'my-agent', systemDatabaseUrl: process.env.DBOS_SYSTEM_DATABASE_URL }); await DBOS.launch(); console.log(await researchAgent('Why did the agent cross the road?')); ``` ### Installation ```sh npm install @dbos-inc/vercel-ai @dbos-inc/dbos-sdk ai ``` Requires DBOS v4.27+ or v5, AI SDK v7+, and a Postgres database for DBOS. ### Durable Model Calls To durably checkpoint each call you make to a model, wrap your model in `durableCalls`. Then, call your model or agent from a workflow: ```ts import { DBOS } from '@dbos-inc/dbos-sdk'; import { ToolLoopAgent, wrapLanguageModel } from 'ai'; import { openai } from '@ai-sdk/openai'; import { durableCalls } from '@dbos-inc/vercel-ai'; const model = wrapLanguageModel({ model: openai('gpt-5'), middleware: durableCalls() }); const agent = new ToolLoopAgent({ model, instructions: 'You are a helpful research assistant.', tools }); const researchAgent = DBOS.registerWorkflow( async (question: string) => { const result = await agent.stream({ prompt: question }); for await (const delta of result.textStream) process.stdout.write(delta); return await result.text; }, { name: 'researchAgent' }, ); ``` You can parameterize `durableCalls` to configure model call retries and timeouts: ```ts durableCalls({ name?: string; // step name (default: "..") retriesAllowed?: boolean; // retry failed model calls (default: true) maxAttempts?: number; // total attempts when retries are allowed (default: 3) intervalSeconds?: number; // delay before first retry (default: 1) backoffRate?: number; // exponential backoff multiplier (default: 2) shouldRetry?: (error: unknown) => boolean; // default: skip provider-declared non-retryable errors and aborts timeoutMS?: number; // per-attempt timeout durableStream?: string; // stream each call's output to this durable stream include?: { requestBody?: boolean; responseBody?: boolean }; // checkpoint raw provider bodies; match generateText's `include` (default: false) }); ``` ### Durable Tools To durably checkpoint your agents' tool calls, wrap them in `durableTools`: ```ts import { tool, stepCountIs } from 'ai'; import { durableTools } from '@dbos-inc/vercel-ai'; import { z } from 'zod'; const tools = durableTools({ getWeather: tool({ description: 'Get the weather for a city', inputSchema: z.object({ city: z.string() }), execute: ({ city }) => fetchWeather(city), }), }); const agent = DBOS.registerWorkflow(async (question: string) => { const result = await generateText({ model, prompt: question, tools, stopWhen: stepCountIs(10) }); return result.text; }, { name: 'weatherAgent' }); ``` You can pass step configuration (such as timeouts or retries) to `durableTools`. You can set defaults for all tools or configure tools individually. Retries are off by default. ```ts const tools = durableTools(myTools, { timeoutMS: 30_000, tools: { getWeather: { retriesAllowed: true, maxAttempts: 3 }, }, }); ``` When using durable tools, to ensure the ordering of parallel tool calls is consistent during recovery, do not await I/O in callbacks that run before a tool executes, such as `onToolExecutionStart`. ### Durable Streams You can [**durably stream**](../typescript/tutorials/workflow-communication.md#workflow-streaming) agent or model output so it can be read by an external client or UI. To do this, configure `durableCalls` or `durableTools`/`durableMCPTools` with a durable stream name: ```ts import { createUIMessageStreamResponse, streamText } from 'ai'; import { durableCalls, durableTools, readDurableStream } from '@dbos-inc/vercel-ai'; const model = wrapLanguageModel({ model: openai('gpt-5'), middleware: durableCalls({ durableStream: 'ui' }) }); const tools = durableTools(myTools, { durableStream: 'ui' }); const chatTurn = DBOS.registerWorkflow(async (messages: ModelMessage[]) => { const result = streamText({ model, messages, tools, stopWhen: stepCountIs(10) }); return await result.text; }, { name: 'chatTurn' }); export async function POST(req: Request) { const { messages, messageId } = (await req.json()) as { messages: ModelMessage[]; messageId: string }; const handle = await DBOS.startWorkflow(chatTurn)(messages); return createUIMessageStreamResponse({ stream: readDurableStream({ workflowID: handle.workflowID, key: 'ui', messageId }), }); } ``` You can read from a durable stream using `readDurableStream`, for example to stream it to a UI. It emits a stream of AI SDK `UIMessageChunk`. You can also pass a [`DBOSClient`](../typescript/reference/client.md) into `readDurableStream` to read it from a different process. You can write your own data to a stream with `writeDurableStream(key, chunks)`. Tools can also write chunks with `toolWriter()`, which returns a `UIMessageStreamWriter` bound to the current tool call. Chunks from a tool call are written when the tool call succeeds (`transient: true` data parts are written live instead). To also include them in the response message your workflow builds, pass your `createUIMessageStream` writer to `durableTools`: ```ts import { consumeStream, createUIMessageStream, streamText } from 'ai'; import { durableTools, toolWriter } from '@dbos-inc/vercel-ai'; const search = tool({ inputSchema: z.object({ q: z.string() }), execute: async ({ q }) => { const writer = toolWriter(); writer.write({ type: 'data-progress', data: { pct: 50 }, transient: true }); writer.write({ type: 'source-url', sourceId: 's1', url: 'https://example.com' }); return runSearch(q); }, }); const chatTurn = DBOS.registerWorkflow(async (messages: ModelMessage[]) => { const stream = createUIMessageStream({ execute: ({ writer }) => { const tools = durableTools({ search }, { durableStream: 'ui', writer }); writer.merge(streamText({ model, messages, tools, stopWhen: stepCountIs(10) }).toUIMessageStream()); }, onFinish: ({ responseMessage }) => saveMessage(responseMessage), }); await consumeStream({ stream }); }, { name: 'chatTurn' }); ``` Your streams are closed when your workflow finishes; you can also close a stream early using `closeDurableStream`. If a workflow is interrupted during a model call, when the workflow recovers, it restarts the model call and streams its output again. Readers that connect afterwards see the model's output once; live readers, including readers resumed from an `offset`, receive a transient `data-dbos-superseded` chunk indicating the model call has been restarted. Likewise, if a tool call is re-executed and writes different chunks, readers that connect afterwards see only the re-execution's chunks; live readers receive a transient `data-dbos-tool-superseded` chunk naming the `toolCallId` whose earlier chunks to discard. #### Durable MCP Tools `durableMCPTools` wraps an [MCP](https://modelcontextprotocol.io/) client (for example, from [`@ai-sdk/mcp`](https://www.npmjs.com/package/@ai-sdk/mcp)) so both the tool listing and every tool call run as durable steps: ```ts import { createMCPClient } from '@ai-sdk/mcp'; import { durableMCPTools } from '@dbos-inc/vercel-ai'; const agent = DBOS.registerWorkflow(async (question: string) => { const mcpClient = await createMCPClient({ transport: { type: 'http', url: MCP_URL } }); try { const tools = await durableMCPTools(mcpClient); const result = await generateText({ model, prompt: question, tools, stopWhen: stepCountIs(10) }); return result.text; } finally { await mcpClient.close(); } }, { name: 'mcpAgent' }); ``` To use the client's explicit-schema mode (tool subsetting, typed inputs, output schemas), pass `toolOptions`; it is forwarded to `client.tools()` for both the listing and each tool call: ```ts const tools = await durableMCPTools(mcpClient, { toolOptions: { schemas: { 'get-weather': { inputSchema: z.object({ city: z.string() }) } } }, }); ``` ### Durable Subagents You can delegate complex tasks to **subagents**, which act as tools for their "parent" agent. To create a durable subagent, wrap your agent in `agentTool`, then pass it into `durableTools` just like any other tool: ```ts import { ToolLoopAgent } from 'ai'; import { agentTool, durableTools } from '@dbos-inc/vercel-ai'; const researcher = new ToolLoopAgent({ model, instructions: 'Research thoroughly.', tools: researchTools }); const research = agentTool({ name: 'research', // subagent name description: 'Research a question in depth', inputSchema: z.object({ question: z.string() }), agent: researcher, prompt: ({ question }) => question, // tool input → prompt (or ModelMessage[]) }); const tools = durableTools({ research, getWeather }, { durableStream: 'ui' }); const orchestrator = new ToolLoopAgent({ model, tools }); ``` Internally, subagents are implemented as child workflows of the parent agent workflow, so each call has its own checkpoints and parallel calls are safe. Call `agentTool` before `DBOS.launch()`, since it registers that workflow. By default, the tool returns the subagent's final text; you can configure this with the `output` parameter. ### Durable Embedding Models `durableEmbeddingCalls` enables durable calls to embedding models: ```ts import { embedMany, wrapEmbeddingModel } from 'ai'; import { durableEmbeddingCalls } from '@dbos-inc/vercel-ai'; const embeddingModel = wrapEmbeddingModel({ model: openai.textEmbeddingModel('text-embedding-3-small'), middleware: durableEmbeddingCalls({ retriesAllowed: true }), }); const embedChunks = DBOS.registerWorkflow(async (chunks: string[]) => { const { embeddings } = await embedMany({ model: embeddingModel, values: chunks }); return embeddings; }, { name: 'embedChunks' }); ``` ### Durable Image Models `durableImageCalls` makes image generation durable: ```ts import { generateImage, wrapImageModel } from 'ai'; import { durableImageCalls } from '@dbos-inc/vercel-ai'; const imageModel = wrapImageModel({ model: openai.imageModel('gpt-image-1'), middleware: durableImageCalls() }); const drawImage = DBOS.registerWorkflow(async (prompt: string) => { const { images } = await generateImage({ model: imageModel, prompt }); return images.map((image) => image.base64); }, { name: 'drawImage' }); ``` ### Learn More For more details on building agents with the Vercel AI SDK, see the [Vercel AI SDK documentation](https://ai-sdk.dev/). For information about durable execution and workflow design, see the [DBOS programming guide](../typescript/programming-guide.md). --- ## Vercel ## Use DBOS On Vercel You can use DBOS to add durable workflows, background jobs, or AI agents to your Next.js app hosted on Vercel. We recommend the following architecture: 1. In your Next.js app, enqueue workflows for execution using the [DBOS client](../typescript/reference/client.md). 2. Create a DBOS worker in a [Vercel Function](https://vercel.com/docs/functions) to serverlessly dequeue and execute your workflows. 3. Configure a [Vercel cron job](https://vercel.com/docs/cron-jobs) to periodically start your worker to poll for new workflows to execute. Your DBOS client and worker should both connect to a Postgres database—for example a Supabase or Neon database configured through [Vercel Postgres](https://vercel.com/docs/postgres). :::info You can check out a working example of this integration [on GitHub](https://github.com/dbos-inc/dbos-vercel-integration). ::: ### 1. Enqueue Workflows From Server Actions First, enqueue workflows from your Next.js app using the [DBOS client](../typescript/reference/client.md) in your [server actions](https://nextjs.org/docs/app/getting-started/updating-data). For example, here's a server action that enqueues a workflow to execute in the background: ```ts title="app/actions.ts" 'use server'; import { DBOSClient } from '@dbos-inc/dbos-sdk'; export async function enqueueWorkflow() { console.log('Enqueueing DBOS workflow'); const client = await DBOSClient.create({ systemDatabaseUrl: process.env.DBOS_SYSTEM_DATABASE_URL! }); await client.enqueue({ workflowName: 'exampleWorkflow', queueName: 'exampleQueue', }); await client.destroy(); } ``` You can also use the client to list past workflows or retrieve their results. ### 2. Create a Worker in a Vercel Function Next, create a DBOS worker in a [Vercel Function](https://vercel.com/docs/functions) to serverlessly dequeue and execute your workflows. In the Vercel Function, define and register your workflows, steps, and queues: ```ts title="app/api/dbos/route.ts" import { DBOS } from '@dbos-inc/dbos-sdk'; import { waitUntil } from '@vercel/functions'; // Define a workflow and steps async function stepOne() { // Sleep 3 seconds await new Promise((resolve) => setTimeout(resolve, 3000)); DBOS.logger.info('Step one completed!'); } async function stepTwo() { // Sleep 3 seconds await new Promise((resolve) => setTimeout(resolve, 3000)); DBOS.logger.info('Step two completed!'); } async function exampleFunction() { await DBOS.runStep(() => stepOne(), { name: 'stepOne' }); await DBOS.runStep(() => stepTwo(), { name: 'stepTwo' }); } // Register your workflow with DBOS DBOS.registerWorkflow(exampleFunction, { name: 'exampleWorkflow', }); ``` Then, configure and launch DBOS on worker startup, and register the queue. When started, the worker will poll your queues and execute your workflows, waiting until all enqueued workflows complete or a timeout is reached. If some workflows are still executing when the worker times out, don't worry—DBOS will automatically recover them when the worker next starts. ```ts title="app/api/dbos/route.ts" // Configure and launch DBOS. It will automatically dequeue and execute workflows. DBOS.setConfig({ name: 'dbos-vercel-integration', applicationVersion: '0.1.0', systemDatabaseUrl: process.env.DBOS_SYSTEM_DATABASE_URL, }); await DBOS.launch(); await DBOS.registerQueue('exampleQueue'); // After the worker Vercel function is launched, // it waits for either all enqueued workflows // to complete or for a timeout to be reached async function waitForQueuedWorkflowsToComplete(timeoutMs: number): Promise { const startTime = Date.now(); const intervalMs = 1000; while (true) { if (Date.now() - startTime >= timeoutMs) { throw new Error(`Timeout reached after ${timeoutMs}ms - queued workflows still exist`); } const queuedWorkflows = await DBOS.listQueuedWorkflows({queueName: 'exampleQueue'}); if (queuedWorkflows.length === 0) { console.log('All queued workflows completed'); return; } console.log(`${queuedWorkflows.length} workflows still queued, waiting...`); await new Promise((resolve) => setTimeout(resolve, intervalMs)); } } export async function GET(request: Request) { waitUntil(waitForQueuedWorkflowsToComplete(300000)); return new Response(`Starting DBOS worker! Request URL: ${request.url}`); } ``` ### 3. Schedule the Worker with Cron Finally, configure a [Vercel cron job](https://vercel.com/docs/cron-jobs) to periodically start your worker to poll for new workflows to execute. Vercel will automatically scale the worker function to handle your workflows. For example, you might configure your worker to poll for new workflows once a minute: ```json title="vercel.json" { "$schema": "https://openapi.vercel.sh/vercel.json", "crons": [ { "path": "/api/dbos", "schedule": "* * * * *" } ] } ``` --- ## Fault-Tolerant Checkout(Examples) :::info This example is also available in [TypeScript](../../typescript/examples/checkout-tutorial), [Go](../../golang/examples/widget-store), and [Python](../../python/examples/widget-store.md). ::: In this example, we use DBOS and Spring Boot to build an online storefront that's resilient to any failure. You can see the application live [here](https://demo-widget-store.cloud.dbos.dev/). Try playing with it and pressing the crash button as often as you want. Within a few seconds, the app will recover and resume as if nothing happened. All source code is [available on GitHub](https://github.com/dbos-inc/dbos-demo-apps/tree/main/java/widget-store). ![Widget store UI](./assets/widget_store_ui.png) ### Building the Checkout Workflow The core of this application is the checkout workflow, which orchestrates the entire purchase process. This workflow is triggered whenever a customer buys a widget and handles the complete order lifecycle: 1. Reserves inventory to ensure the item is available 2. Creates a new order in the system 3. Processes payment 4. Marks the order as paid and initiates fulfillment 5. Handles failures gracefully by releasing reserved inventory and canceling orders when necessary DBOS **durably executes** this workflow. It checkpoints each step in the database so that if the app fails or is interrupted during checkout, it will automatically recover from the last completed step. This means that customers never lose their order progress, no matter what breaks. You can try this yourself! On the [live application](https://demo-widget-store.cloud.dbos.dev/), start an order and press the crash button at any time. Within seconds, your app will recover to exactly the state it was in before the crash and continue as if nothing happened. ```java @Workflow public String checkoutWorkflow(String key) { try { dbos.runStep(() -> repo.subtractInventory(), "subtractInventory"); } catch (RuntimeException e) { logger.error("Failed to reserve inventory for workflow {}", DBOS.workflowId()); dbos.setEvent(PAYMENT_ID, null); return; } var orderId = dbos.runStep(() -> repo.createOrder(), "createOrder"); dbos.setEvent(PAYMENT_ID, DBOS.workflowId()); var payment_status = dbos.recv(PAYMENT_STATUS, Duration.ofSeconds(120)); if (payment_status.map(ps -> ps.equals("paid")).orElse(false)) { logger.info("Payment successful for order {}", orderId); dbos.runStep(() -> repo.markOrderPaid(orderId), "markOrderPaid"); dbos.startWorkflow(() -> self.dispatchOrderWorkflow(orderId)); } else { logger.info("Payment failed for order {}", orderId); dbos.runStep(() -> repo.errorOrder(orderId), "errorOrder"); dbos.runStep(() -> repo.undoSubtractInventory(), "undoSubtractInventory"); } dbos.setEvent(ORDER_ID, String.valueOf(orderId)); } ``` ### The Checkout and Payment Endpoints Now let's implement the HTTP endpoints that handle customer interactions with the checkout system. The checkout endpoint is triggered when a customer clicks the "Buy Now" button. It starts the checkout workflow in the background, then waits for the workflow to generate and send it a unique payment ID. It then returns the payment ID so the browser can redirect the user to the payments page. The endpoint accepts an [idempotency key](../tutorials/workflow-tutorial.md#workflow-ids-and-idempotency) so that even if the customer presses "buy now" multiple times, only one checkout workflow is started. ```java @PostMapping("/checkout/{key}") public ResponseEntity checkout(@PathVariable String key) { logger.info("Checkout requested with key: " + key); var options = new StartWorkflowOptions().withWorkflowId(key); dbos.startWorkflow(() -> service.checkoutWorkflow(), options); var paymentId = dbos.getEvent(key, PAYMENT_ID, Duration.ofSeconds(60)); if (paymentId.isEmpty()) { throw new RuntimeException("Item not available"); } else { return ResponseEntity.ok(paymentId.get()); } } ``` The payment endpoint handles the communication between the payment system and the checkout workflow. It uses the payment ID to signal the checkout workflow whether the payment succeeded or failed. It then retrieves the order ID from the checkout workflow so the browser can redirect the customer to the order status page. ```java @PostMapping("/payment_webhook/{key}/{status}") public ResponseEntity paymentWebhook(@PathVariable String key, @PathVariable String status) { logger.info("Payment webhook called with key: " + key + ", status: " + status); dbos.send(key, status, PAYMENT_STATUS); var orderId = dbos.getEvent(key, ORDER_ID, Duration.ofSeconds(60)); return ResponseEntity.ok(orderId.orElse(null)); } ``` ### Database Operations Now, let's take a look at how the checkout workflow's steps are implemented. Each step performs a database operation, like updating inventory or order status. These are implemented as [@Transactional](https://docs.spring.io/spring-framework/reference/data-access/transaction/declarative/annotations.html) Java functions that interact with the PostgreSQL database. ```java @Transactional public class WidgetStoreRepository { private static final int PRODUCT_ID = 1; // Product and Order Repository classes are JpaRepository implementations private final ProductRepository productRepository; private final OrderRepository orderRepository; public WidgetStoreRepository(ProductRepository productRepository, OrderRepository orderRepository) { this.productRepository = productRepository; this.orderRepository = orderRepository; } public ProductDto retrieveProduct() { return productRepository.findById(PRODUCT_ID).map(ProductDto::fromEntity).orElse(null); } @Transactional public void setInventory(int inventory) { productRepository.setInventory(PRODUCT_ID, inventory); } @Transactional public void subtractInventory() { int updated = productRepository.subtractInventory(PRODUCT_ID); if (updated == 0) { throw new RuntimeException("Insufficient Inventory"); } } @Transactional public void undoSubtractInventory() { productRepository.addInventory(PRODUCT_ID); } @Transactional public Integer createOrder() { Product product = productRepository.getReferenceById(PRODUCT_ID); Order order = new Order(); order.setOrderStatus(OrderStatus.PENDING); order.setProduct(product); order.setLastUpdateTime(LocalDateTime.now()); order.setProgressRemaining(10); return orderRepository.save(order).orderId(); } public OrderDto retrieveOrder(int orderId) { return orderRepository.findById(orderId).map(OrderDto::fromEntity).orElse(null); } public List retrieveOrders() { return orderRepository.findAllByOrderByOrderIdDesc().stream() .map(OrderDto::fromEntity) .toList(); } @Transactional public void markOrderPaid(int orderId) { orderRepository.updateOrderStatus(orderId, OrderStatus.PAID); } @Transactional public void errorOrder(int orderId) { orderRepository.updateOrderStatus(orderId, OrderStatus.CANCELLED); } @Transactional public void updateOrderProgress(int orderId) { Order order = orderRepository .findById(orderId) .orElseThrow(() -> new RuntimeException("Order not found: " + orderId)); order.setProgressRemaining(order.progressRemaining() - 1); order.setLastUpdateTime(LocalDateTime.now()); if (order.progressRemaining() == 0) { order.setOrderStatus(OrderStatus.DISPATCHED); } orderRepository.save(order); } } ``` ### Launching and Serving the App `transact-spring-boot-starter` automatically provides `DBOSConfig` and `DBOS` beans, automatically creates proxies for beans with `@Workflow` or `@Step` annotations and hooks into Spring's lifecycle management to automatically `dbos.launch()` and `dbos.shutdown()`. WidgetStoreConfig only exists to create the app database on startup and run Flyway migrations, making the demo app easier to run. ### Try it Yourself! First, clone and enter the [dbos-demo-apps](https://github.com/dbos-inc/dbos-demo-apps) repository: ```shell git clone https://github.com/dbos-inc/dbos-demo-apps.git cd java/widget-store ``` Then follow the instructions in the README to build and run the app! --- ## Add DBOS To Your App(Java) This guide shows you how to add the open-source [DBOS Transact](https://github.com/dbos-inc/dbos-transact-java) library to your existing application to **durably execute** it and make it resilient to any failure. #### 1. Install DBOS Add DBOS to your application by including it in your build configuration. **Gradle** ```groovy dependencies { implementation 'dev.dbos:transact:1.1.0' } ``` **Maven** ```xml dev.dbos transact 1.1.0 ``` #### 2. Add the DBOS Initializer Add these lines of code to your program's main method. They configure and launch DBOS when your program starts. ```java import dev.dbos.transact.DBOS; import dev.dbos.transact.config.DBOSConfig; public class MyApp { public static void main(String[] args) throws Exception { // Configure DBOS DBOSConfig dbosConfig = DBOSConfig.defaultsFromEnv("dbos-java-starter") .withAppVersion("0.1.0"); DBOS dbos = new DBOS(dbosConfig); // Register your workflows (see step 4) // Launch DBOS dbos.launch(); // Register your queues, which are stored in the system database } } ``` :::info DBOS uses a PostgreSQL database to durably store workflow and step state. You can connect to your database by setting these environment variables: - `DBOS_SYSTEM_JDBC_URL`: The JDBC URL for your PostgreSQL database (e.g., `jdbc:postgresql://localhost:5432/mydb`) - `PGUSER`: Your PostgreSQL username - `PGPASSWORD`: Your PostgreSQL password If you don't have a PostgreSQL database, you can start one locally with Docker: ```shell docker run -d \ --name dbos-postgres \ -e POSTGRES_PASSWORD=dbos \ -p 5432:5432 \ postgres:latest ``` ::: #### 3. Start Your Application Try starting your application. If everything is set up correctly, your app should run normally, but log `DBOS started` on startup. Congratulations! You've integrated DBOS into your application. #### 4. Start Building With DBOS At this point, you can add DBOS workflows and steps to your application. For example, you can annotate one of your methods as a [workflow](./tutorials/workflow-tutorial.md) and the methods it calls as [steps](./tutorials/step-tutorial.md). DBOS durably executes the workflow so if it is ever interrupted, upon restart it automatically resumes from the last completed step. ```java import dev.dbos.transact.DBOS; import dev.dbos.transact.workflow.Workflow; interface Example { public void workflow(); } class ExampleImpl implements Example { private final DBOS dbos; public ExampleImpl(DBOS dbos) { this.dbos = dbos; } private void stepOne() { System.out.println("Step one completed!"); } private void stepTwo() { System.out.println("Step two completed!"); } @Override @Workflow public void workflow() { dbos.runStep(() -> stepOne(), "stepOne"); dbos.runStep(() -> stepTwo(), "stepTwo"); } } ``` To use your workflows, create a proxy before launching DBOS: ```java // Create a DBOS instance and register the workflow proxy (before launching) DBOS dbos = new DBOS(config); Example proxy = dbos.registerProxy(Example.class, new ExampleImpl(dbos)); // Launch DBOS dbos.launch(); // Now you can call your workflows proxy.workflow(); ``` **Important:** You must create all workflow proxies before calling `dbos.launch()`. Workflow recovery begins after `dbos.launch()`, so all workflows must be registered before this point. Queues are the opposite: their configuration is stored in the system database, so register them with [`dbos.registerQueue`](./reference/queues.md#dbosregisterqueue) after `dbos.launch()`. You can add DBOS to your application incrementally—it won't interfere with code that's already there. It's totally okay for your application to have one DBOS workflow alongside thousands of lines of non-DBOS code. To learn more about programming with DBOS, check out [the programming guide](./programming-guide.md). --- ## Learn DBOS Java This guide shows you how to use DBOS to build Java apps that are **resilient to any failure**. :::tip To teach your AI coding assistant to build with DBOS, try out [skills](./prompting.md) and [MCP](../integrations/mcp.md). ::: ### 1. Setting Up Your Environment First, initialize a new project with Gradle (See [the installation instructions](https://docs.gradle.org/current/userguide/installation.html) if you do not have Gradle set up, or do not have Gradle 8 or later): ```shell mkdir myapp && cd myapp gradle init ``` When prompted, just accept the defaults for everything Gradle asks. Then, install DBOS (plus Logback for logging) by adding the following to your `app/build.gradle.kts` dependencies: ```kotlin dependencies { implementation("dev.dbos:transact:1.1.0") implementation("org.slf4j:slf4j-simple:2.0.17") // needed to see DBOS log messages implementation("io.javalin:javalin:7.0.1") // needed for creating HTTP endpoint later in the guide // remaining dependencies that were generated by gradle init } ``` DBOS also requires a PostgreSQL database. If you don't already have PostgreSQL, you can launch it in a Docker container with this command: ```shell docker run -d \ --name dbos-postgres \ -e POSTGRES_PASSWORD=dbos \ -p 5432:5432 \ postgres:latest ``` Then, set the following environment variables to your connection information (later, we'll pass them into DBOS). For example: ```shell export PGUSER=postgres export PGPASSWORD=dbos export DBOS_SYSTEM_JDBC_URL=jdbc:postgresql://localhost:5432/dbos_java_starter ``` ### 2. Workflows and Steps DBOS helps you add reliability to Java programs. The key feature of DBOS is **workflow methods** comprised of **steps**. DBOS automatically provides durability by checkpointing the state of your workflows and steps to its system database. If your program crashes or is interrupted, DBOS uses this saved state to recover each of your workflows from its last completed step. Thus, DBOS makes your application **resilient to any failure**. Let's create a simple DBOS program that runs a workflow of two steps. Add the following code to your `app/src/main/java/org/example/App.java` file: ```java showLineNumbers title="App.java" package org.example; import dev.dbos.transact.DBOS; import dev.dbos.transact.config.DBOSConfig; import dev.dbos.transact.workflow.Workflow; interface Example { public void workflow(); } class ExampleImpl implements Example { private final DBOS dbos; public ExampleImpl(DBOS dbos) { this.dbos = dbos; } private void stepOne() { System.out.println("Step one completed!"); } private void stepTwo() { System.out.println("Step two completed!"); } @Override @Workflow public void workflow() { dbos.runStep(() -> stepOne(), "stepOne"); dbos.runStep(() -> stepTwo(), "stepTwo"); } } public class App { public static void main(String[] args) { var dbosConfig = DBOSConfig.defaultsFromEnv("dbos-java-starter") .withAppVersion("0.1.0"); try (var dbos = new DBOS(dbosConfig)) { Example proxy = dbos.registerProxy(Example.class, new ExampleImpl(dbos)); dbos.launch(); proxy.workflow(); } } } ``` Now, build and run this code with: ```shell ./gradlew run ``` Your program should print output like: ```shell [main] INFO dev.dbos.transact.DBOS - Launching DBOS v1.1.0 [main] INFO dev.dbos.transact.execution.DBOSExecutor - DBOS Executor starting [main] INFO dev.dbos.transact.execution.DBOSExecutor - System Database: jdbc:postgresql://localhost:5432/dbos_java_starter [main] INFO dev.dbos.transact.execution.DBOSExecutor - System Database User name: postgres [main] INFO dev.dbos.transact.execution.DBOSExecutor - Executor ID: local [main] INFO dev.dbos.transact.execution.DBOSExecutor - Application Version: 18d0a53492115958cbeca6ca9cb01b226b11c0cdb9286a91010bc4cba3d8d73f Step one completed! Step two completed! [main] INFO dev.dbos.transact.execution.DBOSExecutor - DBOS Executor stopping [main] INFO dev.dbos.transact.DBOS - DBOS shut down ``` :::info To reduce logging output, we are eliding `com.zaxxer.hikari` and `dev.dbos.transact.migrations` log levels below warn. ::: Of course, an app that runs a single workflow and exits isn't very useful in practice. Let's convert this into a long-running HTTP server so we can trigger workflows on demand and observe DBOS's durable execution in action. Replace the contents of `app/src/main/java/org/example/App.java` with: ```java showLineNumbers title="App.java" package org.example; import dev.dbos.transact.DBOS; import dev.dbos.transact.config.DBOSConfig; import dev.dbos.transact.workflow.Workflow; import java.time.Duration; import io.javalin.Javalin; interface Example { public void workflow(); } class ExampleImpl implements Example { private final DBOS dbos; public ExampleImpl(DBOS dbos) { this.dbos = dbos; } private void stepOne() { System.out.println("\"%s\" Step one completed!".formatted(DBOS.workflowId())); } private void stepTwo() { System.out.println("\"%s\" Step two completed!".formatted(DBOS.workflowId())); } @Override @Workflow public void workflow() { dbos.runStep(() -> stepOne(), "stepOne"); for (int i = 0; i < 5; i++) { System.out.println("Press Control + C to stop \"%s\"...".formatted(DBOS.workflowId())); dbos.sleep(Duration.ofSeconds(1)); } dbos.runStep(() -> stepTwo(), "stepTwo"); } } public class App { public static void main(String[] args) { var dbosConfig = DBOSConfig.defaultsFromEnv("dbos-java-starter") .withAppVersion("0.1.0"); var dbos = new DBOS(dbosConfig); Example proxy = dbos.registerProxy(Example.class, new ExampleImpl(dbos)); Javalin.create(config -> { config.events.serverStarting(dbos::launch); config.events.serverStopping(dbos::shutdown); config.routes.get("/", ctx -> { proxy.workflow(); ctx.result("workflow executed"); }); }).start(8080); } } ``` Then, build and run this code with: ```shell ./gradlew run ``` Next, visit this URL: http://localhost:8080. In your terminal, you should see an output like: ``` [main] INFO io.javalin.Javalin - You are running Javalin 7.0.1 (released February 28, 2026). "477cf44d-3230-4e19-be98-1aef144a3f0c" Step one completed! Press Control + C to stop "477cf44d-3230-4e19-be98-1aef144a3f0c"... Press Control + C to stop "477cf44d-3230-4e19-be98-1aef144a3f0c"... ``` Now, press CTRL+C stop your app. Then, run `./gradlew run` again to restart it. You should see an output like: ``` [main] INFO io.javalin.Javalin - You are running Javalin 7.0.1 (released February 28, 2026). Press Control + C to stop "477cf44d-3230-4e19-be98-1aef144a3f0c"... Press Control + C to stop "477cf44d-3230-4e19-be98-1aef144a3f0c"... Press Control + C to stop "477cf44d-3230-4e19-be98-1aef144a3f0c"... Press Control + C to stop "477cf44d-3230-4e19-be98-1aef144a3f0c"... Press Control + C to stop "477cf44d-3230-4e19-be98-1aef144a3f0c"... "477cf44d-3230-4e19-be98-1aef144a3f0c" Step two completed! ``` You can see how DBOS **recovers your workflow from the last completed step**, executing step two without re-executing step one. Learn more about workflows, steps, and their guarantees [here](./tutorials/workflow-tutorial.md). ### 3. Queues and Parallelism To run many functions concurrently, use DBOS _queues_. To try them out, copy this code into `App.java`: ```java showLineNumbers title="App.java" package org.example; import dev.dbos.transact.DBOS; import dev.dbos.transact.StartWorkflowOptions; import dev.dbos.transact.config.DBOSConfig; import dev.dbos.transact.workflow.QueueOptions; import dev.dbos.transact.workflow.Workflow; import dev.dbos.transact.workflow.WorkflowHandle; import java.time.Duration; import java.util.ArrayList; import java.util.List; import io.javalin.Javalin; interface Example { void taskWorkflow(int i) throws InterruptedException; void queueWorkflow() throws InterruptedException; } class ExampleImpl implements Example { private final DBOS dbos; private Example self; public ExampleImpl(DBOS dbos) { this.dbos = dbos; } public void setSelf(Example self) { this.self = self; } @Override @Workflow public void taskWorkflow(int i) throws InterruptedException { Thread.sleep(Duration.ofSeconds(5)); System.out.printf("Task %d completed!\n", i); } @Override @Workflow public void queueWorkflow() throws InterruptedException { System.out.printf("Starting queueWorkflow!\n"); List> handles = new ArrayList<>(); var options = new StartWorkflowOptions().withQueue("example-queue"); for (int i = 0; i < 10; i++) { final int index = i; var handle = dbos.startWorkflow(() -> self.taskWorkflow(index), options); handles.add(handle); } for (var handle : handles) { handle.getResult(); } System.out.printf("Successfully completed %d workflows!\n", handles.size()); } } public class App { public static void main(String[] args) { var dbosConfig = DBOSConfig.defaultsFromEnv("dbos-java-starter") .withAppVersion("0.1.0"); var dbos = new DBOS(dbosConfig); ExampleImpl impl = new ExampleImpl(dbos); Example proxy = dbos.registerProxy(Example.class, impl); impl.setSelf(proxy); Javalin.create(config -> { config.events.serverStarting(() -> { dbos.launch(); // Queues are stored in the system database, so register them after launch dbos.registerQueue("example-queue", QueueOptions.empty()); }); config.events.serverStopping(dbos::shutdown); config.routes.get("/", ctx -> { proxy.queueWorkflow(); ctx.result("workflow executed"); }); }).start(8080); } } ``` The queue is registered with [`dbos.registerQueue`](./reference/queues.md#dbosregisterqueue) after `dbos.launch()`, because queue configuration is stored in the system database. When you enqueue a function by passing `new StartWorkflowOptions().withQueue("example-queue")` into `dbos.startWorkflow`, DBOS executes it _asynchronously_, running it in the background without waiting for it to finish. `dbos.startWorkflow` returns a handle representing the state of the enqueued function. This example enqueues ten functions, then waits for them all to finish using `.getResult()` to wait for each of their handles. Now, restart your app with: ```shell ./gradlew run ``` Then, visit this URL: http://localhost:8080. Wait five seconds and you should see an output like: ``` Starting queueWorkflow! Task 0 completed! Task 1 completed! Task 2 completed! Task 3 completed! Task 4 completed! Task 5 completed! Task 6 completed! Task 7 completed! Task 8 completed! Task 9 completed! Successfully completed 10 workflows! ``` You can see how all ten steps run concurrently—even though each takes five seconds, they all finish at the same time. Learn more about DBOS queues [here](./tutorials/queue-tutorial.md). ### 4. Connecting to DBOS Conductor [Conductor](../conductor/overview.md) is the control plane for your durable workflows, providing distributed workflow recovery, observability, and management. Once you connect your app to Conductor, you can view and manage all its workflows and queued tasks from the [DBOS Console](https://console.dbos.dev). To connect your app to Conductor, first sign up for an account on the [DBOS Console](https://console.dbos.dev/login-redirect). Then, install [`dbosctl`](../conductor/reference/dbosctl.md), the Conductor command-line client. On Windows, [download a release binary](https://github.com/dbos-inc/dbos-ctl/releases) instead. ```shell curl -sSfL https://raw.githubusercontent.com/dbos-inc/dbos-ctl/main/install.sh | sh ``` Next, configure a `dbosctl` profile and log in. `dbosctl login` prints a URL and a code for you to approve in your browser. ```shell dbosctl config set dbos --managed dbosctl login ``` Then, register your application with Conductor and create an API key. The name you register must match the application name in your DBOS configuration. The key's secret is printed once and cannot be retrieved afterwards, so copy it now. ```shell dbosctl app register dbos-java-starter dbosctl api-key create dbos-java-starter-key ``` Next, supply your API key to your app through the `withConductorKey` configuration option. Update the configuration in `main` to read the key from an environment variable: ```java var dbosConfig = DBOSConfig.defaultsFromEnv("dbos-java-starter") .withAppVersion("0.1.0") .withConductorKey(System.getenv("DBOS_CONDUCTOR_KEY")); ``` Finally, set the `DBOS_CONDUCTOR_KEY` environment variable to the key you created and restart your app: ```shell export DBOS_CONDUCTOR_KEY= ./gradlew run ``` Your app is now connected to Conductor! Launch a workflow by visiting http://localhost:8080, then watch it execute in real time from the [DBOS Console](https://console.dbos.dev). Learn more about Conductor [here](../conductor/overview.md). Congratulations! You've finished the DBOS Java guide. Next, you should: - Learn how to [**add DBOS to your own application**](./integrating-dbos.md). - Check out some [**example applications**](../examples/index.md). --- ## AI Model Prompting(Java) You may want assistance from an AI model in building a DBOS application. To make sure your model has the latest information on how to use DBOS, provide it with this prompt. You may also want to use the [DBOS MCP server](../integrations/mcp.md) so your model can directly access your application's workflows and steps. ### How To Use First, use the click-to-copy button in the top right of the code block to copy the full prompt to your clipboard. Then, paste into your AI tool of choice (for example OpenAI's ChatGPT or Anthropic's Claude). This adds the prompt to your AI model's context, giving it up-to-date instructions on how to build an application with DBOS. If you are using an AI-powered IDE, you can add this prompt to your project's context. For example: - Claude Code: Add the prompt, or a link to it, to your CLAUDE.md file. - Cursor: Add the prompt to [your project rules](https://docs.cursor.com/context/rules-for-ai). - Zed: Copy the prompt to a file in your project, then use the [`/file`](https://zed.dev/docs/assistant/commands?highlight=%2Ffile#file) command to add the file to your context. - GitHub Copilot: Create a [`.github/copilot-instructions.md`](https://docs.github.com/en/copilot/customizing-copilot/adding-repository-custom-instructions-for-github-copilot) file in your repository and add the prompt to it. ### Prompt ````markdown # Build Reliable Applications With DBOS ## Guidelines - Respond in a friendly and concise manner - Ask clarifying questions when requirements are ambiguous - Generate code in Java using the DBOS library. - You MUST import all methods and classes used in the code you generate - DBOS does NOT stand for anything. ## Workflow Guidelines Workflows provide durable execution so you can write programs that are resilient to any failure. In the event of any failure, execution will resume where it left off. Workflows are ordinary Java methods that are comprised of steps, which are also ordinary Java methods. Steps are the restart unit for the workflow, their results are checkpointed to the database. When using DBOS workflows, you should call any function that performs complex operations or accesses external APIs or services as a step using `dbos.runStep` or `@Step`, that way the step will be skipped if execution resumes. Steps that are used once can be made as lambdas and called with `dbos.runStep()`. Steps that are reused can be declared as instance methods annotated with `@Step`. If a workflow is interrupted for any reason (e.g., an executor restarts or crashes), when your program restarts the workflow automatically resumes execution from the last completed step. - If asked to add DBOS to existing code, you MUST ask which function to make a workflow. Do NOT recommend any changes until they have told you what function to make a workflow. Do NOT make a function a workflow unless SPECIFICALLY requested. - When making a function a workflow, you should make all functions it calls steps. Do NOT change the functions in any way. - Do NOT make functions steps unless they are DIRECTLY called by a workflow. - If the workflow function performs a non-deterministic action, you MUST move that action to its own function and make that function a step, or wrap it as a lambda with `dbos.runStep`. Examples of non-deterministic actions include accessing an external API or service, accessing files on disk, generating a random number, or getting the current time. - DBOS workflows and steps should NOT have side effects in memory outside of their own scope. They can access instance or static variables, but they should NOT create or update static, instance, or other variables outside their scope. - Do NOT call any DBOS instance method (`dbos.send`, `dbos.recv`, `dbos.startWorkflow`, `dbos.getEvent`) from a step. Those do their own checkpointing. - Do NOT start workflows from inside a step. - Do NOT call `dbos.setEvent` and `dbos.recv` from outside a workflow function. ## Serializability requirements Workflow method arguments, and step and workflow function return values, are serialized as JSON using jackson. Help the user with any annotations or mapping necessary to get jackson to work seamlessly. For example, the following annotation is sometimes necessary: ```java @com.fasterxml.jackson.annotation.JsonProperty(access = JsonProperty.Access.READ_ONLY) ``` ### Getting set up While it is not necessary to use Gradle, Gradle should be suggested for new projects. Gradle 8 or later should be suggested. Java 21 should be suggested, but any Java 17 or later can be used if the user requests that. The main addition to any project is additional dependencies. For **Gradle** (`build.gradle`): ```groovy dependencies { implementation 'dev.dbos:transact:1.1.0' } ``` For **Maven** (`pom.xml`): ```xml dev.dbos transact 1.1.0 ``` The application will also need a PostgreSQL database. If there is not one already, it can be set up using standard approaches, or with Docker: ```shell docker run -d \ --name dbos-postgres \ -e POSTGRES_PASSWORD=dbos \ -p 5432:5432 \ postgres:17 ``` For convenience, the database credentials should be set in the environment, or passed in to the environment for any shell command that launches DBOS. ``` export PGUSER=postgres export PGPASSWORD=dbos export DBOS_SYSTEM_JDBC_URL=jdbc:postgresql://localhost:5432/ ``` Once some code hase been added to the project, the typical gradlew commands can be used to build and run it. ```shell ./gradlew assemble ./gradlew run ``` ### Use of Spring DBOS examples often use Spring boot, and there are examples of integrating it with the DBOS lifecycle, however DBOS is just a library and can be used by itself, or with other frameworks. The preferred way to use DBOS with Spring Boot is with the `transact-spring-boot-starter` dependency, which auto-configures DBOS from `application.properties`/`application.yml`. See the Spring Boot configuration properties documented in the reference. If integrating manually, expose the `DBOS` instance as a bean and wire your workflow classes to receive it: ```java @Configuration public class DBOSConfig { @Bean public DBOS dbos(DBOSAppService appService) { var config = dev.dbos.transact.config.DBOSConfig.defaults("dbos-starter") .withAppVersion("0.1.0") .withDatabaseUrl(System.getenv("DBOS_SYSTEM_JDBC_URL")) .withDbUser(Objects.requireNonNullElse(System.getenv("PGUSER"), "postgres")) .withDbPassword(Objects.requireNonNullElse(System.getenv("PGPASSWORD"), "dbos")); var dbos = new DBOS(config); var impl = new DBOSAppServiceImpl(dbos); var proxy = dbos.registerProxy(DBOSAppService.class, impl); impl.setProxy(proxy); return dbos; } } // Use SmartLifecycle to control launch/shutdown timing relative to the web server @Component @Lazy(false) public class DBOSLifecycle implements SmartLifecycle { private static final Logger log = LoggerFactory.getLogger(DBOSLifecycle.class); private volatile boolean running = false; @Autowired private DBOS dbos; @Override public void start() { log.info("Launch DBOS"); dbos.launch(); running = true; } @Override public void stop() { log.info("Shut Down DBOS"); try { dbos.shutdown(); } finally { running = false; } } @Override public boolean isRunning() { return running; } @Override public boolean isAutoStartup() { return true; } // Start BEFORE the web server (default is 0). Lower = earlier. @Override public int getPhase() { return -1; } } ``` ### DBOS Lifecycle Guidelines DBOS should be installed and imported from the `dev.dbos.transact` package. Use any of the following imports if they are necessary. ```java import dev.dbos.transact.DBOS; import dev.dbos.transact.DBOSClient; import dev.dbos.transact.EnqueueOptions; import dev.dbos.transact.StartWorkflowOptions; import dev.dbos.transact.config.DBOSConfig; import dev.dbos.transact.exceptions.DBOSQueueDuplicatedException; import dev.dbos.transact.workflow.ForkOptions; import dev.dbos.transact.workflow.ListWorkflowsInput; import dev.dbos.transact.workflow.QueueName; import dev.dbos.transact.workflow.QueueOptions; import dev.dbos.transact.workflow.ScheduleStatus; import dev.dbos.transact.workflow.SerializationStrategy; import dev.dbos.transact.workflow.Step; import dev.dbos.transact.workflow.StepOptions; import dev.dbos.transact.workflow.Timeout; import dev.dbos.transact.workflow.Workflow; import dev.dbos.transact.workflow.WorkflowClassName; import dev.dbos.transact.workflow.WorkflowHandle; import dev.dbos.transact.workflow.WorkflowSchedule; import dev.dbos.transact.workflow.WorkflowState; import dev.dbos.transact.workflow.WorkflowStatus; ``` Any DBOS program MUST create a `DBOS` instance, register workflows, call `launch()`, then register any queues it uses. For simple cases, this can be in its main function, like so. You MUST use this default configuration (changing the name 'dbos-java-starter' to the real app name as appropriate) unless otherwise specified. ```java public static void main(String[] args) throws Exception { Logger root = (Logger) LoggerFactory.getLogger(Logger.ROOT_LOGGER_NAME); root.setLevel(Level.INFO); DBOSConfig config = DBOSConfig.defaults("dbos-java-starter") .withAppVersion("0.1.0") .withDatabaseUrl(System.getenv("DBOS_SYSTEM_JDBC_URL")) .withDbUser(System.getenv("PGUSER")) .withDbPassword(System.getenv("PGPASSWORD")); DBOS dbos = new DBOS(config); ExampleImpl impl = new ExampleImpl(dbos); Example proxy = dbos.registerProxy(Example.class, impl); dbos.launch(); } ``` ### Workflow and Steps Examples Simple example: Use of background execution: Use of queues: #### Scheduled Workflows You can schedule DBOS workflows to run automatically on a cron schedule. Scheduled workflows are **exactly-once**: DBOS assigns each firing a deterministic workflow ID derived from the schedule name and scheduled time. The recommended way to declare schedules is to call `dbos.applySchedules()` after `dbos.launch()`. This atomically creates or replaces the named schedules so your code is always the source of truth: ```java dbos.applySchedules( new WorkflowSchedule("every-minute", "everyMinute", "com.example.ExampleImpl", "0 * * * * *"), new WorkflowSchedule("daily-report", "dailyReport", "com.example.ExampleImpl", "0 0 9 * * *") .withCronTimezone(ZoneId.of("America/New_York")) ); ``` A workflow invoked by a `WorkflowSchedule` must accept exactly two arguments: an `Instant` for the scheduled fire time and an `Object` for the optional context: ```java @Workflow public void everyMinute(Instant scheduled, Object context) { ... } ``` ##### WorkflowSchedule ```java new WorkflowSchedule(String scheduleName, String workflowName, String className, String cron) ``` - **scheduleName**: A unique name for this schedule (used for management operations). - **workflowName**: The name of the workflow method to invoke (as registered, or as set by `@Workflow(name=...)`). - **className**: The fully-qualified class name, or the short name set by `@WorkflowClassName`. - **cron**: A [Spring 5.3+ CronExpression](https://docs.spring.io/spring-framework/docs/current/javadoc-api/org/springframework/scheduling/support/CronExpression.html). Common optional configuration via `with` methods: | Method | Description | |--------|-------------| | `withCronTimezone(ZoneId)` | Interpret the cron in this timezone (default: UTC). | | `withAutomaticBackfill(true)` | Retroactively start any firings missed while the app was down. | | `withQueueName(String)` | Enqueue executions on this queue instead of the default scheduler queue. | | `withStatus(ScheduleStatus.PAUSED)` | Create the schedule in a paused state. | | `withContext(Object)` | Attach a serializable context object passed to the workflow. | ##### Runtime Schedule Management Schedules can be created, paused, resumed, and deleted at runtime: ```java // Create (throws if name already exists) dbos.createSchedule(new WorkflowSchedule("on-demand", "processReport", "com.example.ReportImpl", "0 0 * * * *")); // Pause and resume dbos.pauseSchedule("daily-report"); dbos.resumeSchedule("daily-report"); // Inspect Optional s = dbos.getSchedule("every-minute"); List active = dbos.listSchedules(List.of(ScheduleStatus.ACTIVE), null, null); // Delete dbos.deleteSchedule("on-demand"); // Fire immediately outside its normal cadence WorkflowHandle handle = dbos.triggerSchedule("daily-report"); ``` To retroactively run missed firings for a time range: ```java List> handles = dbos.backfillSchedule("every-minute", Instant.parse("2025-01-01T00:00:00Z"), Instant.parse("2025-01-02T00:00:00Z")); ``` ### Workflow Documentation Workflows provide **durable execution** so you can write programs that are **resilient to any failure**. Workflows are comprised of steps, which wrap ordinary Java functions. If a workflow is interrupted for any reason (e.g., an executor restarts or crashes), when your program restarts the workflow automatically resumes execution from the last completed step. The recovery mechanism requires that workflow methods must be registered. Registration creates a proxy that adds durability to the registered workflow. For Java, this means defining both an interface and implementation class, annotating the implementation with @Workflow and @Step, and then calling `dbos.registerProxy`. #### @Workflow ```java public @interface Workflow { String name(); int maxRecoveryAttempts(); SerializationStrategy serializationStrategy(); } ``` An annotation that can be applied to a class method to mark it as a durable workflow. **Parameters:** - **name**: The workflow name. Must be unique within the class. Defaults to method name. - **maxRecoveryAttempts**: Optionally configure the maximum number of times execution of a workflow may be attempted. Acts as a dead letter queue so that a buggy workflow that crashes its application does not do so infinitely. If a workflow exceeds this limit, its status is set to `MAX_RECOVERY_ATTEMPTS_EXCEEDED`. - **serializationStrategy**: The default serialization strategy for local invocations of this workflow. Set to `SerializationStrategy.PORTABLE` to test cross-language interoperability. ### Methods #### registerProxy ```java T registerProxy(Class interfaceClass, T implementation) T registerProxy(Class interfaceClass, T implementation, String instanceName) ``` Register the workflows in a class, returning a proxy object from which the class methods may be invoked as durable workflows. All workflows must be registered before DBOS is launched. **Example Syntax:** ```java interface Example { public void workflow(); } class ExampleImpl implements Example { @Workflow public void workflow() { return; } } DBOS dbos = new DBOS(config); Example proxy = dbos.registerProxy(Example.class, new ExampleImpl()); dbos.launch(); proxy.workflow(); ``` **Parameters:** - **interfaceClass**: The interface class whose workflows are to be registered. - **implementation**: An instance of the class whose workflows to register. - **instanceName**: A unique name for this class instance. Use only when you are creating multiple instances of a class and your workflow depends on class instance variables. When DBOS needs to recover a workflow belonging to that class, it looks up the class instance using `instanceName` so it can recover the workflow using the right instance of its class. ### Starting Workflows In The Background ```java WorkflowHandle startWorkflow(ThrowingSupplier workflow) WorkflowHandle startWorkflow(ThrowingSupplier workflow, StartWorkflowOptions options) WorkflowHandle startWorkflow(ThrowingRunnable workflow) WorkflowHandle startWorkflow(ThrowingRunnable workflow, StartWorkflowOptions options) ``` Start a workflow in the background and return a handle to it. Optionally enqueue it on a DBOS queue. The `startWorkflow` method resolves after the workflow is durably started; at this point the workflow is guaranteed to run to completion even if the app is interrupted. **Example Syntax**: ```java interface Example { public void workflow(); } class ExampleImpl implements Example { @Workflow public void workflow() { return; } } Example proxy = dbos.registerProxy(Example.class, new ExampleImpl(dbos)); dbos.launch(); dbos.startWorkflow(() -> proxy.workflow(), new StartWorkflowOptions()); ``` ##### StartWorkflowOptions `StartWorkflowOptions` is a with-based configuration record for parameterizing `dbos.startWorkflow`. All fields are optional. **Constructors:** ```java new StartWorkflowOptions() ``` Create workflow options with all fields set to their defaults. **Methods:** - **`withWorkflowId(String workflowId)`** - Set the workflow ID of this workflow. - **`withQueue(QueueName queue)`** / **`withQueue(String queueName)`** - Instead of starting the workflow directly, enqueue it on this queue. The queue must be registered first. `new StartWorkflowOptions(QueueName queue)` is equivalent; do NOT use `new StartWorkflowOptions(String)` for a queue, because its `String` argument is a workflow ID. - **`withTimeout(Duration timeout)`** / **`withTimeout(long value, TimeUnit unit)`** - Set a timeout for this workflow. When the timeout expires, the workflow **and all its children** are cancelled. Cancelling a workflow sets its status to `CANCELLED` and preempts its execution at the beginning of its next step. Timeouts are **start-to-completion**: if a workflow is enqueued, the timeout does not begin until the workflow is dequeued and starts execution. Also, timeouts are **durable**: they are stored in the database and persist across restarts, so workflows can have very long timeouts. Timeout deadlines are propagated to child workflows by default. To detach a child workflow from its parent's timeout, start it with its own explicit timeout or use `withNoTimeout()`. - **`withNoTimeout()`** - Explicitly remove any inherited timeout from this workflow. - **`withDeadline(Instant deadline)`** - Set an absolute deadline for this workflow. The workflow and all its children are cancelled if still running at the deadline. - **`withDeduplicationId(String deduplicationId)`** - May only be used when enqueuing. At any given time, only one workflow with a specific deduplication ID can be enqueued in the specified queue. - **`withPriority(Integer priority)`** - May only be used when enqueuing. Priority values can range from `0` to `2,147,483,647`, where a low number indicates a higher priority. A negative priority throws `IllegalArgumentException`. Workflows without assigned priorities have priority `0`, the highest priority. - **`withQueuePartitionKey(String key)`** - Set a queue partition key. Required on partitioned queues (queues with a per-partition limit), and not allowed on other queues. - **`withDelay(Duration delay)`** - Delay the start of the workflow by the specified duration after it is dequeued. - **`withAppVersion(String appVersion)`** - Tag the workflow with a specific application version. One common use-case for workflows is building reliable background tasks that keep running even when your program is interrupted, restarted, or crashes. You can use startWorkflow to start a workflow in the background. When you start a workflow this way, it returns a workflow handle, from which you can access information about the workflow or wait for it to complete and retrieve its result. Here's an example: ```java class ExampleImpl implements Example { @Workflow public String backgroundTask(String input) { // ... return output; } } public void runWorkflowExample(DBOS dbos, Example proxy) throws Exception { // Start the background task WorkflowHandle handle = dbos.startWorkflow( () -> proxy.backgroundTask("input"), new StartWorkflowOptions() ); // Wait for the background task to complete and retrieve its result String result = handle.getResult(); System.out.println("Workflow result: " + result); } ``` After starting a workflow in the background, you can use retrieveWorkflow to retrieve a workflow's handle from its ID. You can also retrieve a workflow's handle from outside of your DBOS application with DBOSClient.retrieveWorkflow. If you need to run many workflows in the background and manage their concurrency or flow control, use queues. ### Workflow IDs and Idempotency Every time you execute a workflow, that execution is assigned a unique ID, by default a UUID. You can access this ID from the DBOS.workflowId method. Workflow IDs are useful for communicating with workflows and developing interactive workflows. You can set the workflow ID of a workflow using `withWorkflowId` when calling `startWorkflow`. Workflow IDs are **globally unique** within your application. An assigned workflow ID acts as an idempotency key: if a workflow is called multiple times with the same ID, it executes only once. This is useful if your operations have side effects like making a payment or sending an email. For example: ```java class ExampleImpl implements Example { @Workflow public String exampleWorkflow() { System.out.println("Running workflow with ID: " + DBOS.workflowId()); // ... return "success"; } } public void example(DBOS dbos, Example proxy) throws Exception { String myID = "unique-workflow-id-123"; WorkflowHandle handle = dbos.startWorkflow( () -> proxy.exampleWorkflow(), new StartWorkflowOptions().withWorkflowId(myID) ); String result = handle.getResult(); System.out.println("Result: " + result); } ``` ### Determinism Workflows are in most respects normal Java methods. They can have loops, branches, conditionals, and so on. However, a workflow method must be **deterministic**: if called multiple times with the same inputs, it should invoke the same steps with the same inputs in the same order (given the same return values from those steps). If you need to perform a non-deterministic operation like accessing the database, calling a third-party API, generating a random number, or getting the local time, you shouldn't do it directly in a workflow method. Instead, you should do all non-deterministic operations in steps. :::warning Java's threading and concurrency APIs are non-deterministic. You should use them only inside steps. ::: For example, **don't do this**: ```java @Workflow public String exampleWorkflow() { int randomChoice = new Random().nextInt(2); if (randomChoice == 0) { return dbos.runStep(() -> stepOne(), "stepOne"); } else { return dbos.runStep(() -> stepTwo(), "stepTwo"); } } ``` Instead, do this: ```java private int generateChoice() { return new Random().nextInt(2); } @Workflow public String exampleWorkflow() { int randomChoice = dbos.runStep(() -> generateChoice(), "generateChoice"); if (randomChoice == 0) { return dbos.runStep(() -> stepOne(), "stepOne"); } else { return dbos.runStep(() -> stepTwo(), "stepTwo"); } } ``` ### Workflow Timeouts You can set a timeout for a workflow using withTimeout in `StartWorkflowOptions`. When the timeout expires, the workflow and all its children are cancelled. Cancelling a workflow sets its status to CANCELLED and preempts its execution at the beginning of its next step. You can detach a child workflow from its parent's timeout by starting it with a custom timeout using `withTimeout`. Timeouts are **start-to-completion**: if a workflow is enqueued, the timeout does not begin until the workflow is dequeued and starts execution. Also, timeouts are durable: they are stored in the database and persist across restarts, so workflows can have very long timeouts. ```java @Workflow public void exampleWorkflow() throws InterruptedException { // Workflow implementation } WorkflowHandle handle = dbos.startWorkflow( () -> proxy.exampleWorkflow(), new StartWorkflowOptions().withTimeout(Duration.ofHours(12)) ); ``` ### Durable Sleep You can use `dbos.sleep` to put your workflow to sleep for any period of time. This sleep is **durable**. DBOS saves the wakeup time in the database so that even if the workflow is interrupted and restarted multiple times while sleeping, it still wakes up on schedule. Sleeping is useful for scheduling work to run in the future (even days, weeks, or months from now). For example: ```java public String runTask(String task) { // Execute the task... return "task completed"; } @Workflow public String exampleWorkflow(float timeToSleepSeconds, String task) throws InterruptedException { // Sleep for the specified duration dbos.sleep(Duration.ofMillis((long)(timeToSleepSeconds*1000))); // Execute the task after sleeping String result = dbos.runStep( () -> runTask(task), "runTask" ); return result; } ``` ### Workflow Versioning and Recovery Because DBOS recovers workflows by re-executing them using information saved in the database, a workflow cannot safely be recovered if its code has changed since the workflow was started. To guard against this, DBOS _versions_ applications and their workflows. Always set the application version explicitly with `withAppVersion`, and change it whenever you deploy changed workflow code: ```java var config = DBOSConfig.defaultsFromEnv("my-app").withAppVersion("1.0.0"); ``` If no version is set, DBOS computes one at launch from a hash of the DBOS version, the application name, and the name, signature, and bytecode of each registered workflow method. This computed version is only a fallback: it does not change when you change steps or other code a workflow calls, and it changes on every DBOS upgrade even when your workflows did not. All workflows are tagged with the application version on which they started. When DBOS tries to recover workflows, it only recovers workflows whose version matches the current application version. This prevents unsafe recovery of workflows that depend on different code. You cannot change the version of a workflow, but you can use `dbos.forkWorkflow` to restart a workflow from a specific step on a specific code version. ### Workflow Communication DBOS provides a few different ways to communicate with your workflows. You can: - Send messages to workflows - Publish events from workflows for clients to read ### Workflow Messaging and Notifications You can send messages to a specific workflow. This is useful for signaling a workflow or sending notifications to it while it's running. ##### Send ```java void send(String destinationId, Object message, String topic, String idempotencyKey) ``` You can call `dbos.send()` to send a message to a workflow. Messages can optionally be associated with a topic and are queued on the receiver per topic. You can also call `send` from outside of your DBOS application with the DBOS Client. ##### Recv ```java Optional recv(String topic, Duration timeout) ``` Workflows can call `dbos.recv()` to receive messages sent to them, optionally for a particular topic. Each call to `recv()` waits for and consumes the next message to arrive in the queue for the specified topic, returning `null` if the wait times out. If the topic is not specified, this method only receives messages sent without a topic. ##### Messages Example Messages are especially useful for sending notifications to a workflow. For example, in a payments system, after redirecting customers to a payments page, the checkout workflow must wait for a notification that the user has paid. To wait for this notification, the payments workflow uses `recv()`, executing failure-handling code if the notification doesn't arrive in time: ```java interface Checkout { void checkoutWorkflow(); } class CheckoutImpl implements Checkout { private static final String PAYMENT_STATUS = "payment_status"; @Workflow public void checkoutWorkflow() { // Validate the order, redirect the customer to a payments page, // then wait for a notification. String paymentStatus = dbos.recv(PAYMENT_STATUS, Duration.ofSeconds(60)).orElse(null); if (paymentStatus != null && paymentStatus.equals("paid")) { // Handle a successful payment. } else { // Handle a failed payment or timeout. } } } ``` An endpoint waits for the payment processor to send the notification, then uses `send()` to forward it to the workflow: ```java app.post("/payment_webhook/{workflow_id}/{payment_status}", ctx -> { String workflowId = ctx.pathParam("workflow_id"); String paymentStatus = ctx.pathParam("payment_status"); // Send the payment status to the checkout workflow. dbos.send(workflowId, paymentStatus, PAYMENT_STATUS, null); ctx.result("Payment status sent"); }); ``` ##### Reliability Guarantees All messages are persisted to the database, so if `send` completes successfully, the destination workflow is guaranteed to be able to `recv` it. If you're sending a message from a workflow, DBOS guarantees exactly-once delivery. If you're sending a message from normal Java code, you can use a unique workflow ID to guarantee exactly-once delivery. ### Workflow Events Workflows can publish _events_, which are key-value pairs associated with the workflow. They are useful for publishing information about the status of a workflow or to send a result to clients while the workflow is running. ##### setEvent ```java void setEvent(String key, Object value) ``` Any workflow can call `dbos.setEvent` to publish a key-value pair, or update its value if it has already been published. ##### getEvent ```java Optional getEvent(String workflowId, String key, Duration timeout) ``` You can call `dbos.getEvent` to retrieve the value published by a particular workflow identity for a particular key. If the event does not yet exist, this call waits for it to be published, returning `null` if the wait times out. You can also call `getEvent` from outside of your DBOS application with DBOS Client. ##### Events Example Events are especially useful for writing interactive workflows that communicate information to their caller. For example, in a checkout system, after validating an order, the checkout workflow needs to send the customer a unique payment ID. To communicate the payment ID to the customer, it uses events. The payments workflow emits the payment ID using `setEvent()`: ```java interface Checkout { void checkoutWorkflow(); } class CheckoutImpl implements Checkout { private static final String PAYMENT_ID = "payment_id"; @Workflow public void checkoutWorkflow() { // ... validation logic String paymentId = generatePaymentId(); dbos.setEvent(PAYMENT_ID, paymentId); // ... continue processing } } ``` The handler that originally started the workflow uses `getEvent()` to await this payment ID, then returns it: ```java app.post("/checkout/{idempotency_key}", ctx -> { String idempotencyKey = ctx.pathParam("idempotency_key"); // Idempotently start the checkout workflow in the background. WorkflowHandle handle = dbos.startWorkflow( () -> checkoutProxy.checkoutWorkflow(), new StartWorkflowOptions().withWorkflowId(idempotencyKey) ); // Wait for the checkout workflow to send a payment ID, then return it. String paymentId = dbos.getEvent(handle.workflowId(), PAYMENT_ID, Duration.ofSeconds(60)).orElse(null); if (paymentId == null) { ctx.status(404); ctx.result("Checkout failed to start"); } else { ctx.result(paymentId); } }); ``` All events are persisted to the database, so the latest version of an event is always retrievable. Additionally, if `getEvent` is called in a workflow, the retrieved value is persisted in the database so workflow recovery can use that value, even if the event is later updated. #### Reliability Guarantees When using DBOS workflows, you should call any method that performs complex operations or accesses external APIs or services as a _step_. If a workflow is interrupted, upon restart it automatically resumes execution from the **last completed step**. You can use `runStep` to call a method as a step. A step can return any serializable value and may throw checked or unchecked exceptions. Here's a simple example: ```java class ExampleImpl implements Example { private int generateRandomNumber(int n) { return new Random().nextInt(n); } @Workflow public int workflowFunction(int n) { int randomNumber = dbos.runStep( () -> generateRandomNumber(n), // Run generateRandomNumber as a checkpointed step "generateRandomNumber" // A name for the step ); return randomNumber; } } ``` You should make a method a step if you're using it in a DBOS workflow and it performs a **nondeterministic** operation. A nondeterministic operation is one that may return different outputs given the same inputs. Common nondeterministic operations include: - Accessing an external API or service. - Accessing files on disk. - Generating a random number. - Getting the current time. You **cannot** call, start, or enqueue workflows from within steps. These operations should be performed from workflow methods. You can call one step from another step, but the called step becomes part of the calling step's execution rather than functioning as a separate step. ### Configurable Retries You can optionally configure a step to automatically retry any error a set number of times with exponential backoff. This is useful for automatically handling transient failures, like making requests to unreliable APIs. Retries are configurable through step options that can be passed to `runStep`. Available retry configuration options (via `StepOptions`): - `withMaxAttempts(int n)` - Maximum number of attempts (default: 1, i.e. no retries). Set to >1 to enable retries. - `withRetryInterval(Duration interval)` - How long to wait before the first retry (default: 1 second). - `withBackoffRate(double rate)` - Exponential backoff multiplier between retries (default: 2.0). For example, let's write a step that fetches a website, and configure it to retry failures up to 10 times: ```java class ExampleImpl implements Example { private String fetchStep(String url) throws Exception { HttpClient client = HttpClient.newHttpClient(); HttpRequest request = HttpRequest.newBuilder() .uri(URI.create(url)) .build(); HttpResponse response = client.send( request, HttpResponse.BodyHandlers.ofString() ); return response.body(); } @Workflow public String fetchWorkflow(String inputURL) throws Exception { return dbos.runStep( () -> fetchStep(inputURL), new StepOptions("fetchFunction") .withMaxAttempts(10) .withRetryInterval(Duration.ofMillis(500)) .withBackoffRate(2.0) ); } } ``` If a step exhausts all retry attempts, it throws an exception to the calling workflow. ### Queues Workflow queues ensure that workflow functions will be run, without starting them immediately. Queues are useful for controlling the number of workflows run in parallel, or the rate at which they are started. Queues are stored in the system database. You MUST register a queue with `dbos.registerQueue` **after** calling `dbos.launch()`, and before enqueueing workflows on it. Do NOT use `new Queue(...)`, `dbos.registerQueue(Queue)`, `dbos.registerQueues(...)`, or `dbos.getQueue(...)`: they declare deprecated in-memory queues that no other process can see. ```java dbos.launch(); dbos.registerQueue("example-queue", QueueOptions.setWorkerConcurrency(5)); ``` #### dbos.registerQueue ```java void registerQueue(String name, QueueOptions options) void registerQueue(String name, QueueOptions options, QueueConflictResolution onConflict) ``` Register a queue in the system database. Must be called after `dbos.launch()`. Queue names must be unique within the system database; `_dbos_internal_queue` is reserved. To register a queue with default settings (no limits), pass `QueueOptions.empty()`. If a queue with this name already exists in the database, `onConflict` controls whether its configuration is overwritten: - `QueueConflictResolution.UPDATE_IF_LATEST_VERSION` (default): overwrite only when this process runs the latest registered application version. Safe for rolling deploys. - `QueueConflictResolution.ALWAYS_UPDATE`: always overwrite. - `QueueConflictResolution.NEVER_UPDATE`: leave the existing configuration unchanged. #### QueueOptions `QueueOptions` holds a queue's configuration. Build it with a static `set...` factory and chain `and...` methods: ```java QueueOptions.setConcurrency(10) .andWorkerConcurrency(2) .andRateLimit(100, Duration.ofSeconds(60)); ``` Static factories (each has a matching `and...` method for chaining): - **`setConcurrency(Integer value)`**: The maximum number of workflows from this queue that may run concurrently across all DBOS processes. - **`setWorkerConcurrency(Integer value)`**: The maximum number of workflows from this queue that may run concurrently within a single DBOS process. If `concurrency` is also set, must not exceed it. - **`setRateLimit(Integer max, Duration period)`** / **`setRateLimit(int limit, long period, TimeUnit unit)`**: The maximum number of workflows that may be started from this queue in a rolling period, across all processes. - **`setPartitionConcurrency(Integer value)`**: The maximum number of workflows from any one partition that may run concurrently across all processes. - **`setPartitionWorkerConcurrency(Integer value)`**: The maximum number of workflows from any one partition that may run concurrently within a single process. - **`setPartitionRateLimit(Integer max, Duration period)`** / **`setPartitionRateLimit(int limit, long period, TimeUnit unit)`**: The maximum number of workflows that may be started from any one partition in a rolling period. - **`setPriorityEnabled(boolean value)`**: Deprecated since 1.1 and ignored. Every queue already dequeues in priority order; do not set it. - **`setPollingInterval(Duration value)`**: How often workers poll the database for new workflows on this queue. Defaults to 1 second. - **`empty()`**: No settings. Do NOT use `setPartitionQueue`/`andPartitionQueue`; they are deprecated. Setting any per-partition limit is what partitions a queue. #### Enqueueing Workflows Enqueue a workflow by passing a `QueueName` to `StartWorkflowOptions`: ```java WorkflowHandle handle = dbos.startWorkflow( () -> proxy.processTask(task), new StartWorkflowOptions(QueueName.of("example-queue"))); ``` `new StartWorkflowOptions().withQueue("example-queue")` also works. Do NOT write `new StartWorkflowOptions("example-queue")`: the single-`String` constructor takes a **workflow ID**, not a queue name. Enqueueing on a queue that isn't registered throws. Queued workflows are started in priority order, and in first-in, first-out (FIFO) order among workflows of the same priority. To enqueue a workflow by name, including one implemented by another application or in another language that shares the system database, use `dbos.enqueueWorkflow` with the same `EnqueueOptions` as `DBOSClient` (see below): ```java WorkflowHandle enqueueWorkflow(EnqueueOptions options, Object[] args) ``` #### Queue Example Here's an example of a workflow using a queue to process tasks in parallel: ```java interface Example { String processTask(String task); List processTasks(List tasks); } class ExampleImpl implements Example { private final DBOS dbos; private Example proxy; ExampleImpl(DBOS dbos) { this.dbos = dbos; } void setProxy(Example proxy) { this.proxy = proxy; } @Workflow public String processTask(String task) { // ... return task; } @Workflow public List processTasks(List tasks) { var handles = new ArrayList>(); // Enqueue each task so all tasks are processed concurrently. var queueName = QueueName.of("task-queue"); for (var task : tasks) { handles.add(dbos.startWorkflow( () -> proxy.processTask(task), new StartWorkflowOptions(queueName))); } // Wait for each task to complete and retrieve its result. var results = new ArrayList(); for (var handle : handles) { results.add(handle.getResult()); } return results; } } // In main: DBOS dbos = new DBOS(config); ExampleImpl impl = new ExampleImpl(dbos); Example proxy = dbos.registerProxy(Example.class, impl); impl.setProxy(proxy); dbos.launch(); dbos.registerQueue("task-queue", QueueOptions.setWorkerConcurrency(5)); ``` #### Reconfiguring Queues at Runtime Queue configuration lives in the system database, so you can change it at runtime without redeploying or restarting your workers. Workers pick up the new configuration on their next polling iteration. ```java void updateQueue(String name, QueueOptions options) Optional findQueue(String name) List listQueues() boolean deleteQueue(String name) ``` `updateQueue` writes only the fields set in `options`; fields not set are left unchanged. To clear a limit, set it to `null`: ```java dbos.updateQueue("example-queue", QueueOptions.setConcurrency(20)); dbos.updateQueue("example-queue", QueueOptions.setRateLimit(null, null)); // remove the rate limit ``` `findQueue` returns a `Queue` record, a snapshot of the queue's configuration with accessors such as `concurrency()`, `workerConcurrency()`, `rateLimit()`, `partitionConcurrency()`, `partitionWorkerConcurrency()`, `partitionRateLimit()`, and `pollingInterval()`. Do NOT construct `Queue` yourself. **Warning:** workflows already enqueued on a deleted queue can no longer be dequeued or executed until a queue with the same name is registered again. Cancel or drain pending workflows before deleting a queue. If workflows are already stuck on a deleted queue, move them to a registered queue with `dbos.resumeWorkflow(workflowId, "new-queue")` (or `dbos.resumeWorkflows(workflowIds, "new-queue")`); find them with `dbos.listWorkflows(new ListWorkflowsInput().withQueueName("old-queue").withStatus(WorkflowState.ENQUEUED))`. `DBOSClient` has the same `registerQueue`, `updateQueue`, `findQueue`, `listQueues`, and `deleteQueue` methods, for managing queues from outside your application. #### Enqueueing from Another Application Often, you want to enqueue a workflow from outside your DBOS application. For example, let's say you have an API server and a data processing service. You're using DBOS to build a durable data pipeline in the data processing service. When the API server receives a request, it should enqueue the data pipeline for execution on the data processing service. You can use `DBOSClient` to register queues and enqueue workflows from outside your DBOS application by connecting directly to your DBOS application's system database. Since `DBOSClient` is designed to be used from outside your DBOS application, workflow and queue metadata must be specified explicitly. For example, this code registers `pipeline-queue` and enqueues the `dataPipeline` workflow of class `com.example.PipelineImpl` on it with `task` as an argument. ```java var client = new DBOSClient( dbUrl, dbUser, dbPassword, "dbos", null, false, // The name of the application that runs the data pipeline "data-processing-service"); client.registerQueue("pipeline-queue", QueueOptions.empty()); var options = new EnqueueOptions( "dataPipeline", "com.example.PipelineImpl", QueueName.of("pipeline-queue")); WorkflowHandle handle = client.enqueueWorkflow(options, new Object[] {task}); ``` Note: `client.registerQueue` defaults `onConflict` to `QueueConflictResolution.ALWAYS_UPDATE` because clients are not associated with an application version. `UPDATE_IF_LATEST_VERSION` is not supported on the client and throws. See [DBOSClient](#dbosclient) below for the full client API. #### Managing Concurrency You can control how many workflows from a queue run simultaneously by configuring concurrency limits. This helps prevent resource exhaustion when workflows consume significant memory or processing power. ##### Worker Concurrency Worker concurrency sets the maximum number of workflows from a queue that can run concurrently on a single DBOS process. This is particularly useful for resource-intensive workflows to avoid exhausting the resources of any process. For example, this queue has a worker concurrency of 5, so each process will run at most 5 workflows from this queue simultaneously: ```java dbos.registerQueue("example-queue", QueueOptions.setWorkerConcurrency(5)); ``` ##### Global Concurrency Global concurrency limits the total number of workflows from a queue that can run concurrently across all DBOS processes in your application. For example, this queue will have a maximum of 10 workflows running simultaneously across your entire application. :::warning Worker concurrency limits are recommended for most use cases. Take care when using a global concurrency limit as any `PENDING` workflow on the queue counts toward the limit, including workflows from previous application versions. ::: ```java dbos.registerQueue("example-queue", QueueOptions.setConcurrency(10)); ``` ##### In-Order Processing You can use a queue with `concurrency` of 1 to guarantee sequential, in-order processing of events. Only a single event will be processed at a time. For example, this processes events sequentially in the order of their arrival: ```java dbos.launch(); dbos.registerQueue("in-order-queue", QueueOptions.setConcurrency(1)); // Called for each incoming event, for example from an HTTP handler public void onEvent(String event) { dbos.startWorkflow( () -> proxy.processEvent(event), new StartWorkflowOptions(QueueName.of("in-order-queue"))); } ``` #### Rate Limiting You can set _rate limits_ for a queue, limiting the number of workflows that it can start in a given period. Rate limits are global across all DBOS processes using this queue. For example, this queue has a limit of 100 workflows with a period of 60 seconds, so it may not start more than 100 workflows in 60 seconds: ```java dbos.registerQueue("example-queue", QueueOptions.setRateLimit(100, Duration.ofSeconds(60))); ``` Rate limits are especially useful when working with a rate-limited API, such as many LLM APIs. #### Setting Timeouts You can set a timeout for an enqueued workflow with `withTimeout` in `StartWorkflowOptions`. When the timeout expires, the workflow **and all its children** are cancelled. Cancelling a workflow sets its status to `CANCELLED` and preempts its execution at the beginning of its next step. Timeouts are **start-to-completion**: a workflow's timeout does not begin until the workflow is dequeued and starts execution. Also, timeouts are **durable**: they are stored in the database and persist across restarts, so workflows can have very long timeouts. Example syntax: ```java WorkflowHandle handle = dbos.startWorkflow( () -> proxy.processTask(task), new StartWorkflowOptions(QueueName.of("example-queue")).withTimeout(Duration.ofMinutes(10))); ``` #### Partitioning Queues You can partition a queue to apply flow control separately to each partition key, for example to run at most one workflow at a time per user. Setting any per-partition limit (`setPartitionConcurrency`, `setPartitionWorkerConcurrency`, or `setPartitionRateLimit`) partitions the queue. Every workflow enqueued on a partitioned queue must have a partition key, set with `withQueuePartitionKey`; workflows on other queues must not have one. A partitioned queue can also have queue-wide limits, which bound the queue as a whole across all partitions. For example, this queue runs at most one workflow per user at a time, and at most 20 in total: ```java dbos.registerQueue("per-user-queue", QueueOptions.setConcurrency(20).andPartitionConcurrency(1)); dbos.startWorkflow( () -> proxy.processTask(task), new StartWorkflowOptions(QueueName.of("per-user-queue")).withQueuePartitionKey(userId)); ``` A partitioned queue enforces its per-partition limits **and** its queue-wide limits (`concurrency`, `workerConcurrency`, and `rateLimit`) at the same time. This lets you protect your workers from overload while still fairly distributing work between partitions. For example, this "fair queue" runs at most one task per user, but no more than 10 tasks on any single process: ```java dbos.registerQueue("fair-queue", QueueOptions.setPartitionConcurrency(1).andWorkerConcurrency(10)); ``` Each queue-wide limit has a per-partition counterpart, so you can mix and match them freely: ```java // At most 100 tasks running globally and 25 running per tenant, // at most 10 tasks running per process and 2 per tenant per process, // and at most 1000 tasks started per minute globally and 50 per tenant. dbos.registerQueue("tenant-queue", QueueOptions.setConcurrency(100) .andWorkerConcurrency(10) .andRateLimit(1000, Duration.ofSeconds(60)) .andPartitionConcurrency(25) .andPartitionWorkerConcurrency(2) .andPartitionRateLimit(50, Duration.ofSeconds(60))); ``` When both are set, each per-partition concurrency limit must be less than or equal to its queue-wide counterpart, and `partitionWorkerConcurrency` must be less than or equal to `partitionConcurrency`; limits that are not set are not compared. Deduplication is not supported on partitioned queues: setting both `withQueuePartitionKey` and `withDeduplicationId` throws `IllegalArgumentException`. :::warning Workflows already on a queue without a partition key are never dequeued once the queue becomes partitioned. Drain a queue before adding its first per-partition limit. If workflows are already stuck this way, move them to a queue that is not partitioned with `dbos.resumeWorkflow(workflowId, "new-queue")` (or `dbos.resumeWorkflows(workflowIds, "new-queue")`); find them with `dbos.listWorkflows(new ListWorkflowsInput().withQueueName("partitioned-queue").withStatus(WorkflowState.ENQUEUED))`. ::: #### Deduplication You can set a deduplication ID for an enqueued workflow with `withDeduplicationId` when calling `startWorkflow`. At any given time, only one workflow with a specific deduplication ID can be enqueued in the specified queue. If a workflow with a deduplication ID is currently enqueued, delayed, or actively executing (status `ENQUEUED`, `DELAYED`, or `PENDING`), subsequent workflow enqueue attempts with the same deduplication ID in the same queue will throw a `DBOSQueueDuplicatedException`. For example, this is useful if you only want to have one workflow active at a time per user—set the deduplication ID to the user's ID. **Example syntax:** ```java @Workflow public String taskWorkflow(String task) { // Process the task... return "completed"; } public void example(DBOS dbos, Example proxy, String task, String userID) throws Exception { try { // Use user ID for deduplication WorkflowHandle handle = dbos.startWorkflow( () -> proxy.taskWorkflow(task), new StartWorkflowOptions(QueueName.of("example-queue")).withDeduplicationId(userID) ); String result = handle.getResult(); System.out.println("Workflow completed: " + result); } catch (DBOSQueueDuplicatedException e) { // A workflow with this deduplication ID is already enqueued or running } } ``` #### Priority You can set a priority for an enqueued workflow with `withPriority` when calling `startWorkflow`. Workflows with the same priority are dequeued in **FIFO (first in, first out)** order. Priority values can range from `0` to `2,147,483,647`, where **a low number indicates a higher priority**. A negative priority throws `IllegalArgumentException`. Priority is enabled on every queue; no extra configuration is needed. :::tip Workflows without assigned priorities have priority `0`, the highest priority. ::: **Example syntax:** ```java @Workflow public String taskWorkflow(String task) { // Process the task... return "completed"; } public void example(DBOS dbos, Example proxy, String task, int priority) throws Exception { WorkflowHandle handle = dbos.startWorkflow( () -> proxy.taskWorkflow(task), new StartWorkflowOptions(QueueName.of("example-queue")).withPriority(priority) ); String result = handle.getResult(); System.out.println("Workflow completed: " + result); } ``` #### Explicit Queue Listening By default, a process running DBOS listens to (dequeues workflows from) all queues owned by its application in its system database. However, sometimes you only want a process to listen to a specific list of queues. You can use `withListenQueues` in your `DBOSConfig` to explicitly tell a process running DBOS to only listen to a specific set of queues. Each entry is a queue name; names that don't match any queue at launch are deferred until a queue is registered with that name. This is particularly useful when managing heterogeneous workers, where specific tasks should execute on specific physical servers. For example, say you have a mix of CPU workers and GPU workers and you want CPU tasks to only execute on CPU workers and GPU tasks to only execute on GPU workers. You can create separate queues for CPU and GPU tasks and configure each type of worker to only listen to the appropriate queue: ```java String workerType = System.getenv("WORKER_TYPE"); // "cpu" or "gpu" DBOSConfig config = DBOSConfig.defaults("dbos-java-starter").withAppVersion("0.1.0"); if ("gpu".equals(workerType)) { // GPU workers will only dequeue and execute workflows from the GPU queue config = config.withListenQueues("gpu-queue"); } else if ("cpu".equals(workerType)) { // CPU workers will only dequeue and execute workflows from the CPU queue config = config.withListenQueues("cpu-queue"); } DBOS dbos = new DBOS(config); // ... register workflows ... dbos.launch(); dbos.registerQueue("cpu-queue", QueueOptions.empty()); dbos.registerQueue("gpu-queue", QueueOptions.empty()); ``` Note that `withListenQueues` only controls what workflows are dequeued, not what workflows can be enqueued, so you can freely enqueue tasks onto the GPU queue from a CPU worker for execution on a GPU worker, and vice versa. ### DBOSClient ```java DBOSClient(String url, String user, String password) DBOSClient(String url, String user, String password, String schema) DBOSClient(String url, String user, String password, String schema, DBOSSerializer serializer) DBOSClient(String url, String user, String password, String schema, DBOSSerializer serializer, boolean useListenNotify) DBOSClient(String url, String user, String password, String schema, DBOSSerializer serializer, boolean useListenNotify, String applicationName) DBOSClient(DataSource dataSource) DBOSClient(DataSource dataSource, String schema) DBOSClient(DataSource dataSource, String schema, DBOSSerializer serializer) DBOSClient(DataSource dataSource, String schema, DBOSSerializer serializer, String applicationName) DBOSClient(DataSource dataSource, String schema, DBOSSerializer serializer, boolean useListenNotify) DBOSClient(DataSource dataSource, String schema, DBOSSerializer serializer, boolean useListenNotify, String applicationName) ``` Construct the DBOSClient. A client never migrates the system database; it throws if the schema is older than this SDK requires, so launch (and migrate) your DBOS application first. **Parameters:** - **url**: The JDBC URL for your system database. - **user**: Your PostgreSQL username or role. - **password**: The password for your PostgreSQL user or role. - **schema**: The schema DBOS system tables are stored in. Defaults to `dbos`. - **dataSource**: Provide an existing `DataSource` instead of connection URL/credentials. - **serializer**: A custom serializer for workflow inputs/outputs. Must match the serializer used by the DBOS application. - **useListenNotify**: If true, `getEvent` and `readStream` are woken by PostgreSQL notifications instead of polling. Defaults to false. - **applicationName**: The application this client acts for, when several applications share one system database. Without one, the client sees every application's workflows and queues. ### Workflow Interaction Methods #### enqueueWorkflow ```java WorkflowHandle enqueueWorkflow( EnqueueOptions options, Object[] args) ``` Enqueue a workflow and return a handle to it. **Parameters:** - **options**: Configuration for the enqueued workflow, as defined below. - **args**: An array of the workflow's arguments. These will be serialized and passed into the workflow when it is dequeued. **Example Syntax:** This code enqueues workflow `exampleWorkflow` in class `com.example.ExampleImpl` on queue `example-queue` with arguments `argumentOne` and `argumentTwo`. ```java var client = new DBOSClient(dbUrl, dbUser, dbPassword); var options = new EnqueueOptions( "exampleWorkflow", "com.example.ExampleImpl", QueueName.of("example-queue")); var handle = client.enqueueWorkflow(options, new Object[]{"argumentOne", "argumentTwo"}); ``` ##### EnqueueOptions `EnqueueOptions` is a with-based configuration record for parameterizing `client.enqueueWorkflow` and `dbos.enqueueWorkflow`. **Constructors:** ```java public EnqueueOptions(String workflowName, QueueName queue) public EnqueueOptions(String workflowName, String className, QueueName queue) public EnqueueOptions(String workflowName, String className, String instanceName, QueueName queue) ``` The constructors fix what to run and where: the workflow name, optionally its class and named instance, and the queue as a `QueueName` (for example `QueueName.of("my-queue")`). There is no `withClassName` or `withInstanceName`. A Java workflow is identified by its class, so always pass the fully qualified name of the class that implements it (or its `@WorkflowClassName` value) when enqueuing a Java workflow. Omit it only for a workflow that isn't registered on a class, such as a Python workflow function. **Methods:** - **`withWorkflowId(String workflowId)`**: Specify the idempotency ID to assign to the enqueued workflow. - **`withAppVersion(String appVersion)`**: The version of your application that should process this workflow. If not set, only an executor on the owning application's latest registered version dequeues it. - **`withTimeout(Duration timeout)`** / **`withTimeout(Timeout timeout)`** / **`withNoTimeout()`**: Set a timeout for the enqueued workflow. Does not begin until the workflow is dequeued and starts execution. Inside a workflow, an unset timeout inherits the caller's; `withNoTimeout()` declines it. - **`withDeadline(Instant deadline)`**: Set an absolute deadline for the enqueued workflow. - **`withDelay(Duration delay)`**: Delay the start of the workflow by the specified duration after it is dequeued. - **`withDeduplicationId(String deduplicationId)`**: At any given time, only one workflow with a specific deduplication ID can be enqueued in the specified queue. - **`withPriority(Integer priority)`**: Priority values range from `0` to `2,147,483,647`; lower numbers run first, and a negative priority throws. - **`withQueuePartitionKey(String key)`**: Partition key, for partitioned queues. - **`withSerialization(SerializationStrategy serialization)`**: Serialization format for the arguments, for example `SerializationStrategy.PORTABLE` to enqueue a workflow written in another language. Named arguments (`enqueueWorkflow(options, positionalArgs, namedArgs)`) require `PORTABLE`. - **`withApplicationName(String applicationName)`**: Enqueue the workflow for another application sharing the system database. ### Classes and Instances You can use multiple instances of the same class containing workflow methods, but if you do, they must be named at the time they are registered with `registerProxy`. This name allows workflow recovery to be directed to the correct instance. ```java T registerProxy(Class interfaceClass, T implementation, String instanceName) ``` ### Workflow Handles Starting a workflow or retrieving it produces a WorkflowHandle for interacting with the workflow. ```java WorkflowHandle retrieveWorkflow(String workflowId) ``` Retrieve the handle of a workflow. **Parameters**: - **workflowId**: The ID of the workflow whose handle to retrieve. ```java public interface WorkflowHandle { String workflowId(); T getResult() throws E; WorkflowStatus getStatus(); } ``` WorkflowHandle provides methods to interact with a running or completed workflow. The type parameters `T` and `E` represents the expected return type of the workflow and the checked exceptions it may throw. Handles can be used to wait for workflow completion, check status, and retrieve results. ##### WorkflowHandle.getResult ```java T getResult() throws E; ``` Wait for the workflow to complete and return its result. ##### WorkflowHandle.getStatus ```java WorkflowStatus getStatus(); ``` Retrieve the WorkflowStatus of the workflow. ##### WorkflowHandle.workflowId ```java String workflowId(); ``` Return the ID of the workflow underlying this handle. #### Workflow Status Some workflow introspection and management methods return a `WorkflowStatus`. This object has the following definition: ```java public record WorkflowStatus( String workflowId, // The workflow ID WorkflowState status, // PENDING, ENQUEUED, DELAYED, SUCCESS, ERROR, CANCELLED, or MAX_RECOVERY_ATTEMPTS_EXCEEDED String workflowName, // The workflow function name String className, // The class containing the workflow String instanceName, // The named class instance, if any String authenticatedUser, // The authenticated user who initiated the workflow String assumedRole, // The assumed role for the workflow execution List authenticatedRoles, // Roles authenticated for the workflow Object[] input, // The deserialized workflow input Object output, // The workflow's output, if any ErrorResult error, // The error the workflow threw, if any String executorId, // The ID of the executor that most recently ran this workflow Instant createdAt, // When the workflow was created Instant updatedAt, // When the workflow status was last updated String appVersion, // The application version on which this workflow was started String appId, // The application identifier Integer recoveryAttempts, // The number of times this workflow has been started String queueName, // If enqueued, on which queue Duration timeout, // The workflow timeout duration, if any Instant deadline, // The absolute deadline, if any Instant startedAt, // When the workflow started executing (after dequeue) String deduplicationId, // The deduplication ID, if any Integer priority, // The queue priority, if any String queuePartitionKey, // The queue partition key, if any String forkedFrom, // The workflow ID this was forked from, if any String parentWorkflowId, // The parent workflow ID if this is a child workflow Boolean wasForkedFrom, // Whether another workflow was forked from this one Instant delayUntil, // Time until which the workflow is delayed Instant completedAt, // When the workflow reached a terminal state, if it has String serialization, // Serialization format used for inputs/outputs Map attributes, // Custom key-value attributes attached at creation, if any String scheduleName, // The schedule that started this workflow, if any String applicationName // The application that owns this workflow, or null if unclaimed ) ``` ### DBOS Variables #### workflowId ```java static String workflowId() ``` Retrieve the ID of the current workflow. Returns `null` if not called from a workflow or step. #### stepId ```java static Integer stepId() ``` Returns the unique ID of the current step within its workflow. Returns `null` if not called from a step. #### inWorkflow ```java static boolean inWorkflow(); ``` Return `true` if the current calling context is executing a workflow, or `false` otherwise. #### inStep ```java static boolean inStep(); ``` Return `true` if the current calling context is executing a workflow step, or `false` otherwise. ### Workflow Management Methods #### listWorkflows ```java List listWorkflows(ListWorkflowsInput input) ``` Retrieve a list of WorkflowStatus of all workflows matching specified criteria. ##### ListWorkflowsInput `ListWorkflowsInput` is a with-based configuration record for filtering and customizing workflow queries. All fields are optional. **`with` Methods** (most accept a single value or a `List`): - `withWorkflowIds(String id)` / `withWorkflowIds(List)` — filter by workflow ID(s) - `withWorkflowName(String name)` / `withWorkflowName(List)` — filter by workflow function name - `withClassName(String className)` — filter by class name - `withInstanceName(String instanceName)` — filter by instance name - `withAuthenticatedUser(String user)` / `withAuthenticatedUser(List)` — filter by authenticated user - `withStatus(WorkflowState status)` / `withStatus(List)` — filter by status (`ENQUEUED`, `PENDING`, `SUCCESS`, `ERROR`, `CANCELLED`, `DELAYED`, `MAX_RECOVERY_ATTEMPTS_EXCEEDED`) - `withStartTime(Instant startTime)` — workflows created after this time - `withEndTime(Instant endTime)` — workflows created before this time - `withApplicationVersion(String version)` / `withApplicationVersion(List)` — filter by app version - `withLimit(Integer limit)` — max results to return - `withOffset(Integer offset)` — skip this many results (for pagination) - `withSortDesc(Boolean sortDesc)` — sort by creation time descending (true) or ascending (false) - `withExecutorIds(String id)` / `withExecutorIds(List)` — filter by executor process - `withQueueName(String name)` / `withQueueName(List)` — filter by queue - `withWorkflowIdPrefix(String prefix)` / `withWorkflowIdPrefix(List)` — filter by ID prefix - `withQueuesOnly(Boolean queuesOnly)` — only return enqueued workflows - `withLoadInput(Boolean value)` — whether to load workflow inputs (default: true) - `withLoadOutput(Boolean value)` — whether to load workflow outputs (default: true) - `withForkedFrom(String id)` / `withForkedFrom(List)` — filter to workflows forked from these IDs - `withParentWorkflowId(String id)` / `withParentWorkflowId(List)` — filter to child workflows of these parents - `withWasForkedFrom(Boolean value)` — filter to workflows that have been forked from - `withHasParent(Boolean value)` — filter to child workflows #### listWorkflowSteps ```java List listWorkflowSteps(String workflowId) List listWorkflowSteps(String workflowId, Integer limit, Integer offset) ``` Retrieve the execution steps of a workflow (with optional pagination). This is a list of `StepInfo` objects, with the following structure: ```java StepInfo( int functionId, // Sequential step ID within the workflow String functionName, // Name of the step function Object output, // Output returned by the step, if any ErrorResult error, // Error returned by the step, if any String childWorkflowId,// If the step starts a child workflow, its ID Instant startedAt, // When the step started Instant completedAt, // When the step completed String serialization, // Serialization format used for the step's output String applicationName // The application that ran the step ) ``` #### cancelWorkflow ```java void cancelWorkflow(String workflowId) void cancelWorkflows(List workflowIds) ``` Cancel one or more workflows. Sets status to `CANCELLED`, removes from queue, and preempts execution at the next step boundary. #### resumeWorkflow ```java WorkflowHandle resumeWorkflow(String workflowId) WorkflowHandle resumeWorkflow(String workflowId, String queueName) List> resumeWorkflows(List workflowIds) ``` Resume one or more workflows from their last completed step. Optionally re-enqueue on a queue instead of starting immediately. #### forkWorkflow ```java WorkflowHandle forkWorkflow(String workflowId, int startStep) WorkflowHandle forkWorkflow(String workflowId, int startStep, ForkOptions options) ``` ```java public record ForkOptions( String forkedWorkflowId, String applicationVersion, Duration timeout, String queueName, String queuePartitionKey ) { ForkOptions withForkedWorkflowId(String forkedWorkflowId); ForkOptions withApplicationVersion(String applicationVersion); ForkOptions withTimeout(Duration timeout); ForkOptions withTimeout(long value, TimeUnit unit); ForkOptions withQueue(QueueName queue); ForkOptions withQueue(String queueName); ForkOptions withQueuePartitionKey(String queuePartitionKey); } ``` Start a new execution of a workflow from a specific step. Steps before `startStep` are not re-executed. **Parameters:** - **workflowId**: The ID of the workflow to fork - **startStep**: The step from which to fork the workflow - **options**: - **forkedWorkflowId**: Workflow ID for the forked workflow (UUID if not provided) - **applicationVersion**: App version for the forked workflow (inherited if not provided) - **timeout**: A timeout for the forked workflow (`null` for no timeout) - **queueName**: Enqueue the forked workflow on this queue instead of starting immediately - **queuePartitionKey**: Partition key for partitioned queues ### Configuring DBOS Configure and create a DBOS instance. **DBOSConfig** `DBOSConfig` is a with-based configuration record for configuring DBOS. The application name, database URL, database user, and database password are required. **Constructor:** ```java DBOSConfig.defaults(String appName) ``` Create a DBOSConfig object. This configuration can be adjusted by using `with` methods that produce new configurations. **With Methods:** - **`withAppName(String appName)`**: Your application's name. Required. - **`withDatabaseUrl(String databaseUrl)`**: The JDBC URL for your system database. A valid JDBC URL is of the form `jdbc:postgresql://host:port/database`. - **`withDbUser(String dbUser)`**: Your PostgreSQL username or role. - **`withDbPassword(String dbPassword)`**: The password for your PostgreSQL user or role. - **`withDataSource(DataSource dataSource)`**: Provide an existing `DataSource` instead of URL/credentials. - **`withDatabaseSchema(String schema)`**: The schema for DBOS system tables. Defaults to `dbos`. - **`withMigrate(boolean enable)`**: If true, apply migrations to the system database on launch. Defaults to true. If false, launch only checks that the system database schema is new enough, and throws if it isn't; migrate it out-of-band with `dbosctl sysdb migrate`. - **`withConductorKey(String key)`**: An API key for DBOS Conductor. If provided, the application is connected to Conductor. - **`withConductorDomain(String domain)`**: A custom hostname for DBOS Conductor, for self-hosted Conductor. - **`withAppVersion(String appVersion)`**: The code version for this application and its workflows. - **`withExecutorId(String executorId)`**: A unique identifier for this process instance. - **`withEnablePatching(boolean enable)`**: Enable workflow patching support. - **`withListenQueues(String... queues)`** / **`withListenQueues(QueueName... queues)`**: Specify the queues this DBOS process should dequeue and execute workflows from. By default it dequeues from all queues. - **`withSerializer(DBOSSerializer serializer)`**: A custom serializer for the system database. - **`withSchedulerPollingInterval(Duration interval)`**: How often the scheduler polls for due scheduled workflows. - **`withUseListenNotify(boolean enable)`**: Use PostgreSQL LISTEN/NOTIFY to wake waiting workflows instead of polling. Defaults to true; set to false on databases that don't support it. - **`withNotificationCoalesceInterval(Duration interval)`**: How often batched workflow event and stream notifications are sent to other processes. Defaults to 10ms; must be at least 1ms. - **`withDatabasePollingConcurrency(Integer limit)`**: The maximum number of polling reads (from waiting on a result, `recv`, `getEvent`, or reading a stream) that may run against the system database at once. Defaults to half the connection pool (at least one); a non-positive value removes the cap. ```` --- ## DBOS Client `DBOSClient` provides a programmatic way to interact with your DBOS application from external code. ### DBOSClient ```java DBOSClient(String url, String user, String password) DBOSClient(String url, String user, String password, String schema) DBOSClient(String url, String user, String password, String schema, DBOSSerializer serializer) DBOSClient(String url, String user, String password, String schema, DBOSSerializer serializer, boolean useListenNotify) DBOSClient(String url, String user, String password, String schema, DBOSSerializer serializer, boolean useListenNotify, String applicationName) DBOSClient(DataSource dataSource) DBOSClient(DataSource dataSource, String schema) DBOSClient(DataSource dataSource, String schema, DBOSSerializer serializer) DBOSClient(DataSource dataSource, String schema, DBOSSerializer serializer, String applicationName) DBOSClient(DataSource dataSource, String schema, DBOSSerializer serializer, boolean useListenNotify) DBOSClient(DataSource dataSource, String schema, DBOSSerializer serializer, boolean useListenNotify, String applicationName) ``` Construct the DBOSClient. `DBOSClient` implements `AutoCloseable`; call `close()` to release its database resources. :::danger DBOSClient requires a PostgreSQL database. Providing a non-PostgreSQL `DataSource` will throw an exception. ::: The client never creates or migrates the system database. On construction, it checks that the system database schema has been migrated to a version compatible with this DBOS release, and throws `IllegalStateException` if the schema is missing or too old. Launch a DBOS application (or run [`dbosctl sysdb migrate`](../../conductor/reference/dbosctl.md#dbosctl-sysdb-migrate)) against the system database first. **Parameters:** - **url**: The JDBC URL for your system database. - **user**: Your PostgreSQL username or role. - **password**: The password for your PostgreSQL user or role. - **schema**: The schema the DBOS System Database tables are stored in. Defaults to `dbos` if not provided. - **dataSource**: System Database data source. A `HikariDataSource` is created if not provided. - **serializer**: A custom [serializer](./lifecycle.md#custom-serialization) for workflow inputs and outputs. Must match the serializer used by the DBOS application. - **useListenNotify**: If `true`, the client runs a listener thread so [`getEvent`](#getevent) and [`readStream`](#readstream) are woken by PostgreSQL `LISTEN`/`NOTIFY` notifications instead of polling the database. Defaults to `false` on the constructors that do not take it, because it costs a dedicated connection and thread that only those two methods benefit from. Leave it `false` if the system database was migrated with `LISTEN`/`NOTIFY` disabled. - **applicationName**: The application on whose behalf this client acts. Workflows the client enqueues, and queues and schedules it registers, are owned by that application, and the client's listing operations default to that application's rows. Always set this if multiple applications [share a system database](../../explanations/sharing-a-system-database.md). ##### Named and unnamed clients A client constructed with an `applicationName` is a **named** client: it acts as that application. A client constructed without one is an **unnamed** client: the workflows, queues, and schedules it creates are owned by no application (so every application sharing the system database treats them as its own), and its listing operations return every application's rows. Individual operations can still target a specific application, for example with [`EnqueueOptions.withApplicationName`](#enqueueoptions). ##### applicationName ```java String applicationName() ``` Return the application this client acts on behalf of, or `null` for an unnamed client. ### Workflow Interaction Methods #### enqueueWorkflow ```java WorkflowHandle enqueueWorkflow( EnqueueOptions options, Object[] args) WorkflowHandle enqueueWorkflow( EnqueueOptions options, Object[] positionalArgs, Map namedArgs) ``` Enqueue a workflow and return a handle to it. **Parameters:** - **options**: Configuration for the enqueued workflow, as defined below. - **args** / **positionalArgs**: An array of the workflow's arguments. These will be serialized and passed into the workflow when it is dequeued. - **namedArgs**: Named arguments, for targets that take them, such as a Python workflow with keyword arguments. Only portable serialization carries named arguments, so passing any requires `withSerialization(SerializationStrategy.PORTABLE)` on the options; otherwise the call throws `IllegalArgumentException`. **Example Syntax:** This code enqueues workflow `exampleWorkflow` in class `com.example.ExampleImpl` on queue `example-queue` with arguments `argumentOne` and `argumentTwo`. ```java var client = new DBOSClient(dbUrl, dbUser, dbPassword); var options = new EnqueueOptions( "exampleWorkflow", "com.example.ExampleImpl", QueueName.of("example-queue")); var handle = client.enqueueWorkflow(options, new Object[]{"argumentOne", "argumentTwo"}); ``` ##### EnqueueOptions `EnqueueOptions` (`dev.dbos.transact.EnqueueOptions`) is a with-based configuration record for parameterizing `client.enqueueWorkflow`. The same record is used by [`dbos.enqueueWorkflow`](./methods.md#enqueueworkflow) inside a DBOS application. The nested `DBOSClient.EnqueueOptions`, and the `DBOSClient` enqueue overloads that take it, are *(deprecated since 1.1)*. **Constructors:** ```java public EnqueueOptions(String workflowName, QueueName queue) public EnqueueOptions(String workflowName, String className, QueueName queue) public EnqueueOptions(String workflowName, String className, String instanceName, QueueName queue) ``` The constructors fix what to run and where: the workflow name, optionally the class that contains it and the [named instance](../tutorials/workflow-classes.md) to run it on, and the queue, as a [`QueueName`](./queues.md#queuename). A Java workflow is identified by its class, so always pass the fully qualified name of the class that implements it (or its `@WorkflowClassName` value) when enqueuing a Java workflow. Omit it only for a workflow that isn't registered on a class, such as a Python workflow function. The workflow name and queue must not be null or empty. **Methods:** - **`withWorkflowId(String workflowId)`**: Specify the idempotency ID to assign to the enqueued workflow. - **`withAppVersion(String appVersion)`**: The version of your application that should process this workflow. If left undefined, the workflow is enqueued without a version and is only dequeued by an executor running the owning application's latest registered version, which sets the version when it first dequeues it. - **`withTimeout(Duration timeout)`**, **`withTimeout(long value, TimeUnit unit)`**: Set an explicit timeout for the enqueued workflow. When the timeout expires, the workflow and all its children are cancelled. The timeout does not begin until the workflow is dequeued and starts execution. - **`withTimeout(Timeout timeout)`**, **`withNoTimeout()`**: Set the timeout as a [`Timeout`](./methods.md#timeout): explicit, none, or inherit. Inside a workflow, [`dbos.enqueueWorkflow`](./methods.md#enqueueworkflow) resolves it as `startWorkflow` does, so an unset timeout inherits the enqueuing workflow's and `withNoTimeout()` declines it. From a client there is nothing to inherit, so an unset or inherited timeout means no timeout. - **`withDeadline(Instant deadline)`**: Set a deadline for the enqueued workflow. If the workflow is executing when the deadline arrives, the workflow and all its children are cancelled. :::info An explicit timeout and a deadline cannot both be set. ::: - **`withDelay(Duration delay)`**: Delay the start of the workflow by the specified duration after it is dequeued. - **`withDeduplicationId(String deduplicationId)`**: At any given time, only one workflow with a specific deduplication ID can be enqueued in the specified queue. If a workflow with a deduplication ID is currently enqueued or actively executing (status `ENQUEUED`, `PENDING`, or `DELAYED`), subsequent workflow enqueue attempt with the same deduplication ID in the same queue will raise an exception. - **`withPriority(Integer priority)`**: The priority of the enqueued workflow in the specified queue. Workflows with the same priority are dequeued in FIFO (first in, first out) order. Priority values can range from `0` to `2,147,483,647`, where a low number indicates a higher priority. A negative priority throws `IllegalArgumentException`. Workflows without assigned priorities have priority `0`, the highest priority. - **`withSerialization(SerializationStrategy serialization)`**: Specify the [serialization strategy](./lifecycle.md#custom-serialization) for the workflow arguments. Options are `SerializationStrategy.DEFAULT`, `SerializationStrategy.PORTABLE`, or `SerializationStrategy.NATIVE`. - **`withQueuePartitionKey(String partitionKey)`**: Set a queue partition key for the workflow. Use if and only if the queue is partitioned, which it is when any per-partition limit (`partitionConcurrency`, `partitionWorkerConcurrency`, or `partitionRateLimit`) is set, or when it was registered with the deprecated `partitionQueue` flag. Per-partition limits apply to each partition key separately; queue-wide limits still apply to the queue as a whole, except on a queue registered with the deprecated `partitionQueue` flag, where they apply to each partition instead. See [Partitioning Queues](../tutorials/queue-tutorial.md#partitioning-queues). :::info - Partition keys are required when enqueueing to a partitioned queue. - Partition keys cannot be used with non-partitioned queues. - Partition keys and deduplication IDs cannot be used together. ::: - **`withAuthenticatedUser(String user)`**: Set the authenticated user for the enqueued workflow. Stored in the workflow record and accessible via `DBOS.authenticatedUser()` when the workflow executes. - **`withAssumedRole(String role)`**: Set the assumed role for the workflow. Accessible via `DBOS.assumedRole()`. - **`withAuthenticatedRoles(String... roles)`**: Set the list of authenticated roles. Accessible via `DBOS.authenticatedRoles()`. - **`withAuthentication(String user, String... roles)`**: Convenience method to set both `authenticatedUser` and `authenticatedRoles` in one call. - **`withAttributes(Map attributes)`**: Attach custom JSON-serializable key-value metadata to the workflow. Searchable via `ListWorkflowsInput.withAttributes(Map)`. - **`withApplicationName(String applicationName)`**: The application that owns the enqueued workflow. Only executors running that application dequeue and run it, so this is how one application enqueues work for another application sharing its system database. Defaults to the client's own [`applicationName`](#applicationname); on an unnamed client, the workflow is owned by no application. See [Sharing a System Database](../../explanations/sharing-a-system-database.md). #### enqueuePortableWorkflow *(deprecated since 1.1)* `enqueuePortableWorkflow` only takes the deprecated `DBOSClient.EnqueueOptions`. To enqueue a workflow written in another language, set `withSerialization(SerializationStrategy.PORTABLE)` on [`EnqueueOptions`](#enqueueoptions) and call [`enqueueWorkflow`](#enqueueworkflow), passing named arguments if the target takes them. #### send ```java send(String destinationId, Object message, String topic, String idempotencyKey) send(String destinationId, Object message, String topic, String idempotencyKey, SendOptions options) ``` Similar to [`dbos.send`](./methods.md#send). The optional `SendOptions` parameter controls serialization and fork delivery; see [`SendOptions`](#sendoptions) below. #### sendBulk ```java void sendBulk(List messages) void sendBulk(List messages, SendOptions options) ``` Send multiple messages to workflows in a single batch. Each message is delivered to its destination workflow independently; messages need not share the same destination. **Parameters:** - **messages**: A list of [`SendMessage`](./methods.md#sendmessage) records describing each message to send. - **options**: Optional send options controlling serialization and fork delivery; see [`SendOptions`](#sendoptions) below. #### SendOptions ```java SendOptions.defaults() SendOptions.portable() ``` `SendOptions` controls serialization and fork delivery for [`send`](#send) and [`sendBulk`](#sendbulk). **Factory methods:** - **`SendOptions.defaults()`**: Uses the default serialization strategy. - **`SendOptions.portable()`**: Uses portable JSON serialization for cross-language interoperability. **Builder method:** - **`withSendToForks(boolean)`**: Returns a new `SendOptions` with the `sendToForks` flag set. If `true`, the message is also delivered to any forked copies of the destination workflow. #### getEvent ```java Optional getEvent(String targetId, String key, Duration timeout) ``` Similar to [`dbos.getEvent`](./methods.md#getevent). #### readStream ```java Iterator readStream(String workflowId, String key) ``` Similar to [`dbos.readStream`](./methods.md#readstream). Use this from external code that does not have access to a `DBOS` instance. If no workflow with the given ID exists, iterating throws `DBOSNonExistentWorkflowException`. #### findWorkflowIdByDeduplicationId ```java String findWorkflowIdByDeduplicationId(String queueName, String deduplicationId) ``` Return the ID of the active (`ENQUEUED`, `PENDING`, or `DELAYED`) workflow holding the given deduplication ID on the given queue, or `null` if there is none. #### findDeduplicationHolder ```java DeduplicationHolder findDeduplicationHolder(String queueName, String deduplicationId) public record DeduplicationHolder( String workflowId, String applicationName, String workflowName, String className, String instanceName, WorkflowState status, boolean isDebounced) ``` Like [`findWorkflowIdByDeduplicationId`](#findworkflowidbydeduplicationid), but also returns the application that owns the holding workflow, along with its name and status. Deduplication IDs are unique across all applications sharing a system database, so the holder may belong to another application. Returns `null` if no active workflow holds the deduplication ID. ### Workflow Management Methods #### retrieveWorkflow ```java WorkflowHandle retrieveWorkflow(String workflowId) ``` Similar to [`dbos.retrieveWorkflow`](./methods.md#retrieveworkflow). #### getWorkflowStatus ```java Optional getWorkflowStatus(String workflowId) ``` Retrieve the [`WorkflowStatus`](./methods.md#workflowstatus) of a workflow. #### listWorkflows ```java List listWorkflows(ListWorkflowsInput input) ``` Similar to [`dbos.listWorkflows`](./methods.md#listworkflows). #### listWorkflowSteps ```java List listWorkflowSteps(String workflowId) List listWorkflowSteps(String workflowId, Integer limit, Integer offset) ``` Similar to [`dbos.listWorkflowSteps`](./methods.md#listworkflowsteps). #### cancelWorkflow ```java void cancelWorkflow(String workflowId) void cancelWorkflow(String workflowId, boolean cancelChildren) void cancelWorkflows(List workflowIds) void cancelWorkflows(List workflowIds, boolean cancelChildren) ``` Similar to [`dbos.cancelWorkflow`](./methods.md#cancelworkflow). When `cancelChildren` is `true`, also recursively cancels all descendant workflows. #### resumeWorkflow ```java WorkflowHandle resumeWorkflow(String workflowId) WorkflowHandle resumeWorkflow(String workflowId, String queueName) List> resumeWorkflows(List workflowIds) List> resumeWorkflows(List workflowIds, String queueName) ``` Similar to [`dbos.resumeWorkflow`](./methods.md#resumeworkflow). #### deleteWorkflow ```java void deleteWorkflow(String workflowId) void deleteWorkflow(String workflowId, boolean deleteChildren) void deleteWorkflows(List workflowIds) void deleteWorkflows(List workflowIds, boolean deleteChildren) ``` Similar to [`dbos.deleteWorkflow`](./methods.md#deleteworkflow). #### forkWorkflow ```java WorkflowHandle forkWorkflow( String originalWorkflowId, int startStep, ForkOptions options) ``` Similar to [`dbos.forkWorkflow`](./methods.md#forkworkflow). #### updateWorkflowAttributes ```java void updateWorkflowAttributes(String workflowId, Map attributes) ``` Similar to [`dbos.updateWorkflowAttributes`](./methods.md#updateworkflowattributes). #### setWorkflowDelay ```java void setWorkflowDelay(String workflowId, Duration delay) void setWorkflowDelay(String workflowId, Instant delayUntil) ``` Pause a workflow until a delay elapses or a specific time is reached. The workflow will resume from where it left off after the delay. **Parameters:** - **workflowId**: The ID of the workflow to delay. - **delay**: The duration to delay the workflow from now. - **delayUntil**: The absolute time until which to delay the workflow. ### Schedule Management Methods #### createSchedule ```java void createSchedule(WorkflowSchedule schedule) ``` Create a cron schedule. See [`WorkflowSchedule`](./methods.md#workflowschedule) for the schedule configuration. The schedule is owned by, and its workflows run by, the application named by [`WorkflowSchedule.withApplicationName`](./methods.md#workflowschedule), defaulting to the client's own [`applicationName`](#applicationname). A schedule created by an unnamed client without an application name is owned by no application, so every application sharing the system database runs it. #### getSchedule ```java Optional getSchedule(String name) ``` Get a schedule by name. Returns empty if the schedule does not exist. #### listSchedules ```java List listSchedules( List status, List workflowName, List namePrefix) List listSchedules( List status, List workflowName, List namePrefix, List applicationName) ``` List schedules with optional filters. Pass `null` for any parameter to skip that filter. **Parameters:** - **status**: Filter by [`ScheduleStatus`](./methods.md#workflowschedule). Pass `null` for no status filter. - **workflowName**: Filter by workflow name. Pass `null` for no workflow name filter. - **namePrefix**: Filter by schedule name prefix. Pass `null` for no prefix filter. - **applicationName**: List schedules owned by these applications, plus schedules no application owns. If `null` (or omitted), lists the client's own [`applicationName`](#applicationname)'s schedules; an unnamed client lists every application's schedules. Pass an empty list to list every application's schedules. #### deleteSchedule ```java void deleteSchedule(String name) ``` Delete a schedule by name. No-op if the schedule does not exist. #### pauseSchedule ```java void pauseSchedule(String name) ``` Pause a schedule. A paused schedule does not fire. #### resumeSchedule ```java void resumeSchedule(String name) ``` Resume a paused schedule so it begins firing again. #### applySchedules ```java void applySchedules(List schedules) void applySchedules(WorkflowSchedule... schedules) ``` Atomically create or replace a set of schedules. #### backfillSchedule ```java List> backfillSchedule( String scheduleName, Instant start, Instant end) ``` Enqueue all executions of a schedule that would have run between `start` (exclusive) and `end` (exclusive). **Parameters:** - **scheduleName**: Name of an existing schedule. - **start**: Start of the backfill window (exclusive). - **end**: End of the backfill window (exclusive). #### triggerSchedule ```java WorkflowHandle triggerSchedule(String scheduleName) ``` Immediately enqueue the scheduled workflow at the current time. **Parameters:** - **scheduleName**: Name of an existing schedule. ### Queue Management Methods `DBOSClient` can manage queues directly in the system database without a running DBOS executor. See [Queues & Concurrency](../tutorials/queue-tutorial.md) in the tutorial for usage examples. #### registerQueue ```java void registerQueue(String name, QueueOptions options) void registerQueue(String name, QueueOptions options, QueueConflictResolution onConflict) void registerQueue(String name, QueueOptions options, QueueConflictResolution onConflict, String applicationName) ``` Register or update a queue in the system database. The default conflict resolution is `ALWAYS_UPDATE`. The queue is owned by, and polled only by, the application named by `applicationName`, defaulting to the client's own [`applicationName`](#applicationname) (or no application, for an unnamed client). Registering a queue whose name is already owned by a different application throws [`DBOSApplicationNameConflictException`](./methods.md#dbosapplicationnameconflictexception). :::info `QueueConflictResolution.UPDATE_IF_LATEST_VERSION` is not supported for `DBOSClient` because clients are not associated with an application version. Use `ALWAYS_UPDATE` or `NEVER_UPDATE`. ::: #### updateQueue ```java void updateQueue(String name, QueueOptions options) ``` Update the configuration of an existing queue. Only fields set on `options` are modified; absent fields are left unchanged (see [`QueueOptions`](./queues.md#queueoptions) and [`Field`](./queues.md#fieldt)). #### findQueue ```java Optional findQueue(String name) ``` Look up a queue by name. Returns empty if no queue with that name exists. #### listQueues ```java List listQueues() List listQueues(List applicationName) ``` Return the queues registered in the system database. The queues listed are those owned by the given applications, plus queues no application owns. If `applicationName` is `null` (or omitted), lists the client's own [`applicationName`](#applicationname)'s queues; an unnamed client lists every application's queues. Pass an empty list to list every application's queues. #### deleteQueue ```java boolean deleteQueue(String name) ``` Delete a queue from the system database. Returns `true` if the queue was deleted, `false` if it did not exist. ### Application Version Methods #### listApplicationVersions ```java List listApplicationVersions() ``` List all registered application versions, ordered by timestamp descending. #### getLatestApplicationVersion ```java VersionInfo getLatestApplicationVersion() ``` Get the most recently promoted application version. #### setLatestApplicationVersion ```java void setLatestApplicationVersion(String versionName) void setLatestApplicationVersion(String versionName, String applicationName) ``` Promote an existing version to be the latest application version by updating its timestamp. The version must already exist. **Parameters:** - **versionName**: The name of the version to promote. - **applicationName**: The application to act as. Defaults to the client's own [`applicationName`](#applicationname). Promoting a version registered by a different application throws [`DBOSApplicationNameConflictException`](./methods.md#dbosapplicationnameconflictexception). Promoting a version owned by no application claims it for this application. ### Application Rename #### renameApplication ```java ApplicationRowCounts renameApplication(String oldName, String newName) ApplicationRowCounts renameApplication( String oldName, String newName, Integer batchSize, boolean adoptUnclaimedRows) public record ApplicationRowCounts( long queues, long schedules, long versions, long workflows, long steps) ``` Every workflow, step, queue, schedule, and application version is owned by the application that created it. After renaming an application, use this method (or the [`dbosctl sysdb rename-application`](../../conductor/reference/dbosctl.md#dbosctl-sysdb-rename-application) command) to transfer everything owned by the old name to the new name. Returns the number of rows transferred, by table. The operation is idempotent: if interrupted, running it again resumes where it left off. :::warning Stop the application being renamed before running this. A running application would race the rename, creating new work under its old name. ::: **Parameters:** - **oldName**: The application's previous name. If `null`, nothing is transferred except rows owned by no application, so `adoptUnclaimedRows` must be `true`. - **newName**: The application that ends up owning the rows. - **batchSize**: The number of completed workflows and steps transferred per transaction; queues, schedules, versions, and active workflows are transferred together in a single transaction. The two-argument overload uses 10,000. Pass `null` to transfer everything in a single transaction. - **adoptUnclaimedRows**: Also transfer rows owned by no application, such as rows created before upgrading to a DBOS version supporting application ownership. The two-argument overload passes `false`. ### Debouncing Workflows can be debounced from external code using `DebouncerClient`. #### DBOSClient.debouncer ```java DebouncerClient debouncer(String workflowName) ``` Create a `DebouncerClient` for the named workflow. Similar to [`dbos.debouncer()`](./methods.md#debouncer) but operates externally — no running DBOS executor is required on the caller's side. `DebouncerClient` is an immutable builder. Configure it with the following methods before calling `debounce`: - **`withClassName(String className)`**: The fully-qualified Java class name of the workflow implementation. **Required** — must be set before calling `debounce`. - **`withInstanceName(String instanceName)`**: The DBOS instance name of the target workflow implementation. - **`withDebounceTimeout(Duration debounceTimeout)`**: Set an absolute cap on how long the debouncer may keep absorbing calls for a single key. - **`withQueue(QueueName queue)`** / **`withQueue(String queueName)`**: Enqueue the user workflow on the specified queue when the debounce period elapses. `withQueue(Queue queue)` is *(deprecated since 1.1)*. - **`withTimeout(Duration timeout)`**: Set a timeout for the user workflow. - **`withAppVersion(String appVersion)`**: Target a specific application version. - **`withPriority(Integer priority)`**: Set the priority for the user workflow. A priority requires a queue: if a priority is set without `withQueue`, `debounce` throws `IllegalArgumentException`. - **`withAttributes(Map attributes)`**: Attach custom JSON-serializable key-value metadata to the user workflow. - **`withSerialization(SerializationStrategy serialization)`**: The [serialization strategy](./methods.md#serialization-strategy) for the user workflow's arguments. It should match the strategy the workflow is registered with. - **`withDeduplicationId(String deduplicationId)`** *(deprecated since 1.1)*: Set a deduplication ID forwarded to the user workflow. This will be ignored from the next release, where the debouncer sets the deduplication ID itself, and removed in 2.0. #### DebouncerClient.debounce ```java WorkflowHandle debounce(String debounceKey, Duration debouncePeriod, Object... args) ``` Similar to [`debouncer.debounce`](./methods.md#debouncerdebounce) but takes positional arguments directly instead of a workflow proxy lambda. **Parameters:** - **debounceKey**: A key used to group workflow executions that will be debounced together. - **debouncePeriod**: Inactivity window before the user workflow runs; each call resets it. - **args**: Positional arguments to pass to the workflow when it runs. **Example Syntax:** ```java var client = new DBOSClient(url, user, password); var debouncer = client.debouncer("processInput") .withClassName(MyServiceImpl.class.getName()) .withDebounceTimeout(Duration.ofMinutes(5)); // Each time a user submits input, debounce the processInput workflow. // The workflow will run 60 seconds after the user stops submitting. WorkflowHandle handle = debouncer.debounce( userId, Duration.ofSeconds(60), userInput); String result = handle.getResult(); ``` --- ## Kotlin Extensions DBOS ships Kotlin extension functions on the `DBOS` class in the `dev.dbos.transact` package. They place the lambda argument last so Kotlin's trailing lambda syntax works naturally. All extensions are `@JvmSynthetic` and invisible to Java callers. ### Extensions #### runStep ```kotlin fun DBOS.runStep(options: StepOptions, block: () -> T): T fun DBOS.runStep(name: String, block: () -> T): T ``` Kotlin-friendly alternatives to [`dbos.runStep`](./workflows-steps.md#runstep) that place the lambda last. **Parameters:** - **options**: A [`StepOptions`](./workflows-steps.md#stepoptions) object controlling the step name and retry policy. - **name**: The step name as a plain string (shorthand for `StepOptions(name)`). - **block**: The step body. Must be deterministic across retries; side effects and external calls belong here rather than in the workflow body. **Example:** ```kotlin // with a name val result = dbos.runStep("fetchData") { httpClient.get(url) } // with options (e.g. retries) val result = dbos.runStep(StepOptions("fetchData").withMaxAttempts(3)) { httpClient.get(url) } ``` #### startWorkflow ```kotlin fun DBOS.startWorkflow(options: StartWorkflowOptions?, block: () -> T): WorkflowHandle ``` Kotlin-friendly alternative to [`dbos.startWorkflow`](./workflows-steps.md#startworkflow) that places the lambda last. Pass `null` for `options` to use defaults. :::note A zero-argument `startWorkflow { }` extension is not provided because Kotlin's SAM conversion causes the Java member to win before any extension is considered for that signature. Always supply `StartWorkflowOptions()` (or `null`) as the first argument. ::: **Parameters:** - **options**: [`StartWorkflowOptions`](./workflows-steps.md#startworkflowoptions) controlling workflow ID, queue, timeout, etc. Pass `null` to use defaults. - **block**: A call to a registered workflow proxy method. **Returns:** A [`WorkflowHandle`](./workflows-steps.md#workflowhandle) for retrieving the workflow result. **Example:** ```kotlin val handle = dbos.startWorkflow(StartWorkflowOptions()) { orderProxy.processOrder("order-123") } val result = handle.result ``` --- ## DBOS Lifecycle You create a `DBOS` instance exactly once in a program's lifetime, register your workflows, then launch it. Here, we document the constructor, configuration, and lifecycle methods. #### DBOSConfig `DBOSConfig` is a with-based configuration record for configuring DBOS. The application name, and a system db datasource - specified either as database URL, user and password or as a preconstructed `DataSource` are required. :::danger DBOS requires a PostgreSQL-compatible database. PostgreSQL and CockroachDB are supported. Creating a `DBOSConfig` with any other `DataSource` will throw an exception. ::: **Constructors:** ```java DBOSConfig.defaults(String appName) DBOSConfig.defaultsFromEnv(String appName) ``` Create a DBOSConfig object. The `defaults` static method only sets the application name and sets all other config fields to their default values. `appName` must follow the naming rule described under `withAppName` below. The `defaultsFromEnv` static method reads database connection information from environment variables. - **`DBOS_SYSTEM_JDBC_URL`**: the JDBC URL for your system database - **`PGUSER`**: the PostgreSQL username or role. Defaults to `postgres` if `PGUSER` environment variable is missing or empty. - **`PGPASSWORD`**: The password for your PostgreSQL user or role. This configuration can be adjusted by using `with` methods that produce new configurations. **With Methods:** - **`withAppName(String appName)`**: Your application's name. Required. It must be between 3 and 256 characters long and contain only lowercase letters, numbers, dashes, and underscores. An application connecting to [Conductor](../../conductor/overview.md) (with a Conductor key set) or running on DBOS Cloud fails to launch with a name outside that rule, because Conductor refuses to register it: `dbos.launch()` throws `IllegalArgumentException`. A self-hosted application logs a warning and launches. Multiple applications (potentially in different languages) may [share a system database](../../explanations/sharing-a-system-database.md), in which case each must have a distinct name: the name identifies which application owns each workflow, queue, schedule, and application version, and applications only run their own workflows. If you rename an application, transfer ownership of its data with [`DBOSClient.renameApplication`](./client.md#renameapplication) or [`dbosctl sysdb rename-application`](../../conductor/reference/dbosctl.md#dbosctl-sysdb-rename-application). - **`withAppVersion(String appVersion)`**: The code version for this application and its workflows. We recommend always setting it; if it is not set, DBOS computes a version from a hash of your workflow methods, which is only a fallback. Workflow versioning is documented [here](../tutorials/upgrading-workflows.md#versioning). - **`withDatabaseUrl(String databaseUrl)`**: The JDBC URL for your system database. A valid JDBC URL is of the form `jdbc:postgresql://host:port/database`. Required unless valid DataSource is provided. - **`withDbUser(String dbUser)`**: Your PostgreSQL username or role. Required unless valid DataSource is provided. - **`withDbPassword(String dbPassword)`**: The password for your PostgreSQL user or role. Required unless valid DataSource is provided. - **`withDataSource(DataSource v)`**: Instead of providing DBOS with the JDBC URL, username and password, you can provide a configured DataSource for DBOS to use. DBOS uses `HikariDataSource` if a data source is not provided. :::warning Using a data source that doesn't support connection pooling like `PGSimpleDataSource` is not recommended. ::: - **`withDatabaseSchema(String schema)`**: PostgreSQL database schema for system database tables. Defaults to `dbos`. - **`withMigrate(boolean enable)`**: If true, attempt to apply migrations to the system database. Defaults to true. - **`withConductorKey(String key)`**: An API key for [DBOS Conductor](../../conductor/overview.md). If provided, application is connected to Conductor. API keys can be created from the [DBOS Console](https://console.dbos.dev). - **`withConductorDomain(String domain)`**: The domain of the DBOS Conductor instance to connect to. Only needed when using a self-hosted Conductor. - **`withConductorExecutorMetadata(Map metadata)`**: Arbitrary key-value metadata attached to this executor and reported to Conductor. - **`withAdminServer(boolean enable)`** *(deprecated since 0.9)*: Whether to run the built-in HTTP admin server. Use [DBOS Conductor](../../conductor/overview.md) for remote administration instead. - **`enableAdminServer()`** / **`disableAdminServer()`** *(deprecated since 0.9)*: Convenience methods equivalent to `withAdminServer(true)` and `withAdminServer(false)`. - **`withAdminServerPort(int port)`** *(deprecated since 0.9)*: The port on which the admin server runs. Defaults to 3001. - **`withExecutorId(String executorId)`**: A unique process ID used to identify this application instance in distributed environments. If using DBOS Conductor or Cloud, this is set automatically. - **`withEnablePatching(boolean patchEnabled)`**: Enable workflow patching. Defaults to false. - **`withEnablePatching()`** / **`withDisablePatching()`**: Convenience methods equivalent to `withEnablePatching(true)` and `withEnablePatching(false)`. - **`withListenQueue(String queueName)`** / **`withListenQueue(QueueName queue)`**: Add a single queue to the set of queues this DBOS process listens to. [`QueueName`](./queues.md#queuename) is a typed wrapper for a queue name. - **`withListenQueues(String... queues)`** / **`withListenQueues(QueueName... queues)`**: Add multiple queues this DBOS process should dequeue and execute workflows from. Defaults to dequeuing from all registered queues. - **`withListenQueue(Queue queue)`** / **`withListenQueues(Queue... queues)`** *(deprecated since 1.1)*: Equivalent to the `QueueName` overloads, taking the name of each `Queue`. - **`withSchedulerPollingInterval(Duration interval)`**: How frequently the scheduler polls the database for new scheduled workflow firings. Defaults to 30 seconds. - **`withUseListenNotify(boolean enable)`**: Whether to use PostgreSQL `LISTEN`/`NOTIFY` for real-time event delivery (e.g. `recv`, `getEvent`). Defaults to `true`. Automatically set to `false` when CockroachDB is detected, since CockroachDB does not support `LISTEN`/`NOTIFY`. Set this to `false` explicitly if your PostgreSQL configuration does not support it. - **`withSerializer(DBOSSerializer serializer)`**: A custom serializer for the system database. See the [custom serialization section](#custom-serialization) for details. - **`withNotificationCoalesceInterval(Duration interval)`**: Interval at which DBOS batches and sends the notifications that wake processes waiting on [events](./methods.md#getevent) and [streams](./methods.md#readstream) written by this process. Rather than one notifying transaction per write, pending wake-ups are sent once per interval. A shorter interval lowers wake-up latency in other processes; a longer one lowers the rate of notifying commits, each of which takes a global lock in PostgreSQL. Waiters in the same process are woken immediately either way. Defaults to 10 milliseconds; the minimum is 1 millisecond. - **`withDatabasePollingConcurrency(Integer limit)`**: The maximum number of database-backed polling reads from wait operations (such as awaiting a workflow result, [`recv`](./methods.md#recv), [`getEvent`](./methods.md#getevent), and [`readStream`](./methods.md#readstream)) that may run concurrently against the system database connection pool. This prevents many waiters from checking out every connection in the pool and starving control-plane operations such as enqueue and dequeue, status writes, recovery, and cancellation. Defaults to half the pool size (minimum 1); the pool DBOS creates has 10 connections. Set to a non-positive value to remove the limit. #### Cloud Environment Variables When deploying to DBOS Cloud, several environment variables are automatically set and override `DBOSConfig` values: | Variable | Description | |----------|-------------| | `DBOS__CLOUD` | Set to `true` by DBOS Cloud. Enables cloud mode: `DBOS_APP_NAME` becomes required and the admin server is forced to port 3001. | | `DBOS_APP_NAME` | Overrides `DBOSConfig.appName()`. Required when `DBOS__CLOUD=true`; `launch()` throws if absent. | | `DBOS__CONDUCTOR_URL` | URL of DBOS Conductor. Overrides `withConductorDomain(...)`. | | `DBOS__CONDUCTOR_APP_NAME` | Application name used to identify this executor with Conductor. | | `DBOS__CONDUCTOR_KEY` | API key for DBOS Conductor. Overrides `withConductorKey(...)`. Set by the cloud platform; avoids putting credentials in `DBOSConfig`. | | `DBOS__VMID` | The executor ID of this process. Overrides `withExecutorId(...)` when `DBOS__CLOUD=true`. | These variables take precedence over any values set in `DBOSConfig`. In local development you do not need to set them. #### DBOS Constructor ```java new DBOS(DBOSConfig config) ``` Create and configure a DBOS instance. #### DBOS.version ```java static String version() ``` Return the DBOS library version string. #### registerQueue / registerQueues *(deprecated since 1.1)* {#registerqueue--registerqueues} ```java void registerQueue(Queue queue) void registerQueues(Queue... queues) ``` Register one or more in-memory queues before `dbos.launch()`. In-memory queues are deprecated for removal: register database-backed queues after launch with [`dbos.registerQueue(String, QueueOptions)`](./queues.md#dbosregisterqueue) instead. #### dbos.launch ```java void launch() ``` Launch DBOS, initializing database connections and beginning workflow recovery and queue processing. This should be called after all workflows are registered; queues are registered after launch. Launch-time recovery returns this executor's `PENDING` workflows from the current application version to their queues (workflows that were not enqueued go to an internal queue), from which they are dequeued and run again; it is skipped when a Conductor key is configured or on DBOS Cloud, where Conductor manages recovery. **You should not call a DBOS workflow until after DBOS is launched.** #### dbos.shutdown ```java void shutdown() ``` Shut down the DBOS instance, releasing database connections and stopping workflow processing. `DBOS` also implements `AutoCloseable`, so it can be used in a try-with-resources block. This may be useful for testing DBOS applications. #### getQueue *(deprecated since 1.1)* {#getqueue} ```java Optional getQueue(String queueName) ``` Return the in-memory `Queue` registered before launch with the given name, or empty if no such queue is registered. Must be called after `dbos.launch()`. Deprecated for removal: use [`dbos.findQueue`](./queues.md#dbosfindqueue), which reads the queue from the system database. #### Custom Serialization DBOS must serialize data such as workflow inputs and outputs and step return values to store it in the system database. By default, data is serialized with Jackson (format name `java_jackson`), but you can optionally supply a custom serializer through DBOS configuration. A custom serializer must implement the `DBOSSerializer` interface: ```java import dev.dbos.transact.json.DBOSSerializer; public interface DBOSSerializer { /** * Return a name for the serialization format. * This name is stored in the database to identify how data was serialized. */ String name(); /** Serialize a value to a string. */ String serialize(Object value); /** Deserialize a string back to a value, or return null if the input is null. */ Object deserialize(String text); /** Serialize a Throwable to a string. */ String serializeThrowable(Throwable throwable); /** Deserialize a string back to a Throwable, or return null if the input is null. */ Throwable deserializeThrowable(String text); } ``` The `name()` method must return a unique identifier for the serialization format. This name is stored alongside serialized values in the database so that the correct deserializer is used when reading data back. The `noHistoricalWrapper` parameter indicates whether the value is wrapped in an enclosing array. When `noHistoricalWrapper` is `true`, the value is a raw `Object[]` array and should be serialized directly. When it is `false`, the value is a single object. For example, here is a skeleton custom serializer: ```java import dev.dbos.transact.json.DBOSSerializer; public class MyCustomSerializer implements DBOSSerializer { @Override public String name() { return "my_custom"; } @Override public String stringify(Object value, boolean noHistoricalWrapper) { // Serialize value to a string } @Override public Object parse(String text, boolean noHistoricalWrapper) { if (text == null) return null; // Deserialize string back to a value } @Override public String stringifyThrowable(Throwable throwable) { if (throwable == null) return null; // Serialize throwable to a string } @Override public Throwable parseThrowable(String text) { if (text == null) return null; // Deserialize string back to a Throwable } } ``` Configure DBOS to use a custom serializer: ```java DBOSConfig config = DBOSConfig.defaultsFromEnv("myApp") .withAppVersion("0.1.0") .withSerializer(new MyCustomSerializer()); DBOS dbos = new DBOS(config); // register workflows and queues... dbos.launch(); ``` If you use a custom serializer in your DBOS application, you must provide the same serializer to any [`DBOSClient`](./client.md) that interacts with the application: ```java var client = new DBOSClient(dbUrl, dbUser, dbPassword, null, new MyCustomSerializer()); ``` --- ## DBOS Methods & Variables(Reference) ### Workflow Communication Methods #### getEvent ```java Optional getEvent(String workflowId, String key, Duration timeout) ``` Retrieve the latest value of an event published by the workflow identified by `workflowId` to the key `key`. If the event does not yet exist, wait for it to be published, returning empty if the wait times out. If the calling workflow is cancelled while waiting, throws `DBOSWorkflowCancelledException`. **Parameters:** - **workflowId**: The identifier of the workflow whose events to retrieve. - **key**: The key of the event to retrieve. - **timeout**: A timeout duration. If the wait times out, an empty Optional is returned. #### setEvent ```java void setEvent(String key, Object value) void setEvent(String key, Object value, SerializationStrategy serialization) ``` Create and associate with this workflow an event with key `key` and value `value`. If the event already exists, update its value. `setEvent` can only be called from within a workflow. **Parameters:** - **key**: The key of the event. - **value**: The value of the event. Must be serializable. - **serialization**: The [serialization strategy](#serialization-strategy) to use for this event. Defaults to `SerializationStrategy.DEFAULT`, which uses the serialization format recorded for the calling workflow. #### getAllEvents ```java Map getAllEvents(String workflowId) ``` Retrieve all events published by the workflow identified by `workflowId`, returned as a map of event key to deserialized value. **Parameters:** - **workflowId**: The identifier of the workflow whose events to retrieve. #### send ```java void send(String destinationId, Object message, String topic) void send(String destinationId, Object message, String topic, String idempotencyKey) void send(String destinationId, Object message, String topic, String idempotencyKey, SerializationStrategy serialization) ``` Send a message to the workflow identified by `destinationID`. Messages can optionally be associated with a topic. **Parameters:** - **destinationId**: The workflow to which to send the message. - **message**: The message to send. Must be serializable. - **topic**: A topic with which to associate the message. Messages are enqueued per-topic on the receiver. - **idempotencyKey**: If `dbos.send` is called from outside a workflow and an idempotency key is set, the message will only be sent once no matter how many times `dbos.send` is called with this key. - **serialization**: The [serialization strategy](#serialization-strategy) to use for this message. Defaults to `SerializationStrategy.DEFAULT`, which inside a workflow uses the serialization format recorded for the sending workflow (for example, a workflow started with portable serialization sends portable messages), and outside a workflow uses the application's configured serializer. #### sendBulk ```java void sendBulk(List messages) void sendBulk(List messages, boolean sendToForks) void sendBulk(List messages, boolean sendToForks, SerializationStrategy serialization) ``` Send multiple messages to workflows in a single batch. Each message is delivered to its destination workflow independently; messages need not share the same destination. **Parameters:** - **messages**: A list of [`SendMessage`](#sendmessage) records describing each message to send. - **sendToForks**: If `true`, also deliver each message to any forked copies of the destination workflow. Defaults to `false`. - **serialization**: The [serialization strategy](#serialization-strategy) to use for all messages in the batch. Defaults to `SerializationStrategy.DEFAULT`, which behaves as described for [`send`](#send). ##### SendMessage ```java new SendMessage(String destinationId, Object message) new SendMessage(String destinationId, Object message, String topic) new SendMessage(String destinationId, Object message, String topic, String idempotencyKey) ``` A record describing a single message in a [`sendBulk`](#sendbulk) batch. **Parameters:** - **destinationId**: The workflow to which to send the message. - **message**: The message to send. Must be serializable. - **topic**: A topic with which to associate the message. - **idempotencyKey**: Idempotency key for exactly-once delivery; a message with a given key is sent only once. #### recv ```java Optional recv(String topic, Duration timeout) ``` Receive and return a message sent to this workflow. Can only be called from within a workflow. Messages are dequeued first-in, first-out from a queue associated with the topic. Calls to `recv` wait for the next message in the queue, returning null if the wait times out. If the workflow is cancelled while waiting, `recv` throws `DBOSWorkflowCancelledException`. **Parameters:** - **topic**: A topic queue on which to wait. - **timeout**: A timeout duration. If the wait times out, return an empty Optional. #### sleep ```java void sleep(Duration sleepduration) ``` Sleep for the given duration. If called from within a workflow, this sleep is durable—it records its intended wake-up time in the database so if it is interrupted and recovers, it still wakes up at the intended time. If called from outside a workflow, or from within a step, it behaves like a regular `Thread.sleep`. **Parameters:** - **duration**: The duration to sleep. #### writeStream ```java void writeStream(String key, Object value) void writeStream(String key, Object value, SerializationStrategy serialization) ``` Append a value to a named stream owned by the current workflow. Must be called from within a workflow or step. Consumers read the stream in order via [`readStream`](#readstream). **Parameters:** - **key**: The stream name within this workflow. A workflow may have multiple independent streams identified by different keys. - **value**: A serializable value to write. - **serialization**: The [serialization strategy](./methods.md#serialization-strategy) to use. Defaults to `SerializationStrategy.DEFAULT`, which uses the serialization format recorded for the calling workflow. Use `SerializationStrategy.PORTABLE` for cross-language consumers. #### closeStream ```java void closeStream(String key) ``` Close a stream, signalling to consumers that no more values will be written. Must be called from within a workflow (not a step). After closing, [`readStream`](#readstream) iterators will drain any remaining values and then stop. **Parameters:** - **key**: The stream key to close. #### readStream ```java Iterator readStream(String workflowId, String key) ``` Read all values written to a stream by the specified workflow. Returns a blocking iterator that stops when the stream is closed or the workflow terminates. Can be called from outside the workflow — typically from a separate thread or external process. If no workflow with the given ID exists, iterating throws `DBOSNonExistentWorkflowException`. On PostgreSQL, the iterator wakes up immediately when a new value is written (via `LISTEN`/`NOTIFY`). On CockroachDB, it falls back to polling once per second. **Parameters:** - **workflowId**: The ID of the workflow that owns the stream. - **key**: The stream key to read. #### retrieveWorkflow ```java WorkflowHandle retrieveWorkflow(String workflowId) ``` Retrieve the [handle](./workflows-steps.md#workflowhandle) of a workflow. **Parameters**: - **workflowId**: The ID of the workflow whose handle to retrieve. #### getResult ```java T getResult(String workflowId) throws E ``` Wait for the workflow to complete and return its result, or rethrow the exception it threw. This is a convenience alternative to `retrieveWorkflow(workflowId).getResult()`. **Parameters:** - **workflowId**: The ID of the workflow whose result to retrieve. #### patch ```java boolean patch(String patchName) ``` Insert a patch marker at the current point in workflow history. Returns `true` if it was successfully inserted or `false` if there is already a checkpoint present at this point in history. Used to safely upgrade workflow code, see the [patching tutorial](../tutorials/upgrading-workflows.md) for more detail. #### deprecatePatch ```java boolean deprecatePatch(String patchName) ``` Safely bypass a patch marker at the current point in workflow history if present. Always returns `true`. Used to safely deprecate patches, see the [patching tutorial](../tutorials/upgrading-workflows.md) for more detail. ### Enqueueing Workflows by Name #### enqueueWorkflow ```java WorkflowHandle enqueueWorkflow( EnqueueOptions options, Object[] args) WorkflowHandle enqueueWorkflow( EnqueueOptions options, Object[] positionalArgs, Map namedArgs) ``` Enqueue a workflow by name, without a reference to its function, and return a handle to it. This takes the same [`EnqueueOptions`](./client.md#enqueueoptions) as [`DBOSClient.enqueueWorkflow`](./client.md#enqueueworkflow) and writes the same database record, so the enqueued workflow may be implemented by another process, another application sharing this system database, or an application written in another language. Unlike [`startWorkflow`](./workflows-steps.md#startworkflow), the options are not validated against this process's registered workflows and queues. If no application version is set, the workflow is only dequeued by an executor running the owning application's latest version. The enqueued workflow is owned by this application unless [`EnqueueOptions.withApplicationName`](./client.md#enqueueoptions) names another one; see [Sharing a System Database](../../explanations/sharing-a-system-database.md). `enqueueWorkflow` may be called from inside a workflow, where the enqueued workflow is recorded as a child: if the calling workflow is recovered, it gets a handle to the original child rather than enqueueing a second one. It may not be called from inside a step; doing so throws `IllegalStateException`. Inside a workflow, the enqueued workflow's timeout is resolved as for `startWorkflow`: unless `EnqueueOptions` sets one, it inherits the calling workflow's timeout; `withNoTimeout()` declines it. **Parameters:** - **options**: The workflow name, queue, and other options; see [`EnqueueOptions`](./client.md#enqueueoptions). - **args** / **positionalArgs**: The workflow's positional arguments. - **namedArgs**: The workflow's named arguments, for targets that take them (for example, a Python workflow with keyword arguments). Only [portable serialization](../../explanations/portable-workflows.md) carries named arguments, so passing any requires `withSerialization(SerializationStrategy.PORTABLE)` on the options; otherwise the call throws `IllegalArgumentException`. **Example Syntax:** ```java var options = new EnqueueOptions("processOrder", "com.example.OrderServiceImpl", QueueName.of("orders")) .withApplicationName("order-service"); WorkflowHandle handle = dbos.enqueueWorkflow(options, new Object[] {"order-123"}); ``` ### Workflow Management Methods #### WorkflowStatus Some workflow introspection and management methods return a `WorkflowStatus`. This object has the following definition: ```java public record WorkflowStatus( // The workflow ID String workflowId, // The workflow status: ENQUEUED, PENDING, DELAYED, SUCCESS, ERROR, CANCELLED, or MAX_RECOVERY_ATTEMPTS_EXCEEDED WorkflowState status, // The name of the workflow function String workflowName, // The class of the workflow function String className, // The name given to the class instance, if any String instanceName, // The authenticated user who initiated the workflow, if any String authenticatedUser, // The assumed role for the workflow execution, if any String assumedRole, // Roles authenticated for the workflow List authenticatedRoles, // The deserialized workflow input Object[] input, // The workflow's output, if any Object output, // The error the workflow threw, if any ErrorResult error, // The ID of the executor (process) that most recently executed this workflow String executorId, // When the workflow was created Instant createdAt, // The last time the workflow status was updated Instant updatedAt, // The application version on which this workflow was started String appVersion, // The application identifier String appId, // The number of times this workflow has been started Integer recoveryAttempts, // If this workflow was enqueued, on which queue String queueName, // The workflow timeout duration, if any Duration timeout, // The absolute deadline for the workflow, if any Instant deadline, // When the workflow started executing (after being dequeued), if applicable Instant startedAt, // The deduplication ID assigned to this workflow, if any String deduplicationId, // The priority assigned to this workflow in its queue, if any Integer priority, // The queue partition key, if any String queuePartitionKey, // The ID of the workflow this was forked from, if any String forkedFrom, // The parent workflow ID if this is a child workflow, if any String parentWorkflowId, // Whether another workflow has been forked from this one Boolean wasForkedFrom, // Time until which the workflow is delayed before starting Instant delayUntil, // When the workflow completed (terminal states only) Instant completedAt, // The serialization format used for the workflow's inputs/outputs String serialization, // Custom key-value metadata attached to the workflow Map attributes, // The name of the schedule that started this workflow, if any String scheduleName, // The application that owns this workflow, or null if no application owns it String applicationName ) ``` See [Sharing a System Database](../../explanations/sharing-a-system-database.md) for how applications own workflows. #### listWorkflows ```java List listWorkflows(ListWorkflowsInput input) ``` Retrieve a list of [`WorkflowStatus`](#workflowstatus) of all workflows matching specified criteria. #### ListWorkflowsInput `ListWorkflowsInput` is a with-based configuration record for filtering and customizing workflow queries. All fields are optional. **`with` Methods:** Many filters accept either a single value or a list. Single-value overloads are provided for convenience. ##### withWorkflowIds ```java ListWorkflowsInput withWorkflowIds(String workflowId) ListWorkflowsInput withWorkflowIds(List workflowIds) ``` Filter by one or more workflow IDs. ##### withClassName ```java ListWorkflowsInput withClassName(String className) ``` Filter workflows by the class name containing the workflow function. ##### withInstanceName ```java ListWorkflowsInput withInstanceName(String instanceName) ``` Filter workflows by the instance name of the class. ##### withWorkflowName ```java ListWorkflowsInput withWorkflowName(String workflowName) ListWorkflowsInput withWorkflowName(List workflowNames) ``` Filter workflows by the workflow function name. ##### withAuthenticatedUser ```java ListWorkflowsInput withAuthenticatedUser(String authenticatedUser) ListWorkflowsInput withAuthenticatedUser(List authenticatedUsers) ``` Filter workflows run by this authenticated user. ##### withStartTime ```java ListWorkflowsInput withStartTime(Instant startTime) ``` Retrieve workflows created after this timestamp. ##### withEndTime ```java ListWorkflowsInput withEndTime(Instant endTime) ``` Retrieve workflows created before this timestamp. ##### withStatus ```java ListWorkflowsInput withStatus(WorkflowState... statuses) ListWorkflowsInput withStatus(List statuses) ``` Filter workflows by status. Status must be one of: `ENQUEUED`, `PENDING`, `DELAYED`, `SUCCESS`, `ERROR`, `CANCELLED`, or `MAX_RECOVERY_ATTEMPTS_EXCEEDED`. ##### withApplicationVersion ```java ListWorkflowsInput withApplicationVersion(String applicationVersion) ListWorkflowsInput withApplicationVersion(List applicationVersions) ``` Retrieve workflows tagged with this application version. ##### withLimit ```java ListWorkflowsInput withLimit(Integer limit) ``` Retrieve up to this many workflows. ##### withOffset ```java ListWorkflowsInput withOffset(Integer offset) ``` Skip this many workflows from the results returned (for pagination). ##### withSortDesc ```java ListWorkflowsInput withSortDesc(Boolean sortDesc) ``` Sort the results in descending (true) or ascending (false) order by workflow creation time. ##### withExecutorIds ```java ListWorkflowsInput withExecutorIds(String executorId) ListWorkflowsInput withExecutorIds(List executorIds) ``` Retrieve workflows that ran on these executor processes. ##### withQueueName ```java ListWorkflowsInput withQueueName(String queueName) ListWorkflowsInput withQueueName(List queueNames) ``` Retrieve workflows that were enqueued on these queues. ##### withWorkflowIdPrefix ```java ListWorkflowsInput withWorkflowIdPrefix(String workflowIdPrefix) ListWorkflowsInput withWorkflowIdPrefix(List workflowIdPrefixes) ``` Filter workflows whose IDs start with the specified prefix. ##### withQueuesOnly ```java ListWorkflowsInput withQueuesOnly(Boolean queuedOnly) ``` Select only workflows that were enqueued. ##### withLoadInput ```java ListWorkflowsInput withLoadInput(Boolean value) ``` Controls whether to load workflow input data (default: true). ##### withLoadOutput ```java ListWorkflowsInput withLoadOutput(Boolean value) ``` Controls whether to load workflow output data (results and errors) (default: true). ##### withForkedFrom ```java ListWorkflowsInput withForkedFrom(String workflowId) ListWorkflowsInput withForkedFrom(List workflowIds) ``` Filter to workflows that were forked from the specified workflow(s). ##### withParentWorkflowId ```java ListWorkflowsInput withParentWorkflowId(String parentWorkflowId) ListWorkflowsInput withParentWorkflowId(List parentWorkflowIds) ``` Filter to workflows that are children of the specified parent workflow(s). ##### withWasForkedFrom ```java ListWorkflowsInput withWasForkedFrom(Boolean wasForkedFrom) ``` Filter to workflows from which another workflow was forked. ##### withHasParent ```java ListWorkflowsInput withHasParent(Boolean hasParent) ``` Filter to workflows that have a parent workflow. ##### withCompletedAfter ```java ListWorkflowsInput withCompletedAfter(Instant completedAfter) ``` Filter to workflows that completed after this timestamp. ##### withCompletedBefore ```java ListWorkflowsInput withCompletedBefore(Instant completedBefore) ``` Filter to workflows that completed before this timestamp. ##### withDequeuedAfter ```java ListWorkflowsInput withDequeuedAfter(Instant dequeuedAfter) ``` Filter to workflows that were dequeued (started execution) after this timestamp. ##### withDequeuedBefore ```java ListWorkflowsInput withDequeuedBefore(Instant dequeuedBefore) ``` Filter to workflows that were dequeued (started execution) before this timestamp. ##### withScheduleName ```java ListWorkflowsInput withScheduleName(String scheduleName) ListWorkflowsInput withScheduleName(List scheduleNames) ``` Retrieve workflows started by these [schedules](#schedule-management-methods). ##### withApplicationName ```java ListWorkflowsInput withApplicationName(String applicationName) ListWorkflowsInput withApplicationName(List applicationNames) ``` Retrieve workflows owned by these applications, plus workflows no application owns. If unset, only this application's workflows (and unowned ones) are listed, unless the query is filtered by [workflow ID](#withworkflowids), which is never narrowed by default; pass an empty list to list every application's workflows. See [Sharing a System Database](../../explanations/sharing-a-system-database.md). ##### withAttributes ```java ListWorkflowsInput withAttributes(Map attributes) ``` Filter to workflows whose custom attributes contain all the specified key-value pairs (PostgreSQL `@>` containment check using a GIN index). Pass a map with the subset of attributes to match. #### listWorkflowSteps ```java List listWorkflowSteps(String workflowId) List listWorkflowSteps(String workflowId, Integer limit, Integer offset) ``` Retrieve the execution steps of a workflow. The `limit` and `offset` parameters support pagination over large step lists. This is a list of `StepInfo` objects, with the following structure: ```java StepInfo( // The sequential ID of the step within the workflow int functionId, // The name of the step function String functionName, // The output returned by the step, if any Object output, // The error returned by the step, if any ErrorResult error, // If the step starts or retrieves the result of a workflow, its ID String childWorkflowId, // When the step started executing Instant startedAt, // When the step completed Instant completedAt, // The serialization format used for the step's output String serialization, // The application that ran this step, or null if no application owns it. // Usually the owner of the step's workflow, but a workflow resumed or forked // by another application records its new steps under that application. String applicationName ) ``` #### getWorkflowStatus ```java Optional getWorkflowStatus(String workflowId) ``` Retrieve the [`WorkflowStatus`](#workflowstatus) of a single workflow by ID. #### cancelWorkflow ```java void cancelWorkflow(String workflowId) void cancelWorkflow(String workflowId, boolean cancelChildren) void cancelWorkflows(List workflowIds) void cancelWorkflows(List workflowIds, boolean cancelChildren) ``` Cancel one or more workflows. This sets their status to `CANCELLED`, removes them from their queue (if enqueued) and preempts execution (interrupting at the beginning of the next step). **Parameters:** - **cancelChildren**: If `true`, recursively cancel all descendant workflows spawned by the cancelled workflow(s). Defaults to `false`. #### resumeWorkflow ```java WorkflowHandle resumeWorkflow(String workflowId) WorkflowHandle resumeWorkflow(String workflowId, String queueName) List> resumeWorkflows(List workflowIds) List> resumeWorkflows(List workflowIds, String queueName) ``` Resume one or more workflows from their last completed step. You can use this to resume workflows that are cancelled or have exceeded their maximum recovery attempts. Resuming a workflow sets it back to `ENQUEUED`, on `queueName` if given and otherwise on the DBOS internal queue, which has no flow control, so the workflow starts as soon as an executor dequeues it. This works on a workflow that is already `ENQUEUED`, so you can also use it to start an enqueued workflow without waiting on its queue, or to move a workflow stranded on a deleted or [newly partitioned](./queues.md#queueoptions) queue. A resumed workflow keeps its application version, so it runs only on an executor of that version. Workflows that already completed with `SUCCESS` or `ERROR` are left unchanged. **Parameters:** - **queueName**: The queue to enqueue the resumed workflow on. If omitted or `null`, it is enqueued on the DBOS internal queue. #### deleteWorkflow ```java void deleteWorkflow(String workflowId) void deleteWorkflow(String workflowId, boolean deleteChildren) void deleteWorkflows(List workflowIds) void deleteWorkflows(List workflowIds, boolean deleteChildren) ``` Permanently delete one or more workflows and their recorded steps from the database. **Parameters:** - **deleteChildren**: If `true`, also delete all child workflows spawned by the deleted workflow(s). Defaults to `false`. #### forkWorkflow ```java WorkflowHandle forkWorkflow(String workflowId, int startStep) WorkflowHandle forkWorkflow(String workflowId, int startStep, ForkOptions options) ``` ```java public record ForkOptions( String forkedWorkflowId, String applicationVersion, Duration timeout, // null means no timeout String queueName, String queuePartitionKey ) { ForkOptions withForkedWorkflowId(String forkedWorkflowId); ForkOptions withApplicationVersion(String applicationVersion); ForkOptions withTimeout(Duration timeout); ForkOptions withTimeout(long value, TimeUnit unit); ForkOptions withQueue(QueueName queue); ForkOptions withQueue(String queueName); ForkOptions withQueue(Queue queue); // deprecated since 1.1 ForkOptions withQueuePartitionKey(String queuePartitionKey); } ``` Start a new execution of a workflow from a specific step. The input step ID (`startStep`) must match the step number of the step returned by workflow introspection. The specified `startStep` is the step from which the new workflow will start, so any steps whose ID is less than `startStep` will not be re-executed. **Parameters:** - **workflowId**: The ID of the workflow to fork - **startStep**: The step from which to fork the workflow - **options**: - **forkedWorkflowId**: The workflow ID for the newly forked workflow (if not provided, generate a UUID) - **applicationVersion**: The application version for the forked workflow (inherited from the original if not provided) - **timeout**: A `Duration` timeout for the forked workflow. Pass `null` for no timeout. - **queueName**: Enqueue the forked workflow on this queue instead of starting it immediately. Set it with `withQueue(QueueName)` or `withQueue(String)`; `withQueue(Queue)` is *(deprecated since 1.1)*. - **queuePartitionKey**: Partition key for the queue (only for partitioned queues). #### updateWorkflowAttributes ```java void updateWorkflowAttributes(String workflowId, Map attributes) ``` Replace the custom attributes attached to a workflow. Pass `null` or an empty map to clear all attributes. Safe to call from within a workflow — the update is recorded as a step so it executes exactly once even if the workflow is recovered. **Parameters:** - **workflowId**: The ID of the workflow whose attributes to update. - **attributes**: The new JSON-serializable key-value attributes, or `null`/empty map to clear. #### setWorkflowDelay ```java void setWorkflowDelay(String workflowId, Duration delay) void setWorkflowDelay(String workflowId, Instant delayUntil) ``` Pause an enqueued or pending workflow so it will not be dequeued until after the specified duration or absolute time. The workflow's status changes to `DELAYED` while waiting. **Parameters:** - **workflowId**: The ID of the workflow to delay. - **delay**: How long from now to delay the workflow. - **delayUntil**: The absolute time until which to delay the workflow. #### listApplicationVersions ```java List listApplicationVersions() ``` Return all registered application versions, ordered by timestamp descending. ```java public record VersionInfo( // The generated version ID String versionId, // The human-readable version name (e.g., "v1.2.3") String versionName, // When this version was promoted Instant versionTimestamp, // When this version record was created Instant createdAt, // The application that registered this version, or null if no application owns it String applicationName ) ``` #### getLatestApplicationVersion ```java VersionInfo getLatestApplicationVersion() ``` Return the most recently promoted application version. #### setLatestApplicationVersion ```java void setLatestApplicationVersion(String versionName) ``` Promote a version by name to be the latest application version. Used during blue-green deployments to control which version new workflows are assigned to. See [upgrading workflows](../tutorials/upgrading-workflows.md) for more detail. ### Schedule Management Methods These methods manage [`WorkflowSchedule`](#workflowschedule) records that periodically invoke workflows on a cron schedule. All schedule methods require DBOS to be launched first. #### WorkflowSchedule ```java public record WorkflowSchedule( String id, // Generated schedule ID (read-only; null on creation) String scheduleName, // Unique name for this schedule String workflowName, // Name of the workflow function to invoke String className, // Class containing the workflow (use @WorkflowClassName for a portable name) String cron, // Cron expression (Spring 5.3+ format) ScheduleStatus status, // ACTIVE or PAUSED Object context, // Optional context object passed to the workflow Instant lastFiredAt, // When the schedule last fired (read-only) boolean automaticBackfill,// If true, missed firings are retroactively started on launch ZoneId cronTimezone, // Timezone for interpreting the cron expression (defaults to UTC) String queueName, // Queue to enqueue scheduled workflows on String applicationName // Application that owns the schedule and runs its workflows ) ``` **Constructors:** ```java new WorkflowSchedule(String scheduleName, String workflowName, String className, String cron) ``` Creates an `ACTIVE` schedule with no backfill, UTC timezone, and no queue override. **`with` Methods:** - **`withScheduleName(String name)`** — Change the schedule's unique name. - **`withWorkflowName(String name)`** — Change the target workflow function name. - **`withClassName(String name)`** — Change the target class name. - **`withCron(String cron)`** — Change the cron expression. - **`withStatus(ScheduleStatus status)`** — Set `ACTIVE` or `PAUSED`. - **`withContext(Object context)`** — Attach a context object, serialized and passed to the workflow as its third argument. - **`withAutomaticBackfill(boolean value)`** — If `true`, any firings missed while the app was down are retroactively started when the app launches. - **`withCronTimezone(ZoneId timezone)`** — Interpret the cron expression in this timezone instead of UTC. - **`withQueueName(String queueName)`** — Enqueue scheduled executions on this queue. - **`withApplicationName(String applicationName)`** — The application that owns the schedule and runs its workflows. Defaults to the creating application. Schedule names are unique across all applications sharing a system database: creating or applying a schedule whose name a different application owns throws [`DBOSApplicationNameConflictException`](#dbosapplicationnameconflictexception). See [Sharing a System Database](../../explanations/sharing-a-system-database.md). ```java public enum ScheduleStatus { ACTIVE, PAUSED } ``` #### applySchedules ```java void applySchedules(WorkflowSchedule... schedules) void applySchedules(List schedules) ``` Atomically create or update a set of schedules in one transaction. Existing schedules are upserted by name: all definition fields are replaced with the new declaration's values (so any optional setting left unset is cleared, e.g. an omitted queue name reverts the schedule to the default scheduler queue), while the schedule's status and last-fired time are preserved. This is the recommended way to declare schedules at application startup — call it once after `dbos.launch()` and it will always reflect your current schedule definitions. #### createSchedule ```java void createSchedule(WorkflowSchedule schedule) ``` Create a single schedule. Throws if a schedule with the same `scheduleName` already exists. #### getSchedule ```java Optional getSchedule(String name) ``` Retrieve a schedule by name. Returns empty if not found. #### listSchedules ```java List listSchedules(List status, List workflowName, List namePrefix) List listSchedules(List status, List workflowName, List namePrefix, List applicationName) ``` List schedules with optional filters. Pass `null` for any filter to skip it. **Parameters:** - **status**: Filter by `ScheduleStatus.ACTIVE` or `ScheduleStatus.PAUSED`. - **workflowName**: Filter by workflow function name. - **namePrefix**: Filter by schedule name prefix. - **applicationName**: List schedules owned by these applications, plus schedules no application owns. If `null` (or omitted), lists this application's schedules; pass an empty list to list every application's schedules. #### pauseSchedule ```java void pauseSchedule(String name) ``` Pause a schedule. A paused schedule does not fire until resumed. #### resumeSchedule ```java void resumeSchedule(String name) ``` Resume a paused schedule. #### deleteSchedule ```java void deleteSchedule(String name) ``` Delete a schedule by name. No-op if the schedule does not exist. #### backfillSchedule ```java List> backfillSchedule(String scheduleName, Instant start, Instant end) ``` Manually enqueue all executions of a schedule that would have fired between `start` (exclusive) and `end` (exclusive). Uses the same deterministic workflow IDs as the live scheduler, so already-executed times are skipped. Useful for recovering from outages when automatic backfill is not enabled. #### triggerSchedule ```java WorkflowHandle triggerSchedule(String scheduleName) ``` Immediately fire a scheduled workflow outside its normal cron cadence. Returns a handle to the enqueued execution. ### Debouncing You can create a `Debouncer` to debounce your workflows. Debouncing delays workflow execution until some time has passed since the workflow was last called. This is useful for preventing wasted work when a workflow may be triggered multiple times in quick succession. For example, if a user is editing an input field, you can debounce their changes to execute a processing workflow only after they haven't edited the field for some time. #### Debouncer Obtain a `Debouncer` via `dbos.debouncer()`: ```java Debouncer debouncer() ``` `Debouncer` is an immutable builder. Configure it with the following methods before calling `debounce`: - **`withDebounceTimeout(Duration debounceTimeout)`**: Set an absolute cap on how long the debouncer may keep absorbing calls for a single key. After this duration elapses from the first call, the user workflow starts regardless of further incoming calls. - **`withQueue(QueueName queue)`** / **`withQueue(String queueName)`**: Enqueue the user workflow on the specified queue when the debounce period elapses instead of starting it directly. `withQueue(Queue queue)` is *(deprecated since 1.1)*. - **`withAppVersion(String appVersion)`**: Target a specific application version for the user workflow. - **`withPriority(Integer priority)`**: Set the priority for the user workflow. A priority requires a queue: if a priority is set without `withQueue`, `debounce` throws `IllegalArgumentException`. - **`withDeduplicationId(String deduplicationId)`** *(deprecated since 1.1)*: Set a deduplication ID forwarded to the user workflow. This will be ignored from the next release, where the debouncer sets the deduplication ID itself, and removed in 2.0. #### debouncer.debounce ```java WorkflowHandle debounce( String debounceKey, Duration debouncePeriod, ThrowingRunnable wfLambda) WorkflowHandle debounce( String debounceKey, Duration debouncePeriod, ThrowingSupplier wfLambda) ``` Submit a workflow for execution but delay it by `debouncePeriod`. Returns a handle to the workflow. The workflow may be debounced again, which further delays its execution (up to `debounceTimeout`). When the workflow eventually executes, it uses the **last** set of inputs passed into `debounce`. After the workflow begins execution, the next call to `debounce` starts the debouncing process again for a new workflow execution. **Parameters:** - **debounceKey**: A key used to group workflow executions that will be debounced together. For example, if the debounce key is set to customer ID, each customer's workflows are debounced separately. - **debouncePeriod**: Inactivity window before the user workflow runs; each call resets it. Must be positive. - **wfLambda**: A lambda calling exactly one `@Workflow` method on a registered proxy. **Example Syntax:** ```java var dbos = new DBOS(config); MyService svc = dbos.registerProxy(MyService.class, new MyServiceImpl()); dbos.launch(); var debouncer = dbos.debouncer() .withDebounceTimeout(Duration.ofMinutes(5)); // Each time a user submits input, debounce the processInput workflow. // The workflow will run 60 seconds after the user stops submitting. WorkflowHandle handle = debouncer.debounce( userId, Duration.ofSeconds(60), () -> svc.processInput(userInput)); String result = handle.getResult(); ``` ### DBOS Context Variables #### workflowId ```java static String workflowId() ``` Retrieve the ID of the current workflow. Returns `null` if not called from a workflow or step. #### stepId ```java static Integer stepId() ``` Returns the unique ID of the current step within its workflow. Returns `null` if not called from a step. #### inWorkflow ```java static boolean inWorkflow(); ``` Return `true` if the current calling context is executing a workflow, or `false` otherwise. #### inStep ```java static boolean inStep(); ``` Return `true` if the current calling context is executing a workflow step, or `false` otherwise. #### authenticatedUser ```java static @Nullable String authenticatedUser() ``` Return the authenticated user associated with the current workflow, or `null` if none was set. Must be called from within a workflow or step. #### assumedRole ```java static @Nullable String assumedRole() ``` Return the assumed role for the current workflow execution, or `null` if none was set. Must be called from within a workflow or step. #### authenticatedRoles ```java static @Nullable List authenticatedRoles() ``` Return the list of authenticated roles for the current workflow, or `null` if none were set. Must be called from within a workflow or step. ### Timeout `Timeout` is a sealed interface used by `StartWorkflowOptions` and `WorkflowOptions` to control how inherited timeouts are handled. ```java import dev.dbos.transact.workflow.Timeout; ``` **Factory methods:** - **`Timeout.of(Duration duration)`** — Set an explicit timeout of the given duration. - **`Timeout.of(long value, TimeUnit unit)`** — Set an explicit timeout. - **`Timeout.none()`** — Opt out of any inherited timeout. The workflow will run without a timeout regardless of what the calling context specifies. - **`Timeout.inherit()`** — Explicitly inherit the timeout from the calling context (the default behavior when no timeout is set). **Example:** ```java // Detach a child workflow from the parent's timeout dbos.startWorkflow(() -> proxy.longRunningChild(), new StartWorkflowOptions().withTimeout(Timeout.none())); ``` ### WorkflowState `WorkflowState` is an enum representing the possible states of a workflow. It is used in [`WorkflowStatus`](#workflowstatus) and [`ListWorkflowsInput`](#listworkflowsinput). ```java import dev.dbos.transact.workflow.WorkflowState; ``` ```java public enum WorkflowState { PENDING, // Currently executing ENQUEUED, // Waiting on a queue to be dequeued DELAYED, // Waiting until a delay expires before being dequeued SUCCESS, // Completed successfully ERROR, // Threw an unhandled exception CANCELLED, // Cancelled by cancelWorkflow() MAX_RECOVERY_ATTEMPTS_EXCEEDED // Crashed too many times; requires manual intervention } ``` `WorkflowState.isActive()` returns `true` for `PENDING`, `ENQUEUED`, and `DELAYED`. ### Serialization Strategy Several DBOS methods accept an optional `SerializationStrategy` parameter that controls how data is serialized. This is useful for cross-language interoperability—for example, if a TypeScript or Python DBOS application needs to read events or messages set by a Java application. ```java import dev.dbos.transact.workflow.SerializationStrategy; ``` The available strategies are: - **`SerializationStrategy.DEFAULT`**: Uses the serializer configured in [`DBOSConfig`](./lifecycle.md#custom-serialization) (defaults to Jackson). Inside a workflow, `send`, `sendBulk`, `setEvent`, and `writeStream` instead use the serialization format recorded for the running workflow, so a workflow started with portable serialization writes portable messages, events, and stream values by default. - **`SerializationStrategy.PORTABLE`**: Uses a portable JSON format (`portable_json`) that can be deserialized by DBOS applications in any language. - **`SerializationStrategy.NATIVE`**: Explicitly uses the native Java Jackson serializer (`java_jackson`). ### Alert Handlers Alert handlers let you receive internal DBOS alerts — such as workflow recovery failures or system warnings — and route them to your observability infrastructure (logging, metrics, PagerDuty, etc.). #### registerAlertHandler ```java void registerAlertHandler(AlertHandler handler) ``` Register a handler to receive alerts generated by DBOS. Must be called **before** `dbos.launch()`. ```java @FunctionalInterface public interface AlertHandler { void invoke(String name, String message, Map metadata); } ``` **Parameters:** - **name**: A short identifier for the alert type. - **message**: A human-readable description of the alert. - **metadata**: Additional key-value context about the alert. **Example:** ```java DBOS dbos = new DBOS(config); dbos.registerAlertHandler((name, message, metadata) -> logger.warn("DBOS alert [{}]: {} {}", name, message, metadata)); dbos.launch(); ``` ### Exceptions #### DBOSApplicationNameConflictException ```java public class DBOSApplicationNameConflictException extends RuntimeException { String kind(); // "Queue", "Schedule", or "Application version" String name(); // The name both applications claim String owner(); // The application that already owns the name } ``` Thrown when registering a queue, creating or applying a schedule, or promoting an application version whose name is already owned by a different application sharing the system database. Queue, schedule, and version names are unique across all applications sharing a system database. Either choose a different name or, if the owning application was renamed, transfer its rows first with [`dbosctl sysdb rename-application`](../../conductor/reference/dbosctl.md#dbosctl-sysdb-rename-application) or [`DBOSClient.renameApplication`](./client.md#renameapplication). See [Sharing a System Database](../../explanations/sharing-a-system-database.md). #### DBOSSystemDatabaseException ```java public class DBOSSystemDatabaseException extends RuntimeException { String sqlState(); Throwable databaseException(); // deprecated since 1.1 } ``` Thrown when a system database operation fails and DBOS will not retry it any further. This covers two different situations, which `sqlState()` distinguishes: - **Connectivity failures** are retried first, so one arriving here means the database stayed unreachable across many attempts. - **Failures that retrying cannot fix**, such as a constraint violation or a missing table, are not retried and are thrown immediately. **Methods:** - **`sqlState()`**: The SQLSTATE of the underlying failure, or `null` if it carried none. - **`getCause()`**: The underlying `SQLException`. It is the only level of wrapping, so the database's own exception is always one level down. - **`databaseException()`** *(deprecated since 1.1)*: Returns the same object as `getCause()`; use `getCause()` instead. --- ## DBOS Plugin Architecture :::tip Unless you intend to extend the DBOS Transact library, you can ignore this topic. ::: DBOS Transact for Java provides extension mechanisms for integrating with Java frameworks and external event sources. **Framework integrations** — If you are using a supported framework, use its dedicated integration instead of building on this API directly: - **Spring Boot**: Use the [`transact-spring-boot-starter`](./spring-boot-starter.md), which auto-configures DBOS, registers workflows from Spring beans, and manages the DBOS lifecycle alongside the Spring application context. **Custom integrations** — If you are building a new framework integration or event receiver (Kafka, SQS, a custom scheduler, etc.), this page documents the lower-level extension points: lifecycle listeners, workflow registration, and external state storage. ### Lifecycle Listeners Lifecycle listeners are a broad category of extensions that need to be notified when DBOS launches or shuts down. Often, these listeners are connecting to extenal systems to integrate outside events into DBOS. Examples include: - Kafka message consumers - SQS/JMS message receivers - Clock-based schedulers What event receivers have in common is that they run in the background and execute DBOS functions in response to externally-triggered circumstances. #### Lifecycle Listener Architecture During program initialization, event receivers are constructed and registered with the DBOS lifecycle. Configuration may be collected during initialization, but no actions are taken until `dbos.launch()` is called. Upon `launch`, event receivers should review their registrations and connect to any outside resources and report clear error messages if this fails. After any initialization is complete, event receivers should commence processing events and initiating DBOS workflow method calls in response. Event receivers are also told to deinitialize when DBOS shuts down. #### Listener Lifecycle DBOS event receiver objects should implement the `DBOSLifecycleListener` interface: ```java /** * For registering callbacks that hear about `dbos.launch()` and `dbos.shutdown()`. At this point, * DBOS is ready to run workflows, and no additional registrations are allowed. */ public interface DBOSLifecycleListener { /** Called from within dbos.launch, after workflow processing is allowed */ void dbosLaunched(DBOS dbos); /** Called from within dbos.shutdown, before workflow processing is stopped */ void dbosShutDown(); } ``` Upon construction, event receivers should register themselves via `dbos.integration().registerLifecycleListener(listener)`. Upon `dbos.launch()`, all registered `dbosLaunched()` methods will be called. Upon `dbos.shutdown()`, all registered `dbosShutDown()` methods will be called. ### DBOSIntegration `DBOSIntegration` is the interface for framework integrators, AOP aspects, and event listeners. Obtain it via: ```java DBOSIntegration integration = dbos.integration(); ``` :::warning `DBOSIntegration` is not part of the primary public API and may change without notice. Application code should use methods on `DBOS` directly. ::: #### registerLifecycleListener ```java void registerLifecycleListener(DBOSLifecycleListener listener) ``` Register a lifecycle listener. Must be called before `dbos.launch()`. See [Listener Lifecycle](#listener-lifecycle) for the `DBOSLifecycleListener` interface. #### getRegisteredWorkflows ```java Collection getRegisteredWorkflows() ``` Return all workflow methods registered with DBOS. May be called before or after `dbos.launch()`. #### getRegisteredWorkflowInstances ```java Collection getRegisteredWorkflowInstances() ``` Return all class instances containing registered workflow methods. May be called before or after `dbos.launch()`. #### getRegisteredWorkflow ```java Optional getRegisteredWorkflow(String workflowName, String className) Optional getRegisteredWorkflow(String workflowName, String className, String instanceName) ``` Find a specific registered workflow by its workflow name, class name, and optional instance name. Returns empty if no matching workflow is found. May be called before or after `dbos.launch()`. **Parameters:** - **workflowName**: The name of the workflow function. - **className**: The name of the class containing the workflow. - **instanceName**: The instance name of the class (defaults to `""` if omitted). #### startRegisteredWorkflow ```java WorkflowHandle startRegisteredWorkflow( RegisteredWorkflow regWorkflow, Object[] args, StartWorkflowOptions options) ``` Start or enqueue a workflow by its `RegisteredWorkflow` registration. Use this when dispatching workflows by registration rather than by direct invocation (e.g., from an event listener). Must be called after `dbos.launch()`. **Parameters:** - **regWorkflow**: The registered workflow to start. Obtain via `getRegisteredWorkflows()` or `getRegisteredWorkflow(...)`. - **args**: Arguments to pass to the workflow function. - **options**: Execution options such as workflow ID, queue, and timeout. Pass `null` to use defaults. #### runWorkflow ```java Object runWorkflow(Object target, String instanceName, Method method, Object[] args, Workflow wfTag) throws Exception ``` Execute a workflow method via its reflective `Method` handle. Intended for use by AOP interceptors that capture workflow invocations at the proxy boundary. Must be called after `dbos.launch()`. **Parameters:** - **target**: The object instance on which the workflow method is declared. - **instanceName**: The DBOS instance name for `target`, or `null` for the default. - **method**: The workflow `Method` to invoke. - **args**: Arguments to pass to the workflow method. - **wfTag**: The `@Workflow` annotation present on `method`. #### registerWorkflow :::info `DBOS.registerWorkflow` was removed in 0.9. Use `dbos.integration().registerWorkflow(...)` instead. ::: ```java RegisteredWorkflow registerWorkflow(Workflow wfTag, Object target, Method method, String instanceName) ``` Register a single workflow method via reflection, deriving the workflow name and class name from the `@Workflow` annotation and target object. Returns the `RegisteredWorkflow` record. Must be called before `dbos.launch()`. Intended for framework integrators (e.g., AOP-based integrations) that construct proxies themselves. **Parameters:** - **wfTag**: The `@Workflow` annotation instance from the method. - **target**: The object instance that owns the method. - **method**: The `java.lang.reflect.Method` to register as a workflow. - **instanceName**: Optional instance name (can be null). ```java RegisteredWorkflow registerWorkflow( String workflowName, String className, String instanceName, Object target, Method method, Integer maxRecoveryAttempts, SerializationStrategy serializationStrategy) ``` Overload that takes explicit field values rather than deriving them from a `@Workflow` annotation. Use when you need to supply names or options that differ from the annotation. Returns the `RegisteredWorkflow` record. Must be called before `dbos.launch()`. **Parameters:** - **workflowName**: Logical name of the workflow. - **className**: Name of the class that declares the workflow method. - **instanceName**: Optional instance name distinguishing multiple registrations of the same class; may be `null`. - **target**: The object instance on which the method will be invoked. - **method**: The workflow `Method`. - **maxRecoveryAttempts**: Maximum number of recovery attempts; `null` uses the default. - **serializationStrategy**: Strategy used to serialize and deserialize workflow arguments and return values; `null` uses the default. #### upsertExternalState ```java ExternalState upsertExternalState(ExternalState state) ``` :::warning Deprecated `upsertExternalState`, `getExternalState`, and `ExternalState` are *(deprecated since 1.1)* and will be removed in DBOS Java 2.0. The system database table behind them, `event_dispatch_kv`, is being retired: it holds dispatch bookkeeping for in-memory event receivers that the other DBOS SDKs have removed or never implemented, and a shared system database migration will drop it sometime after Java 2.0. Store integration state in your own table instead. ::: Insert or update a value in the DBOS system database for the given key. If a value already exists, it is updated unless `updateTime` or `updateSeq` on the new value is less than what is already stored. Returns the current stored state (which may have a higher version than what was submitted). #### getExternalState ```java Optional getExternalState(String service, String workflowName, String key) ``` *(deprecated since 1.1)* Retrieve a value stored by an external service. Returns empty if no value exists for the key. ##### ExternalState ```java public record ExternalState( String service, String workflowName, String key, String value, Instant updateTime, BigInteger updateSeq) {} ``` *(deprecated since 1.1)* Key fields — together these form a unique record per plugin entry: - **service**: Unique identifier for the event receiver; separates entries from different plugins. - **workflowName**: Fully qualified workflow function name; separates entries by workflow. - **key**: Allows multiple records per service/workflow combination. Value fields: - **value**: A string value. Use JSON for structured data. - **updateTime**: `java.time.Instant` timestamp when the value was set. Upserts with an earlier `updateTime` have no effect. - **updateSeq**: Monotonic sequence number. Upserts with a smaller `updateSeq` have no effect. --- ## Queues(Reference) Workflow queues ensure that workflow functions will be run, without starting them immediately. Queues are useful for controlling the number of workflows run in parallel, or the rate at which they are started. Queue configuration is persisted to the system database, so any DBOS process or [`DBOSClient`](./client.md) connected to the same system database can register, retrieve, and reconfigure queues. All queue management methods below are instance methods on your `DBOS` object and require `dbos.launch()` to have been called. ### Queue Management #### dbos.registerQueue ```java void registerQueue(String name, QueueOptions options) void registerQueue(String name, QueueOptions options, QueueConflictResolution onConflict) ``` Register a queue and persist its configuration to the system database. The queue is owned by, and polled only by, this application. If the queue already exists in the database, the `onConflict` parameter controls whether its configuration is overwritten; it defaults to [`QueueConflictResolution.UPDATE_IF_LATEST_VERSION`](#queueconflictresolution). **Parameters:** - **name**: The name of the queue. Queue names are unique across every application that shares the system database: registering a queue whose name is already owned by a different application throws `DBOSApplicationNameConflictException`. See [Sharing a System Database](../../explanations/sharing-a-system-database.md). - **options**: Initial configuration; see [`QueueOptions`](#queueoptions). - **onConflict**: How to behave when a queue with this name already exists in the system database; see [`QueueConflictResolution`](#queueconflictresolution). **Example syntax:** ```java dbos.registerQueue("email", QueueOptions.setConcurrency(10) .andRateLimit(100, Duration.ofSeconds(60))); ``` #### dbos.updateQueue ```java void updateQueue(String name, QueueOptions options) ``` Update the configuration of an existing queue. Only fields set on `options` are modified; absent fields are left unchanged (see [`QueueOptions`](#queueoptions) and [`Field`](#fieldt)). **Example syntax:** ```java // Change only the concurrency — rate limit and other fields are untouched dbos.updateQueue("email", QueueOptions.setConcurrency(20)); ``` The updated configuration is validated as a whole (see [`QueueOptions`](#queueoptions)), and an update that would produce an invalid queue throws `IllegalArgumentException`. The limits of a queue registered with the deprecated `partitionQueue` option cannot be updated; re-register the queue with per-partition limits instead. #### dbos.findQueue ```java Optional findQueue(String name) ``` Retrieve a queue by name from the system database. Returns empty if no queue with that name has been registered. #### dbos.listQueues ```java List listQueues() List listQueues(List applicationName) ``` Return this application's queues, plus queues no application owns. The second overload returns the queues owned by the named applications, plus queues no application owns; pass `null` for this application's queues, or an empty list for every application's queues. #### dbos.deleteQueue ```java boolean deleteQueue(String name) ``` Delete a queue from the system database. Returns `true` if the queue was deleted, `false` if it did not exist. :::warning Workflows already enqueued on a deleted queue can no longer be dequeued, executed, or recovered. However, if a queue with the same name is later registered, it will dequeue the leftover workflows. Do not rely on this: stale workflows unexpectedly resuming on a future queue is rarely the intended behavior. Instead, cancel or drain pending workflows on the queue before deleting it. Workflows already stuck on a deleted queue can be moved to a registered queue with [`dbos.resumeWorkflow(workflowId, queueName)`](./methods.md#resumeworkflow). ::: ### QueueOptions `QueueOptions` configures a queue for registration or partial update. Each field uses [`Field`](#fieldt) tri-state semantics: absent fields are ignored (leave the current database value unchanged), a present field with a value sets it, and a present field with `null` clears it. ```java // Empty options (all fields absent) QueueOptions.empty() // Static factories — each creates options with a single field set QueueOptions.setConcurrency(Integer value) QueueOptions.setWorkerConcurrency(Integer value) QueueOptions.setRateLimit(Integer max, Duration period) QueueOptions.setRateLimit(int max, long period, TimeUnit unit) QueueOptions.setPartitionConcurrency(Integer value) QueueOptions.setPartitionWorkerConcurrency(Integer value) QueueOptions.setPartitionRateLimit(Integer max, Duration period) QueueOptions.setPartitionRateLimit(int max, long period, TimeUnit unit) QueueOptions.setPollingInterval(Duration value) // Chainable setters — start from any factory and chain additional fields QueueOptions andConcurrency(Integer value) QueueOptions andWorkerConcurrency(Integer value) QueueOptions andRateLimit(Integer max, Duration period) QueueOptions andRateLimit(int max, long period, TimeUnit unit) QueueOptions andPartitionConcurrency(Integer value) QueueOptions andPartitionWorkerConcurrency(Integer value) QueueOptions andPartitionRateLimit(Integer max, Duration period) QueueOptions andPartitionRateLimit(int max, long period, TimeUnit unit) QueueOptions andPollingInterval(Duration value) // Deprecated since 1.1: set a per-partition limit instead QueueOptions.setPartitionQueue(boolean value) QueueOptions andPartitionQueue(boolean value) // Deprecated since 1.1: every queue is a priority queue QueueOptions.setPriorityEnabled(boolean value) QueueOptions andPriorityEnabled(boolean value) ``` **Parameters:** - **concurrency**: The maximum number of workflows from this queue that may run concurrently across all DBOS processes. Pass `null` to remove the limit. - **workerConcurrency**: The maximum number of workflows from this queue that may run concurrently within a single DBOS process. Pass `null` to remove the limit. - **rateLimit**: A limit on the maximum number of workflows (`max`) that may be started in a given `period`. Pass `null` for both to remove the limit. - **partitionConcurrency**: The maximum number of workflows from any one partition of this queue that may run concurrently across all DBOS processes. Pass `null` to remove the limit. - **partitionWorkerConcurrency**: The maximum number of workflows from any one partition of this queue that may run concurrently within a single DBOS process. Pass `null` to remove the limit. - **partitionRateLimit**: A limit on the maximum number of workflows (`max`) that may be started from any one partition in a given `period`. Pass `null` for both to remove the limit. - **priorityEnabled** *(deprecated since 1.1)*: Ignored. Every queue dequeues workflows in priority order, so priority needs no queue configuration, and the queue is always stored as a priority queue. - **pollingInterval**: How often DBOS polls the database for new workflows to dequeue. Defaults to 1 second. - **partitionQueue** *(deprecated since 1.1)*: Enable [partitioning](../tutorials/queue-tutorial.md#partitioning-queues) with the queue-wide limits (`concurrency`, `workerConcurrency`, `rateLimit`) enforced per partition rather than across the queue. Set a per-partition limit instead. If a per-partition limit is also set at registration, this flag has no effect. Setting any per-partition limit [partitions](../tutorials/queue-tutorial.md#partitioning-queues) the queue: every workflow enqueued on it must supply a partition key, and the per-partition limits are enforced for each partition key alongside the queue-wide limits. Partitioning a queue that already has enqueued workflows strands them: they have no partition key, so they are never dequeued. Drain a queue before partitioning it. Workflows stranded this way can be moved to a queue that is not partitioned with [`dbos.resumeWorkflow(workflowId, queueName)`](./methods.md#resumeworkflow). The limits are validated when a queue is registered or updated, and an invalid combination throws `IllegalArgumentException`: - Every concurrency limit, rate-limit `max`, rate-limit `period`, and `pollingInterval` must be greater than zero. - A rate limit's `max` and `period` go together: registering a queue with only one of them set, or an update that would leave only one set, throws. Pass `null` for both to register without a limit or to clear one. - A concurrency limit must be less than or equal to any wider limit that is also set; limits that are not set are not compared: - `workerConcurrency` and `partitionConcurrency` must each be less than or equal to `concurrency`. - `partitionWorkerConcurrency` must be less than or equal to `partitionConcurrency`, `workerConcurrency`, and `concurrency`. ### QueueConflictResolution ```java public enum QueueConflictResolution { ALWAYS_UPDATE, NEVER_UPDATE, UPDATE_IF_LATEST_VERSION } ``` Controls how `dbos.registerQueue` behaves when a queue with the same name already exists in the database: - **`ALWAYS_UPDATE`** — overwrite the existing configuration unconditionally. Default for [`DBOSClient.registerQueue`](./client.md#registerqueue). - **`NEVER_UPDATE`** — leave the existing configuration unchanged; no-op if the queue already exists. - **`UPDATE_IF_LATEST_VERSION`** — overwrite the existing configuration only if the current application version is the latest registered version. Default for `dbos.registerQueue`. Not available on `DBOSClient`. ### Field\ ```java public sealed interface Field permits Field.Absent, Field.Present { record Absent() implements Field {} record Present(T value) implements Field {} static Field absent() // field not specified static Field of(T value) // field set to value (or null to clear) default boolean isPresent() default T get() } ``` `Field` is the tri-state wrapper used by each field of [`QueueOptions`](#queueoptions): - `Field.Absent` — the field was not specified; the current database value is left unchanged. - `Field.Present(value)` — the field was specified with a non-null value; the database value is set to `value`. - `Field.Present(null)` — the field was specified with `null`; the database value is cleared (removed). Use `Field.absent()` and `Field.of(value)` to construct values directly. The `QueueOptions` convenience methods (`set*` / `and*`) call these automatically, so you rarely need to construct `Field` values by hand. ### Queue ```java public record Queue( String name, Integer concurrency, Integer workerConcurrency, boolean priorityEnabled, boolean partitioningEnabled, RateLimit rateLimit, Integer partitionConcurrency, Integer partitionWorkerConcurrency, RateLimit partitionRateLimit, Duration pollingInterval, String applicationName ) { public QueueName queueName(); // the name as a QueueName public boolean hasLimiter(); // a queue-wide rate limit is set public boolean hasPartitionLimits(); // any per-partition limit is set public boolean isPartitioned(); // the queue dequeues per partition key public boolean isLegacyPartitioned(); // partitioned by the deprecated partitionQueue option } // Queue.RateLimit public record RateLimit(int limit, Duration period) {} ``` A queue's configuration as stored in the system database, returned by [`dbos.findQueue`](#dbosfindqueue) and [`dbos.listQueues`](#dboslistqueues). Nullable fields are `null` when the corresponding limit is not set. `applicationName` is the application that owns the queue, or `null` if no application owns it. `priorityEnabled` and `partitioningEnabled` are deprecated since 1.1: `priorityEnabled` is always `true`, and `isPartitioned()` / `isLegacyPartitioned()` replace `partitioningEnabled`. ### QueueName ```java public record QueueName(String value) { public static QueueName of(String value) } ``` A typed wrapper for a queue name, which must not be null or blank. Several APIs take both a workflow ID and a queue name as a `String`, which makes them easy to confuse. In particular, `new StartWorkflowOptions("example-queue")` sets the **workflow ID** to `"example-queue"`; it does not enqueue the workflow. Pass a `QueueName` to say unambiguously that a string names a queue: ```java // Enqueue on "example-queue" var options = new StartWorkflowOptions(QueueName.of("example-queue")); // Equivalent var options2 = new StartWorkflowOptions().withQueue("example-queue"); ``` `QueueName` is accepted by the `StartWorkflowOptions(QueueName)` constructor, `StartWorkflowOptions.withQueue`, `ForkOptions.withQueue`, `Debouncer.withQueue`, `DebouncerClient.withQueue`, and `DBOSConfig.withListenQueue(s)`. The `String` overloads remain available; `QueueName` replaces the deprecated overloads that take a `Queue` object. ### Legacy: In-Memory Queues :::warning Deprecated In-memory queues are deprecated since 1.1 and will be removed in a future release. This covers the `Queue` constructors and `with*` methods, [`dbos.registerQueue(Queue)` and `dbos.registerQueues`](#dbosregisterqueue-legacy), [`dbos.getQueue`](./lifecycle.md#getqueue), and the overloads that take a `Queue` object (`StartWorkflowOptions(Queue)`, `StartWorkflowOptions.withQueue(Queue)`, `ForkOptions.withQueue(Queue)`, `DBOSConfig.withListenQueue(Queue)` / `withListenQueues(Queue...)`, `Debouncer.withQueue(Queue)`, and `DebouncerClient.withQueue(Queue)`). Register database-backed queues with [`dbos.registerQueue(String, QueueOptions)`](#dbosregisterqueue) after launch, look them up with [`dbos.findQueue`](#dbosfindqueue) (the replacement for `getQueue`), and refer to them by name or [`QueueName`](#queuename). ::: #### Queue constructors and `with*` methods ```java new Queue(String name) public Queue withName(String name); public Queue withConcurrency(Integer concurrency); public Queue withWorkerConcurrency(Integer workerConcurrency); public Queue withRateLimit(RateLimit rateLimit); public Queue withRateLimit(int limit, Duration period); public Queue withRateLimit(int limit, long period, TimeUnit unit); public Queue withPriorityEnabled(boolean priorityEnabled); public Queue withPartitioningEnabled(boolean partitioningEnabled); public Queue withPollingInterval(Duration pollingInterval); ``` Construct an in-memory queue at configuration time. In-memory queues must be registered with [`dbos.registerQueue(Queue)`](#dbosregisterqueue-legacy) before `dbos.launch()`. `withPartitioningEnabled` is the in-memory equivalent of the deprecated [`partitionQueue`](#queueoptions) option. **Example Syntax:** ```java // Deprecated, before dbos.launch() Queue queue = new Queue("example-queue").withWorkerConcurrency(5); dbos.registerQueue(queue); // Replacement, after dbos.launch() dbos.registerQueue("example-queue", QueueOptions.setWorkerConcurrency(5)); ``` #### dbos.registerQueue {#dbosregisterqueue-legacy} ```java void registerQueue(Queue queue) void registerQueues(Queue... queues) ``` Register one or more in-memory queues. Must be called before `dbos.launch()`. Replaced by [`dbos.registerQueue(String, QueueOptions)`](#dbosregisterqueue). --- ## Spring Boot Starter The `transact-spring-boot-starter` provides Spring Boot auto-configuration for DBOS Transact. :::danger DBOS requires a PostgreSQL database. If a non-PostgreSQL `DataSource` is provided or auto-detected, startup will throw an `IllegalStateException`. ::: ### Auto-Configured Beans | Bean | Description | |------|-------------| | `DBOSConfig` | Built from `dbos.*` properties. Declare your own to replace it; use `DBOSConfigCustomizer` to extend it. | | `DBOS` | The main DBOS instance, injected from `DBOSConfig`. | | `DBOSLifecycle` | `SmartLifecycle` that calls `dbos.launch()` on context start and `dbos.shutdown()` on context stop. Uses `DEFAULT_PHASE` (`Integer.MAX_VALUE`), so DBOS starts last (after all other beans) and stops first. | | `DBOSAspect` | AOP aspect that intercepts `@Workflow` and `@Step` calls on Spring-managed beans. | | `DBOSWorkflowRegistrar` | Scans all singleton beans after initialization and registers those with `@Workflow` methods. | All beans are `@ConditionalOnMissingBean` — declare your own to replace any of them. ### Configuration Properties All properties are in the `dbos.*` namespace. #### Application | Property | Type | Default | Description | |----------|------|---------|-------------| | `dbos.application.name` | `String` | — | Application name. Falls back to `spring.application.name`. One of the two must be set. | | `dbos.application.version` | `String` | — | Application version string used for version management. | #### Datasource | Property | Type | Default | Description | |----------|------|---------|-------------| | `dbos.datasource.url` | `String` | — | JDBC URL for the DBOS system database. If unset, DBOS uses the application's primary `DataSource` bean. | | `dbos.datasource.username` | `String` | — | Database username. | | `dbos.datasource.password` | `String` | — | Database password. | | `dbos.datasource.schema` | `String` | `dbos` | Schema for DBOS system tables. | | `dbos.datasource.migrate` | `boolean` | `true` | Whether to run database migrations on startup. | #### DBOS Conductor | Property | Type | Default | Description | |----------|------|---------|-------------| | `dbos.conductor.key` | `String` | — | DBOS Conductor API key. | | `dbos.conductor.domain` | `String` | — | DBOS Conductor domain. | #### Admin Server :::warning `dbos.admin-server.*` properties are deprecated since 0.9 and will be removed before 1.0. Use [DBOS Conductor](../../conductor/overview.md) instead. ::: | Property | Type | Default | Description | |----------|------|---------|-------------| | `dbos.admin-server.enabled` | `boolean` | `false` | Whether to enable the admin HTTP server. | | `dbos.admin-server.port` | `int` | `3001` | Port for the admin HTTP server. | #### Other | Property | Type | Default | Description | |----------|------|---------|-------------| | `dbos.executor-id` | `String` | — | Unique executor ID for this instance. | | `dbos.enable-patching` | `boolean` | `false` | Enable [workflow patching](./methods.md#patch). | | `dbos.listen-queues` | `List` | `[]` | Queues this executor should dequeue from. | | `dbos.scheduler-polling-interval` | `Duration` | — | Override the default scheduler polling interval. | | `dbos.use-listen-notify` | `boolean` | `true` | Whether to use PostgreSQL `LISTEN`/`NOTIFY` for real-time event delivery. Automatically set to `false` when CockroachDB is detected. | ### DBOSConfigCustomizer ```java @FunctionalInterface public interface DBOSConfigCustomizer { DBOSConfig customize(DBOSConfig config); } ``` Declare one or more `DBOSConfigCustomizer` beans to modify the auto-configured `DBOSConfig` without replacing it. Customizers run in `@Order` / `Ordered` sequence after the base config is assembled from `dbos.*` properties. ```java @Bean @Order(1) public DBOSConfigCustomizer myCustomizer() { return config -> config.withAdminServer(true).withAdminServerPort(8081); } ``` ### DBOSWorkflowRegistrar Implements `SmartInitializingSingleton`. After all singletons are created, it scans every bean for methods annotated with `@Workflow` or `@Step`. Beans containing `@Workflow` methods are registered with DBOS for durable execution and recovery. `@Step` methods are not registered directly — they are intercepted at runtime by `DBOSAspect`. **Requirements:** - Beans with `@Workflow` or `@Step` methods must be **singletons**. Prototype-scoped beans throw `IllegalStateException`. - DBOS registers the raw (unwrapped) bean target. Calls made via `this` inside a workflow body bypass the Spring proxy and are not intercepted by `DBOSAspect`. Use a self-injected proxy instead (see [Spring Boot Integration tutorial](../tutorials/spring-boot-integration.md#defining-workflows-and-steps)) **Multiple beans of the same class:** | Situation | Instance name | |-----------|---------------| | Only one bean of the class | null string (default) | | Multiple beans — the `@Primary` one | null string (default) | | Multiple beans — non-primary | Spring bean name | ### @TransactionalStep Module: `dev.dbos:transact-spring-txstep-starter` (separate from `transact-spring-boot-starter`). Package: `dev.dbos.transact.spring.txstep`. ```java public @interface TransactionalStep { String name() default ""; Isolation isolationLevel() default Isolation.DEFAULT; } ``` Marks a Spring-managed method as a step factory step. The behaviour depends on calling context: | Context | Behaviour | |---------|-----------| | Inside a `@Workflow`, not inside a step | Full step factory behaviour: runs in a `REQUIRES_NEW` transaction, output written to `tx_step_outputs` atomically with user database work. On workflow retry the recorded output is replayed without re-executing the method body. | | Outside a workflow, or inside any step (including another `@TransactionalStep`) | Behaves like `@Transactional`: runs with `PROPAGATION_REQUIRED` (joins an existing transaction or starts a new one), no DBOS checkpoint recorded. | The annotated method must be called through a Spring proxy — calls via `this` bypass the aspect. The step automatically retries on PostgreSQL serialization failures (SQL state `40001`) and deadlocks (`40P01`) when called from within a workflow. **Parameters:** - **name**: Stable name for this step within the workflow. Defaults to the method name. - **isolationLevel**: Transaction isolation level for this step. Uses `org.springframework.transaction.annotation.Isolation`. Defaults to `Isolation.DEFAULT`, which leaves the datasource/connection-pool default unchanged. Example: `@TransactionalStep(isolationLevel = Isolation.SERIALIZABLE)`. **Supported stacks**: Spring JDBC / `JdbcTemplate`, JDBI (`jdbi3-spring`), jOOQ (`spring-boot-starter-jooq`), JPA / Hibernate. #### Installation ```kotlin title="build.gradle.kts" implementation("dev.dbos:transact-spring-boot-starter:") implementation("dev.dbos:transact-spring-txstep-starter:") ``` ```xml title="pom.xml" dev.dbos transact-spring-boot-starter VERSION dev.dbos transact-spring-txstep-starter VERSION ``` #### Auto-Configured Beans | Bean | Description | |------|-------------| | `TransactionalStepAspect` | AOP aspect that intercepts `@TransactionalStep` calls and delegates to `TransactionalStepFactory`. | | `TransactionalStepFactory` | Manages step lifecycle: idempotency check, transaction, and output recording. | | `TransactionalStepRegistrar` | Scans the Spring context post-startup and calls `factory.initialize()` (creates `tx_step_outputs`) only when annotated methods are found — no DB contact for apps that don't use the annotation. | All beans are `@ConditionalOnMissingBean`. #### Configuration | Property | Default | Description | |----------|---------|-------------| | `dbos.txstep.schema` | DBOS system schema | PostgreSQL schema for the `tx_step_outputs` table | #### How it works 1. `TransactionalStepAspect` intercepts every `@TransactionalStep` call and delegates to `TransactionalStepFactory`. 2. The factory calls `DBOS.runStep()`, which checks `tx_step_outputs` for a prior result. If one exists, it is returned immediately (idempotent replay). 3. Otherwise, a `REQUIRES_NEW` Spring transaction is started. The method body runs, and the result is written to `tx_step_outputs` using `DataSourceUtils.getConnection()` — the same connection the transaction holds. 4. The transaction commits, making the user's write and the step output record atomic. If the method throws, the transaction rolls back and the error is recorded separately so retries can replay it. See the [Step Factory tutorial](../tutorials/step-factory-tutorial.md#transactionalstep) for per-stack examples (JDBC, JDBI, jOOQ, JPA). ### DBOSAspect An `@Aspect` that intercepts: - **`@Workflow` methods** — routes the call through `dbos.integration().runWorkflow(...)` for durable execution and recovery. - **`@Step` methods** — delegates to `dbos.runStep(...)` when called inside a workflow context; executes directly otherwise. Spring AOP intercepts only calls that go through the Spring proxy. `this.someStep()` calls inside a workflow body are **not** intercepted. Inject a self-reference to ensure step calls are durable: ```java @Service public class MyService { @Autowired MyService self; @Workflow public String myWorkflow() { return self.myStep(); // intercepted — durable // this.myStep(); // NOT intercepted — runs outside DBOS } @Step public String myStep() { return "result"; } } ``` --- ## Workflows & Steps(Reference) ### Annotations #### @Workflow ```java public @interface Workflow { String name(); int maxRecoveryAttempts(); SerializationStrategy serializationStrategy(); } ``` An annotation that can be applied to a class method to mark it as a durable workflow. :::info Workflow methods must be invoked via the proxy object returned by [`registerProxy`](#registerproxy) in order to be durable. ::: **Parameters:** - **name**: The workflow name. Must be unique within the class. Defaults to method name if not provided. - **maxRecoveryAttempts**: Optionally configure the maximum number of times execution of a workflow may be attempted. This acts as a [dead letter queue](https://en.wikipedia.org/wiki/Dead_letter_queue) so that a buggy workflow that crashes its application (for example, by running it out of memory) does not do so infinitely. If a workflow exceeds this limit, its status is set to `MAX_RECOVERY_ATTEMPTS_EXCEEDED` and it may no longer be executed. - **serializationStrategy**: The default [serialization strategy](../reference/methods.md#serialization-strategy) to use for local invocations of this workflow. Set to `SerializationStrategy.PORTABLE` to test [cross-language interoperability](../../explanations/portable-workflows.md). Defaults to `SerializationStrategy.DEFAULT`. :::tip When a workflow uses portable serialization, Java automatically coerces JSON arguments to match the method's parameter types. For example, JSON integers are widened to `long`, and ISO-8601 date strings are parsed to `Instant` or `OffsetDateTime`. See [Input Validation and Coercion](#input-validation-and-coercion) below for details. ::: #### @Step ```java public @interface Step { String name(); int maxAttempts(); double intervalSeconds(); double backOffRate(); Class shouldRetry(); } ``` An annotation that can be applied to a class method to mark it as a step in a durable workflow. :::info Reminder, step methods must be invoked via the proxy object returned by [`registerProxy`](#registerproxy) in order to be durable. ::: **Parameters:** - **name**: The step name. Must be unique within the class. Defaults to method name if not provided. - **maxAttempts**: Maximum number of times this step is retried on failure. Must be greater than zero. Defaults to one. - **intervalSeconds**: Initial delay between retries in seconds. Must be positive. Defaults to one second. - **backOffRate**: Exponential backoff multiplier between retries. Must be greater than or equal to one. Defaults to two. - **shouldRetry**: A class implementing `StepShouldRetry` that decides per-exception whether to retry. The class must have a public no-arg constructor. Defaults to always retrying. Use this to short-circuit retries for non-transient errors without exhausting `maxAttempts`. See [`StepShouldRetry`](#stepshouldretry). ##### StepShouldRetry ```java @FunctionalInterface public interface StepShouldRetry { boolean shouldRetry(Throwable e); } ``` Implement this interface and pass the class to `@Step(shouldRetry = MyPolicy.class)` to control whether a given exception should trigger a retry. Return `true` to retry, `false` to fail immediately. The class must be `public` and have a public no-arg constructor. **Example:** ```java public class NoRetryOnValidation implements StepShouldRetry { @Override public boolean shouldRetry(Throwable e) { return !(e instanceof ValidationException); } } // Usage: @Step(maxAttempts = 5, shouldRetry = NoRetryOnValidation.class) public String fetchData(String id) { ... } ``` #### @WorkflowClassName ```java public @interface WorkflowClassName { String value(); } ``` An annotation applied to a workflow **implementation class** (not the interface) to assign it a stable, portable class name for workflow registration. Without this annotation, workflows are registered under their fully-qualified Java class name (e.g., `com.example.MyServiceImpl`). Using `@WorkflowClassName` replaces that with a shorter, language-agnostic name that survives refactoring and enables cross-language interoperability. **Example:** ```java @WorkflowClassName("MyService") public class MyServiceImpl implements MyService { @Workflow public String processOrder(String orderId) { ... } } ``` This workflow is registered as `processOrder/MyService/` instead of `processOrder/com.example.MyServiceImpl/`. :::tip Use `@WorkflowClassName` whenever a workflow may be invoked from another language (Python, TypeScript) or when you want workflow IDs to remain stable across package renames. ::: ### Input Validation and Coercion Java automatically coerces portable JSON arguments to match the workflow method's parameter types. No opt-in is required—this happens transparently for all portable workflows. Common coercions include: | JSON Type | Java Method Parameter | Coercion | |-----------|----------------------|----------| | Integer (`30000`) | `long` | Integer → long | | Double (`1.01`) | `double` | Double → double | | String (`"2025-06-15T10:30:00Z"`) | `Instant` | ISO-8601 parsing | | String (`"2025-06-15T10:30:00+02:00"`) | `OffsetDateTime` | ISO-8601 parsing | | JSON array | `ArrayList` | Element-wise coercion | | JSON object | `Map` | LinkedHashMap → Map | If coercion fails (for example, a JSON object where a `String` is expected), the workflow is marked as `ERROR` with a descriptive message. ```java // This workflow expects (String, long), but portable JSON delivers (String, Integer). // Java coerces the Integer to long automatically. @Workflow public String processOrder(String orderId, long quantity) { return "order:" + orderId + " qty:" + quantity; } ``` For more context on why input coercion matters for cross-language workflows, see [Input Validation and Coercion](../../explanations/portable-workflows.md#input-validation-and-coercion). ### Methods #### registerProxy ```java T registerProxy(Class interfaceClass, T implementation) T registerProxy(Class interfaceClass, T implementation, String instanceName) ``` Register the workflows in a class, returning a proxy object from which the class methods may be invoked as durable workflows. All workflows must be registered before DBOS is launched. :::info `transact-spring-boot-starter` handles proxy creation and workflow registration automatically. If you're using `transact-spring-boot-starter`, you _don't_ need to call registerProxy manually. ::: **Example Syntax:** ```java interface Example { public void workflow(); } class ExampleImpl implements Example { @Workflow public void workflow() { return; } } DBOS dbos = new DBOS(config); Example proxy = dbos.registerProxy(Example.class, new ExampleImpl()); dbos.launch(); proxy.workflow(); ``` **Parameters:** - **interfaceClass**: The interface class whose workflows are to be registered. - **implementation**: An instance of the class whose workflows to register. - **instanceName**: A unique name for this class instance. Use only when you are creating multiple instances of a class and your workflow depends on class instance variables. When DBOS needs to recover a workflow belonging to that class, it looks up the class instance using `instanceName` so it can recover the workflow using the right instance of its class. #### startWorkflow ```java WorkflowHandle startWorkflow(ThrowingSupplier workflow) WorkflowHandle startWorkflow(ThrowingSupplier workflow, StartWorkflowOptions options) WorkflowHandle startWorkflow(ThrowingRunnable workflow) WorkflowHandle startWorkflow(ThrowingRunnable workflow, StartWorkflowOptions options) ``` Start a workflow in the background and return a handle to it. Optionally enqueue it on a DBOS queue. The `startWorkflow` method resolves after the workflow is durably started; at this point the workflow is guaranteed to run to completion even if the app is interrupted. **Example Syntax**: ```java interface Example { public void workflow(); } class ExampleImpl implements Example { @Workflow public void workflow() { return; } } DBOS dbos = new DBOS(config); Example proxy = dbos.registerProxy(Example.class, new ExampleImpl()); dbos.launch(); dbos.startWorkflow(() -> proxy.workflow(), new StartWorkflowOptions()); ``` #### StartWorkflowOptions `StartWorkflowOptions` is a with-based configuration record for parameterizing `dbos.startWorkflow`. All fields are optional. **Constructors:** ```java new StartWorkflowOptions() ``` Create workflow options with all fields set to their defaults. ```java new StartWorkflowOptions(String workflowId) ``` Shortcut for `new StartWorkflowOptions().withWorkflowId(workflowId)`. Note that the `String` argument is a **workflow ID**, not a queue name: `new StartWorkflowOptions("my-queue")` starts the workflow with ID `my-queue` rather than enqueueing it. To enqueue, use the `QueueName` constructor below. ```java new StartWorkflowOptions(QueueName queue) ``` Shortcut for `new StartWorkflowOptions().withQueue(queue)`. [`QueueName`](./queues.md#queuename) wraps a queue's name so it cannot be confused with a workflow ID; create one with `QueueName.of("my-queue")`. ```java new StartWorkflowOptions(Queue queue) ``` *(deprecated since 1.1)* Use `new StartWorkflowOptions(QueueName.of(queue.name()))` instead. **Methods:** - **`withWorkflowId(String workflowId)`** - Set the workflow ID of this workflow. - **`withQueue(QueueName queue)`** / **`withQueue(String queueName)`** - Instead of starting the workflow directly, enqueue it on this queue. `withQueue(Queue queue)` is *(deprecated since 1.1)*; pass the queue's name instead. - **`withTimeout(Timeout timeout)`** - Set a timeout using a [`Timeout`](./methods.md#timeout) object. Use this overload to pass `Timeout.none()` (opt out of any inherited timeout) or `Timeout.inherit()` (explicitly inherit from the calling context). - **`withTimeout(Duration timeout)`** / **`withTimeout(long value, TimeUnit unit)`** - Set an explicit timeout duration for this workflow. When the timeout expires, the workflow **and all its children** are cancelled. Cancelling a workflow sets its status to `CANCELLED` and preempts its execution at the beginning of its next step. Timeouts are **start-to-completion**: if a workflow is enqueued, the timeout does not begin until the workflow is dequeued and starts execution. Also, timeouts are **durable**: they are stored in the database and persist across restarts, so workflows can have very long timeouts. Timeout deadlines are propagated to child workflows by default, so when a workflow's deadline expires all of its child workflows (and their children, and so on) are also cancelled. If you want to detach a child workflow from its parent's timeout, you can start it with its own explicit timeout (or `Timeout.none()`) to override the propagated timeout. - **`withNoTimeout()`** - Explicitly remove any inherited timeout or deadline from this workflow. - **`withDeadline(Instant deadline)`** - Set a deadline for this workflow. If the workflow is executing at the time of the deadline, the workflow **and all its children** are cancelled. Cancelling a workflow sets its status to `CANCELLED` and preempts its execution at the beginning of its next step. Deadlines are **durable**: they are stored in the database and persist across restarts. Deadlines are propagated to child workflows by default, so when a workflow's deadline expires all of its child workflows (and their children, and so on) are also cancelled. If you want to detach a child workflow from its parent's deadline, you can start it with a different explicit deadline. :::info An explicit timeout and deadline cannot both be set. ::: - **`withPriority(int priority)`** - May only be used when enqueuing. The priority of the enqueued workflow in the specified queue. Workflows with the same priority are dequeued in FIFO (first in, first out) order. Priority values can range from `0` to `2,147,483,647`, where a low number indicates a higher priority. A negative priority throws `IllegalArgumentException`. Workflows without assigned priorities have priority `0`, the highest priority. Priority works on every queue; no queue configuration is needed. - **`withDeduplicationId(String deduplicationId)`** - May only be used when enqueuing. At any given time, only one workflow with a specific deduplication ID can be enqueued in the specified queue. If a workflow with a deduplication ID is currently enqueued, delayed, or actively executing (status `ENQUEUED`, `PENDING`, or `DELAYED`), subsequent workflow enqueue attempts with the same deduplication ID in the same queue will raise an exception. - **`withQueuePartitionKey(String queuePartitionKey)`** - Set a queue partition key for the workflow. Use if and only if the queue is partitioned, which it is when any per-partition limit (`partitionConcurrency`, `partitionWorkerConcurrency`, or `partitionRateLimit`) is set, or when it was registered with the deprecated `partitionQueue` flag. Per-partition limits apply to each partition key separately; queue-wide limits still apply to the queue as a whole, except on a queue registered with the deprecated `partitionQueue` flag, where they apply to each partition instead. See [Partitioning Queues](../tutorials/queue-tutorial.md#partitioning-queues). :::info - Partition keys are required when enqueueing to a partitioned queue. - Partition keys cannot be used with non-partitioned queues. - Partition keys and deduplication IDs cannot be used together. ::: - **`withDelay(Duration delay)`** - Delay the start of the workflow by the specified duration after it is dequeued. Only applicable when enqueuing. - **`withAppVersion(String appVersion)`** - Tag the workflow with a specific application version, overriding the default version detected at runtime. - **`withAuthenticatedUser(String user)`** - Set the authenticated user for the workflow. - **`withAssumedRole(String role)`** - Set the assumed role for the workflow. - **`withAuthenticatedRoles(String... roles)`** - Set the list of authenticated roles. - **`withAuthentication(String user, String... roles)`** - Convenience method to set both `authenticatedUser` and `authenticatedRoles` in one call. - **`withAttributes(Map attributes)`** - Attach custom JSON-serializable key-value metadata to the workflow. #### WorkflowHandle ```java public interface WorkflowHandle { String workflowId(); T getResult() throws E; WorkflowStatus getStatus(); } ``` WorkflowHandle provides methods to interact with a running or completed workflow. The type parameters `T` and `E` represents the expected return type of the workflow and the checked exceptions it may throw. Handles can be used to wait for workflow completion, check status, and retrieve results. ##### WorkflowHandle.getResult ```java T getResult() throws E; ``` Wait for the workflow to complete and return its result. ##### WorkflowHandle.getStatus ```java WorkflowStatus getStatus(); ``` Retrieve the WorkflowStatus of the workflow. ##### WorkflowHandle.workflowId ```java String workflowId(); ``` Return the ID of the workflow underlying this handle. #### runStep ```java T runStep(ThrowingSupplier stepfunc, StepOptions opts) throws E T runStep(ThrowingSupplier stepfunc, String stepName) throws E void runStep(ThrowingRunnable stepfunc, StepOptions opts) throws E void runStep(ThrowingRunnable stepfunc, String stepName) throws E ``` Run a function as a step. If called from within a workflow, the result is durably stored. Returns the output of the step. **Example Syntax:** ```java class ExampleImpl implements Example { private final DBOS dbos; public ExampleImpl(DBOS dbos) { this.dbos = dbos; } private void stepOne() { System.out.println("Step one completed!"); } private void stepTwo() { System.out.println("Step two completed!"); } @Workflow public void workflow() throws InterruptedException { dbos.runStep(() -> stepOne(), "stepOne"); dbos.runStep(() -> stepTwo(), new StepOptions("stepTwo").withMaxAttempts(3)); } } ``` #### StepOptions `StepOptions` is a with-based configuration record for parameterizing `dbos.runStep`. All fields except step name are optional. **Constructors:** ```java new StepOptions(String name) ``` Create step options and provide a name for this step. By default the step runs once (no retries). **Methods:** - **`withMaxAttempts(int n)`** - Maximum number of times to attempt the step. Must be greater than zero. Defaults to `1` (no retries). Set to a value greater than `1` to enable retries on exception. - **`withRetryInterval(Duration interval)`** - How long to wait before the first retry. Must be positive. Defaults to 1 second. - **`withBackoffRate(double rate)`** - Exponential backoff multiplier between retries. Must be greater than or equal to 1.0. Defaults to 2.0. - **`withShouldRetry(Predicate shouldRetry)`** - Supply a predicate that decides per-exception whether to retry. Return `true` to retry, `false` to fail immediately without consuming remaining attempts. Equivalent to the annotation `shouldRetry` parameter but accepts a lambda directly. See [`StepShouldRetry`](#stepshouldretry). #### WorkflowOptions When workflow functions are called directly, they take their options from the current DBOS context. Setting options into the context is done with a `WorkflowOptions` object and a `setContext()` `try` block on the calling thread: ```java try (var _opts = new WorkflowOptions(wfId).setContext()) { // This workflow is within the `try` and will get `wfId` from the context result = workflowClass.workflowMethod(args); } ``` Workflow options will be restored to prior values at the end of the `try` block. :::info When using background execution or queues, workflow options are passed as a `StartWorkflowOptions` argument to `startWorkflow`. `WorkflowOptions` context does not affect `startWorkflow`. ::: **Constructors:** ```java new WorkflowOptions() ``` Create workflow options with no workflow ID and no timeout. ```java new WorkflowOptions(String workflowId) ``` Shortcut for `new WorkflowOptions().withWorkflowId(workflowId)`. **Fields:** - **workflowId**: The ID to be assigned to a workflow called within the `try` block - **timeout**: The timeout to be assigned to all workflows called within the `try` block - **deadline**: The deadline to be assigned to all workflows called within the `try` block **Methods:** - **`withWorkflowId(String workflowId)`** - Set the [workflow ID](../tutorials/workflow-tutorial.md#workflow-ids-and-idempotency) of the next workflow run. - **`withTimeout(Timeout timeout)`** / **`withTimeout(Duration timeout)`** / **`withTimeout(long value, TimeUnit unit)`** - Set a timeout for all enclosed workflow invocations. When the timeout expires, the workflow **and all its children** are cancelled. Timeouts are **start-to-completion**: the timeout does not begin until the workflow starts execution. Timeouts are also **durable**: they persist across restarts. Timeout deadlines are propagated to child workflows by default. - **`withNoTimeout()`** - Explicitly opt out of any inherited timeout. The workflow called within the `try` block will run without a timeout regardless of any timeout in the surrounding context. - **`withDeadline(Instant deadline)`** - Set an absolute deadline for all enclosed workflow invocations. At the deadline time, the workflow **and all its children** are cancelled. Deadlines are propagated to child workflows by default. :::info An explicit timeout and deadline cannot both be set. ::: - **`withAuthenticatedUser(String user)`** - Set the authenticated user for the workflow. Stored in the workflow record and accessible via `DBOS.authenticatedUser()`. - **`withAssumedRole(String role)`** - Set the assumed role for the workflow. Accessible via `DBOS.assumedRole()`. - **`withAuthenticatedRoles(String... roles)`** - Set the list of authenticated roles. Accessible via `DBOS.authenticatedRoles()`. - **`withAuthentication(String user, String... roles)`** - Convenience method to set both `authenticatedUser` and `authenticatedRoles` in one call. - **`withAttributes(Map attributes)`** - Attach custom JSON-serializable key-value metadata to the workflow. Stored in the workflow record and searchable via `ListWorkflowsInput.withAttributes(Map)`. Pass `null` or an empty map to set no attributes. ### Step Factories A step factory commits your database work and the DBOS step checkpoint in the **same transaction**, making the step exactly-once even for database writes. On workflow retry, the recorded output is replayed without re-executing the callback. Step factories must be constructed before `dbos.launch()`. Each constructor immediately opens a connection to verify the datasource is PostgreSQL and to create the `tx_step_outputs` table in the target schema if it does not already exist. **Common constructor parameters:** - **dbos**: The DBOS runtime instance. - **schema**: The PostgreSQL schema to use for `tx_step_outputs`. Defaults to the DBOS system schema from configuration. - **serializer**: The serializer to use for step outputs. Defaults to the serializer from `dbos` configuration. Every step factory method takes a `stepName` used for observability. Steps are identified by their sequential position within the workflow execution, not by name. See the [Step Factory tutorial](../tutorials/step-factory-tutorial.md) for full examples. #### JdbcStepFactory Module: `dev.dbos:transact` (core module, no extra dependency). Package: `dev.dbos.transact.txstep`. ```java new JdbcStepFactory(DBOS dbos, DataSource dataSource) new JdbcStepFactory(DBOS dbos, DataSource dataSource, String schema) new JdbcStepFactory(DBOS dbos, DataSource dataSource, DBOSSerializer serializer) new JdbcStepFactory(DBOS dbos, DataSource dataSource, String schema, DBOSSerializer serializer) ``` ##### txStep ```java R txStep(TransactionalFunction callback, String stepName) throws X R txStep(TransactionalFunction callback, StepFactoryOptions options) throws X void txStep(TransactionalRunnable callback, String stepName) throws X void txStep(TransactionalRunnable callback, StepFactoryOptions options) throws X ``` Executes `callback` as an idempotent DBOS step inside a JDBC transaction. If a result is already recorded for this step (e.g. on workflow retry), the callback is skipped and the cached result is returned. The callback receives an open `Connection` with autocommit disabled. It must **not** call `commit`, `rollback`, or `close` — the factory manages the transaction lifecycle. The step automatically retries on PostgreSQL serialization failures (SQL state `40001`) and deadlocks (`40P01`). ##### StepFactoryOptions ```java new StepFactoryOptions(String name) new StepFactoryOptions(String name, IsolationLevel isolationLevel) ``` Options for `txStep`. Pass instead of a bare `String` step name to control transaction isolation. **`IsolationLevel`** enum (`dev.dbos.transact.txstep`): - `DEFAULT` — leave the connection pool's default isolation level unchanged - `READ_COMMITTED` - `REPEATABLE_READ` - `SERIALIZABLE` - `READ_UNCOMMITTED` **Example:** ```java factory.txStep(conn -> { // critical section — use serializable isolation conn.prepareStatement("INSERT INTO orders(id) VALUES (?)").executeUpdate(); return orderId; }, new StepFactoryOptions("insertOrder", IsolationLevel.SERIALIZABLE)); ``` **`TransactionalFunction`** — database work that returns a result: ```java @FunctionalInterface interface TransactionalFunction { R execute(Connection conn) throws X; } ``` **`TransactionalRunnable`** — database work with no return value: ```java @FunctionalInterface interface TransactionalRunnable { void execute(Connection conn) throws X; } ``` --- #### JdbiStepFactory Module: `dev.dbos:transact-jdbi-step-factory`. Package: `dev.dbos.transact.jdbi`. ```java new JdbiStepFactory(DBOS dbos, Jdbi jdbi) new JdbiStepFactory(DBOS dbos, Jdbi jdbi, String schema) new JdbiStepFactory(DBOS dbos, Jdbi jdbi, DBOSSerializer serializer) new JdbiStepFactory(DBOS dbos, Jdbi jdbi, String schema, DBOSSerializer serializer) ``` ##### inStep ```java R inStep(HandleCallback callback, String stepName) throws X R inStep(HandleCallback callback, JdbiStepOptions options) throws X ``` Executes `callback` as an idempotent DBOS step inside a JDBI transaction, returning a result. The callback receives an open `Handle`. It must **not** call `commit`, `rollback`, or `close`. The step automatically retries on PostgreSQL serialization failures (SQL state `40001`) and deadlocks (`40P01`). ##### useStep ```java void useStep(HandleConsumer callback, String stepName) throws X void useStep(HandleConsumer callback, JdbiStepOptions options) throws X ``` Void variant of `inStep`. Accepts a `HandleConsumer` for callers that do not need to return a value. ##### JdbiStepOptions ```java new JdbiStepOptions(String name) new JdbiStepOptions(String name, TransactionIsolationLevel isolationLevel) ``` Options for `inStep`/`useStep`. Pass instead of a bare `String` step name to control transaction isolation using JDBI's `TransactionIsolationLevel` enum (`org.jdbi.v3.core.transaction`). --- #### JooqStepFactory Module: `dev.dbos:transact-jooq-step-factory`. Package: `dev.dbos.transact.jooq`. ```java new JooqStepFactory(DBOS dbos, DSLContext dsl) new JooqStepFactory(DBOS dbos, DSLContext dsl, String schema) new JooqStepFactory(DBOS dbos, DSLContext dsl, DBOSSerializer serializer) new JooqStepFactory(DBOS dbos, DSLContext dsl, String schema, DBOSSerializer serializer) ``` ##### txStepResult ```java T txStepResult(TransactionalCallable callback, String stepName) ``` Executes `callback` as an idempotent DBOS step inside a jOOQ transaction, returning a result. The callback receives a jOOQ `Configuration` with an open transaction. It must **not** commit or close the underlying connection. ##### txStep ```java void txStep(TransactionalRunnable transactional, String stepName) ``` Void variant of `txStepResult`. Accepts a jOOQ `TransactionalRunnable` for callers that do not need to return a value. --- #### @TransactionalStep Spring Boot annotation from the `transact-spring-txstep-starter` module. See [`@TransactionalStep`](./spring-boot-starter.md#transactionalstep) in the Spring Boot Starter reference. --- ## Using DBOS with Kotlin DBOS works with Kotlin out of the box — all Java APIs are accessible from Kotlin. The `transact` artifact also ships Kotlin extension functions that make the most common calls idiomatic by placing the lambda last, enabling Kotlin's trailing lambda syntax. ### Dependency The Kotlin extensions are included in the main `transact` artifact alongside the Java code; no extra dependency is needed. **Gradle** ```kotlin dependencies { implementation("dev.dbos:transact:1.1.0") } ``` **Maven** ```xml dev.dbos transact 1.1.0 ``` ### Writing Workflows and Steps Define your interface and implementation exactly as you would in Java, using the same `@Workflow` annotation: ```kotlin interface OrderService { fun processOrder(orderId: String): String } class OrderServiceImpl(private val dbos: DBOS) : OrderService { private lateinit var self: OrderService fun setSelf(proxy: OrderService) { self = proxy } @Workflow override fun processOrder(orderId: String): String { val result = dbos.runStep("fetchOrder") { fetchFromApi(orderId) // trailing lambda — idiomatic Kotlin } dbos.runStep("saveOrder") { saveToDatabase(result) } return result } } ``` The `dbos.runStep(name) { ... }` form is a Kotlin extension function that puts the lambda last, avoiding the `ThrowingSupplier` wrapping you'd need in Java. ### Starting Workflows Use `dbos.startWorkflow(options) { ... }` with trailing lambda syntax: ```kotlin val dbos = DBOS(config) val service = dbos.registerProxy(OrderService::class.java, OrderServiceImpl(dbos)) dbos.launch() // Start a workflow in the background val handle = dbos.startWorkflow(StartWorkflowOptions()) { service.processOrder("order-123") } // Or with a specific workflow ID val handle = dbos.startWorkflow(StartWorkflowOptions("my-workflow-id")) { service.processOrder("order-456") } val result = handle.result ``` :::info Kotlin's SAM conversion means a plain `dbos.startWorkflow { }` call (with no options argument) would be ambiguous with the Java overload. Always pass `StartWorkflowOptions()` (or `null`) as the first argument when using the trailing lambda form. ::: ### Step Options Pass a `StepOptions` object when you need retry configuration: ```kotlin val result = dbos.runStep(StepOptions("fetchOrder").withMaxAttempts(3)) { fetchFromApi(orderId) } ``` ### Registration and Lifecycle Registration and lifecycle are identical to Java: ```kotlin val config = DBOSConfig.defaultsFromEnv("my-app") .withAppVersion("0.1.0") val dbos = DBOS(config) val service = dbos.registerProxy(OrderService::class.java, OrderServiceImpl(dbos).also { it.setSelf(/* proxy set below */) }) dbos.launch() // DBOS is AutoCloseable dbos.use { service.processOrder("order-789") } ``` --- ## Queues & Concurrency(Tutorials) You can use queues to run many workflows at once with managed concurrency. Queues provide _flow control_, letting you manage how many workflows run at once or how often workflows are started. Register a queue with [`dbos.registerQueue`](../reference/queues.md#dbosregisterqueue), specifying its name and options. Queues must be registered **after** [`dbos.launch()`](../reference/lifecycle.md). Queue configuration is persisted to the system database, so queues are visible to every DBOS process connected to that database. ```java dbos.launch(); dbos.registerQueue("example-queue", QueueOptions.empty()); ``` You can then enqueue any workflow using [`withQueue`](../reference/workflows-steps.md#startworkflow) when calling `startWorkflow`. Enqueuing a workflow submits it for execution and returns a [handle](../reference/workflows-steps.md#workflowhandle) to it. Queued tasks are started in [priority](#priority) order, and in first-in, first-out (FIFO) order among tasks of the same priority. ```java class ExampleImpl implements Example { @Workflow public String processTask(String task) { // Process the task... return "Processed: " + task; } } public String example(DBOS dbos, String queue, Example proxy, String task) throws Exception { // Enqueue a workflow WorkflowHandle handle = dbos.startWorkflow( () -> proxy.processTask(task), new StartWorkflowOptions().withQueue(queue) ); // Get the result String result = handle.getResult(); System.out.println("Task result: " + result); return result; } ``` :::tip `new StartWorkflowOptions("example-queue")` sets the workflow **ID**, not the queue. To name a queue in the constructor, pass a [`QueueName`](../reference/queues.md#queuename): `new StartWorkflowOptions(QueueName.of("example-queue"))`. ::: #### Queue Example Here's an example of a workflow using a queue to process tasks concurrently: ```java interface Example { public String taskWorkflow(String task); public List queueWorkflow(String[] tasks) throws Exception; } class ExampleImpl implements Example { private final DBOS dbos; private String queueName; private Example self; public ExampleImpl(DBOS dbos) { this.dbos = dbos; } public void setQueueName(String queueName) { this.queueName = queueName; } public void setSelf(Example self) { this.self = self; } @Workflow public String taskWorkflow(String task) { // Process the task... return "Processed: " + task; } @Workflow public List queueWorkflow(String[] tasks) throws Exception { // Enqueue each task so all tasks are processed concurrently List> handles = new ArrayList<>(); for (String task : tasks) { WorkflowHandle handle = dbos.startWorkflow( () -> self.taskWorkflow(task), new StartWorkflowOptions().withQueue(queueName) ); handles.add(handle); } // Wait for each task to complete and retrieve its result List results = new ArrayList<>(); for (WorkflowHandle handle : handles) { String result = handle.getResult(); results.add(result); } return results; } } public class App { public static void main(String[] args) throws Exception { DBOSConfig config = ... DBOS dbos = new DBOS(config); // Instantiate an Example and register its workflows ExampleImpl impl = new ExampleImpl(dbos); Example proxy = dbos.registerProxy(Example.class, impl); // Provide the workflow proxy to the class so its methods can invoke workflows impl.setSelf(proxy); dbos.launch(); // Register the queue dbos.registerQueue("example-queue", QueueOptions.empty()); impl.setQueueName("example-queue"); // Run the queue workflow String[] tasks = {"task1", "task2", "task3", "task4", "task5"}; List results = proxy.queueWorkflow(tasks); for (String result : results) { System.out.println(result); } } } ``` Sometimes, you may wish to receive the result of each task as soon as it's ready instead of waiting for all tasks to complete. You can do this using [`send` and `recv`](./workflow-communication.md#workflow-messaging-and-notifications). Each enqueued workflow sends a message to the main workflow when it's done processing its task. The main workflow awaits those messages, retrieving the result of each task as soon as the task completes. ```java interface Example { public String processTask(String parentWorkflowId, int taskId, String task); public List processTasks(String[] tasks) throws Exception; } class ExampleImpl implements Example { private static final String TASK_COMPLETE_TOPIC = "task_complete"; private final DBOS dbos; private final String queueName; private Example self; public ExampleImpl(DBOS dbos, String queueName) { this.dbos = dbos; this.queueName = queueName; } public void setSelf(Example self) { this.self = self; } @Workflow public String processTask(String parentWorkflowId, int taskId, String task) { String result = "Processed: " + task; // Process the task // Notify the main workflow this task is complete dbos.send(parentWorkflowId, taskId, TASK_COMPLETE_TOPIC, null); return result; } @Workflow public List processTasks(String[] tasks) throws Exception { String parentWorkflowId = DBOS.workflowId(); List> handles = new ArrayList<>(); for (int i = 0; i < tasks.length; i++) { final int taskId = i; final String task = tasks[i]; WorkflowHandle handle = dbos.startWorkflow( () -> self.processTask(parentWorkflowId, taskId, task), new StartWorkflowOptions().withQueue(queueName) ); handles.add(handle); } List results = new ArrayList<>(); while (results.size() < tasks.length) { // Wait for a notification that a task is complete Integer completedTaskId = dbos.recv(TASK_COMPLETE_TOPIC, Duration.ofMinutes(5)).orElse(null); if (completedTaskId == null) { throw new RuntimeException("Timeout waiting for task completion"); } // Retrieve result of the completed task WorkflowHandle completedTaskHandle = handles.get(completedTaskId); String result = completedTaskHandle.getResult(); System.out.println("Task " + completedTaskId + " completed. Result: " + result); results.add(result); } return results; } } ``` ##### Enqueueing from Another Application Often, you want to enqueue a workflow from another DBOS application or from outside your DBOS application. For example, let's say you have an API server and a data processing service. You're using DBOS to build a durable data pipeline in the data processing service. When the API server receives a request, it should enqueue the data pipeline for execution on the data processing service. If both applications [share a system database](../../explanations/sharing-a-system-database.md), a DBOS application can enqueue another application's workflow with `dbos.enqueueWorkflow`. Because the workflow is implemented elsewhere, workflow and queue metadata must be specified explicitly with [`EnqueueOptions`](../reference/client.md#enqueueoptions), and `withApplicationName` names the application that should run it: ```java var options = new EnqueueOptions( "dataPipeline", // Workflow name "com.example.DataPipelineImpl", // Class name QueueName.of("pipelineQueue") // Queue name ).withApplicationName("data-processing-service"); WorkflowHandle handle = dbos.enqueueWorkflow( options, new Object[]{"task-123", "data"} // Workflow arguments ); ``` When the target takes named arguments, such as a Python workflow with keyword arguments, set `withSerialization(SerializationStrategy.PORTABLE)` on the options and call `dbos.enqueueWorkflow(options, positionalArgs, namedArgs)`. From outside any DBOS application, use the [DBOS Client](../reference/client.md), which connects directly to the system database and takes the same `EnqueueOptions`: ```java var client = new DBOSClient(dbUrl, dbUser, dbPassword); var options = new EnqueueOptions( "dataPipeline", // Workflow name "com.example.DataPipelineImpl", // Class name QueueName.of("pipelineQueue") // Queue name ).withApplicationName("data-processing-service"); var handle = client.enqueueWorkflow( options, new Object[]{"task-123", "data"} // Workflow arguments ); // Optionally wait for the result Object result = handle.getResult(); ``` ##### Enqueueing from PL/pgSQL You can also enqueue a workflow from a PostgreSQL trigger or stored procedure. The DBOS System Database includes an [`enqueue_workflow`](../../explanations/system-tables.md#dbosenqueue_workflow) method for this scenario. For example, here is the previous example of enqueing the `dataPipeline` workflow on the `pipelineQueue` queue with arguments, but using PL/pgSQL. ```sql DECLARE workflow_id text; workflow_id := dbos.enqueue_workflow( workflow_name => 'dataPipeline', class_name => 'com.example.DataPipelineImpl', queue_name => 'pipelineQueue', positional_args => ARRAY[ '"task-123"'::json, '"data"'::json ] ); ``` #### Managing Concurrency You can control how many workflows from a queue run simultaneously by configuring concurrency limits. This helps prevent resource exhaustion when workflows consume significant memory or processing power. ##### Worker Concurrency Worker concurrency sets the maximum number of workflows from a queue that can run concurrently on a single DBOS process. This is particularly useful for resource-intensive workflows to avoid exhausting the resources of any process. For example, this queue has a worker concurrency of 5, so each process will run at most 5 workflows from this queue simultaneously: ```java dbos.registerQueue("example-queue", QueueOptions.setWorkerConcurrency(5)); ``` ##### Global Concurrency Global concurrency limits the total number of workflows from a queue that can run concurrently across all DBOS processes in your application. For example, this queue will have a maximum of 10 workflows running simultaneously across your entire application. :::warning Worker concurrency limits are recommended for most use cases. Take care when using a global concurrency limit as any `PENDING` workflow on the queue counts toward the limit, including workflows from previous application versions. ::: ```java dbos.registerQueue("example-queue", QueueOptions.setConcurrency(10)); ``` #### Rate Limiting You can set _rate limits_ for a queue, limiting the number of workflows that it can start in a given period. Rate limits are global across all DBOS processes using this queue. For example, this queue has a limit of 100 workflows with a period of 60 seconds, so it may not start more than 100 workflows in 60 seconds: ```java dbos.registerQueue("example-queue", QueueOptions.setRateLimit(100, 60, TimeUnit.SECONDS)); ``` Rate limits are especially useful when working with a rate-limited API. #### Reconfiguring Queues at Runtime Because queue configuration lives in the system database, you can change a queue's configuration at runtime without redeploying or restarting your workers. Use `dbos.updateQueue` to modify a queue's configuration. Workers pick up the new configuration on their next polling iteration. ```java // Change the queue's concurrency dbos.updateQueue("example-queue", QueueOptions.setConcurrency(20)); // Change its rate limit dbos.updateQueue("example-queue", QueueOptions.setRateLimit(25, 30, TimeUnit.SECONDS)); ``` :::warning If your application calls `dbos.registerQueue` on startup, the next process to start can overwrite settings you applied at runtime via `updateQueue`. Either update the `registerQueue` call to match the new configuration, or pass `QueueConflictResolution.NEVER_UPDATE` to preserve the runtime changes. ::: You can also find, list, and delete queues: ```java // Find a specific queue by name Optional queue = dbos.findQueue("example-queue"); // List all queues in the system database List queues = dbos.listQueues(); // Delete a queue boolean deleted = dbos.deleteQueue("example-queue"); ``` You can do all of this from a [`DBOSClient`](../reference/client.md#queue-management-methods) as well, which is useful for managing queues from an admin tool or another service. ```java var client = new DBOSClient(dbUrl, dbUser, dbPassword); // Register or update a queue client.registerQueue("example-queue", QueueOptions.setConcurrency(10)); // Register only if it doesn't already exist client.registerQueue("example-queue", QueueOptions.setConcurrency(10), QueueConflictResolution.NEVER_UPDATE); // Update only the concurrency of an existing queue client.updateQueue("example-queue", QueueOptions.setConcurrency(20)); // Find a queue by name Optional queue = client.findQueue("example-queue"); // List all queues List queues = client.listQueues(); // Delete a queue boolean deleted = client.deleteQueue("example-queue"); ``` #### Setting Timeouts You can set a timeout for an enqueued workflow via the `withTimeout` function on `StartWorkflowOptions`. When the timeout expires, the workflow **and all its children** are cancelled. Cancelling a workflow sets its status to `CANCELLED` and preempts its execution at the beginning of its next step. Timeouts are **start-to-completion**: a workflow's timeout does not begin until the workflow is dequeued and starts execution. Also, timeouts are **durable**: they are stored in the database and persist across restarts, so workflows can have very long timeouts. Example syntax: ```java // use StartWorkflowOptions.withTimeout with dbos.startWorkflow var options = new StartWorkflowOptions().withQueue("example-queue").withTimeout(Duration.ofSeconds(10)); var handle = dbos.startWorkflow(() -> proxy.workflow(), options); ``` #### Setting Deadlines You can set a deadline for an enqueued workflow via the `withDeadline` function on `StartWorkflowOptions`. A deadline is an **absolute point in time** by which the workflow must complete; if the deadline passes before the workflow finishes, the workflow **and all its children** are cancelled. Cancelling a workflow sets its status to `CANCELLED` and preempts its execution at the beginning of its next step. :::warning You cannot set both an explicit timeout and a deadline on the same workflow — use one or the other. ::: Like timeouts, deadlines are **durable**: they are stored in the database and persist across restarts. Example syntax: ```java // Use StartWorkflowOptions.withDeadline with dbos.startWorkflow Instant deadline = Instant.now().plus(Duration.ofHours(1)); var options = new StartWorkflowOptions().withQueue("example-queue").withDeadline(deadline); var handle = dbos.startWorkflow(() -> proxy.workflow(), options); ``` #### Delaying Execution You can delay when an enqueued workflow starts executing via the `withDelay` function on `StartWorkflowOptions`. The delay is a `Duration` that must be a positive, non-zero value. The workflow remains `DELAYED` until the delay has elapsed, after which its status changes to `ENQUEUED` and it is eligible to be dequeued and started normally. Example syntax: ```java // Delay the workflow's execution by 30 seconds var options = new StartWorkflowOptions().withQueue("example-queue").withDelay(Duration.ofSeconds(30)); var handle = dbos.startWorkflow(() -> proxy.workflow(), options); ``` #### Partitioning Queues You can **partition** queues to distribute work across dynamically created queue partitions. A queue is partitioned if you register it with any per-partition flow control limit: | Option | Meaning | | --- | --- | | `partitionConcurrency` | Maximum workflows from any one partition running at once across all processes. | | `partitionWorkerConcurrency` | Maximum workflows from any one partition running at once on a single process. | | `partitionRateLimit` | Maximum workflows that may be started from any one partition in a given period. | When you enqueue a workflow on a partitioned queue, you must supply a queue partition key. Essentially, you can think of each partition as a "subqueue" you dynamically create by enqueueing a workflow with a partition key. For example, suppose you want your users to each be able to run at most one task at a time. You can do this with a queue whose `partitionConcurrency` is 1, where the partition key is user ID. **Example Syntax** ```java dbos.registerQueue("example-queue", QueueOptions.setPartitionConcurrency(1)); void onUserTaskSubmission(String userID, Task task) { // Partition the task queue by user ID. As the queue has a // per-partition concurrency of 1, this means that at most one // task can run at once per user (but tasks from different // users can run concurrently). var options = new StartWorkflowOptions().withQueue("example-queue").withQueuePartitionKey(userID); dbos.startWorkflow(() -> proxy.taskWorkflow(task), options); } ``` :::warning Every enqueue on a partitioned queue must supply a partition key. `dbos.startWorkflow` throws if you omit it, but a workflow enqueued without a partition key by other means (such as a [`DBOSClient`](../reference/client.md)) stays `ENQUEUED` and is never dequeued. For the same reason, partitioning a queue that already has enqueued workflows strands them; drain the queue first. Workflows stranded this way can be moved to a queue that is not partitioned with [`dbos.resumeWorkflow(workflowId, queueName)`](../reference/methods.md#resumeworkflow). ::: ##### Combining Queue-Wide and Per-Partition Limits A partitioned queue enforces its per-partition limits **and** its queue-wide limits ([`concurrency`](#global-concurrency), [`workerConcurrency`](#worker-concurrency), and [`rateLimit`](#rate-limiting)) at the same time. This lets you protect your workers from overload while still fairly distributing work between partitions. For example, this "fair queue" runs at most one task per user, but no more than 10 tasks on any single process: ```java dbos.registerQueue("fair-queue", QueueOptions.setPartitionConcurrency(1).andWorkerConcurrency(10)); ``` Each queue-wide limit has a per-partition counterpart, so you can mix and match them freely: ```java // At most 100 tasks running globally and 25 running per tenant, // at most 10 tasks running per process and 2 per tenant per process, // and at most 1000 tasks started per minute globally and 50 per tenant. dbos.registerQueue("tenant-queue", QueueOptions.setConcurrency(100) .andWorkerConcurrency(10) .andRateLimit(1000, Duration.ofSeconds(60)) .andPartitionConcurrency(25) .andPartitionWorkerConcurrency(2) .andPartitionRateLimit(50, Duration.ofSeconds(60))); ``` When both are set, each per-partition concurrency limit must be less than or equal to its queue-wide counterpart, and `partitionWorkerConcurrency` must be less than or equal to `partitionConcurrency`; limits that are not set are not compared. See [`QueueOptions`](../reference/queues.md#queueoptions) for the full set of rules. :::note [Deduplication](#deduplication) is not supported on partitioned queues. ::: :::info Deprecated `partitionQueue` option Before 1.1, queues were partitioned with `QueueOptions.setPartitionQueue(true)` (or `andPartitionQueue(true)`), which applies the queue-wide limits to each partition instead of to the queue as a whole. This option is deprecated since 1.1: set per-partition limits instead. The limits of a queue registered with `partitionQueue` cannot be changed with `updateQueue`; re-register it with per-partition limits. ::: #### Deduplication You can set a deduplication ID for an enqueued workflow using [`withQueue`](../reference/workflows-steps.md#startworkflow) when calling `startWorkflow`. At any given time, only one workflow with a specific deduplication ID can be enqueued in the specified queue. If a workflow with a deduplication ID is currently enqueued or actively executing (status `DELAYED`, `ENQUEUED`, or `PENDING`), subsequent workflow enqueue attempts with the same deduplication ID in the same queue will raise an exception. For example, this is useful if you only want to have one workflow active at a time per user—set the deduplication ID to the user's ID. **Example syntax:** ```java @Workflow public String taskWorkflow(String task) { // Process the task... return "completed"; } public void example(DBOS dbos, Example proxy, String task, String userID) throws Exception { // Use user ID for deduplication WorkflowHandle handle = dbos.startWorkflow( () -> proxy.taskWorkflow(task), new StartWorkflowOptions().withQueue("example-queue").withDeduplicationId(userID) ); String result = handle.getResult(); System.out.println("Workflow completed: " + result); } ``` #### Priority You can set a priority for an enqueued workflow using [`withQueue`](../reference/workflows-steps.md#startworkflow) when calling `startWorkflow`. Workflows with the same priority are dequeued in **FIFO (first in, first out)** order. Priority values can range from `0` to `2,147,483,647`, where **a low number indicates a higher priority**. A negative priority throws `IllegalArgumentException`. Priority is enabled on every queue; no extra configuration is needed. :::tip Workflows without assigned priorities have priority `0`, the highest priority. ::: **Example syntax:** ```java @Workflow public String taskWorkflow(String task) { // Process the task... return "completed"; } public void example(DBOS dbos, Example proxy, String task, int priority) throws Exception { WorkflowHandle handle = dbos.startWorkflow( () -> proxy.taskWorkflow(task), new StartWorkflowOptions().withQueue("example-queue").withPriority(priority) ); String result = handle.getResult(); System.out.println("Workflow completed: " + result); } ``` #### Explicit Queue Listening By default, a process running DBOS listens to (dequeues workflows from) all queues registered in its system database. However, sometimes you only want a process to listen to a specific list of queues. You use the `withListenQueues` method on [DBOSConfig](../reference/lifecycle.md#dbosconfig) to explicitly tell a process running DBOS to only listen to a specific set of queues. This is particularly useful when managing heterogeneous workers, where specific tasks should execute on specific physical servers. For example, say you have a mix of CPU workers and GPU workers and you want CPU tasks to only execute on CPU workers and GPU tasks to only execute on GPU workers. You can create separate queues for CPU and GPU tasks and configure each type of worker to only listen to the appropriate queue: ```java var workerType = System.getenv("WORKER_TYPE"); // "cpu" or "gpu" var config = DBOSConfig.defaults("my-dbos-app") .withAppVersion("0.1.0"); if (workerType.equals("gpu")) { config = config.withListenQueues("gpuQueue"); } else if (workerType.equals("cpu")) { config = config.withListenQueues("cpuQueue"); } DBOS dbos = new DBOS(config); // register workflows... dbos.launch(); dbos.registerQueue("cpuQueue", QueueOptions.empty()); dbos.registerQueue("gpuQueue", QueueOptions.empty()); ``` Note that `withListenQueues` only controls what workflows are dequeued, not what workflows can be enqueued, so you can freely enqueue tasks onto the GPU queue from a CPU worker for execution on a GPU worker, and vice versa. --- ## Scheduled Workflows You can schedule DBOS [workflows](./workflow-tutorial.md) to run automatically on a cron schedule. Scheduled workflows are **exactly-once**: DBOS assigns each firing a deterministic workflow ID derived from the schedule name and scheduled time, so even if your application restarts mid-execution, each scheduled invocation runs exactly once. Each time a scheduled fires, its workflow is executed by exactly one worker process. ### Declaring Schedules with `applySchedules` The recommended way to declare schedules is to call `dbos.applySchedules()` once after `dbos.launch()`. This atomically creates or replaces the named schedules, so your code is always the source of truth: ```java DBOS dbos = new DBOS(config); Example proxy = dbos.registerProxy(Example.class, new ExampleImpl(dbos)); dbos.launch(); dbos.applySchedules( new WorkflowSchedule("every-minute", "everyMinute", "com.example.ExampleImpl", "0 * * * * *"), new WorkflowSchedule("daily-report", "dailyReport", "com.example.ExampleImpl", "0 0 9 * * *") .withCronTimezone(ZoneId.of("America/New_York")) ); ``` A workflow invoked by a `WorkflowSchedule` must accept exactly two arguments: an `Instant` for the scheduled fire time and an `Object` for the optional context attached via `withContext(...)`: ```java @Workflow public void everyMinute(Instant scheduled, Object context) { // scheduled: the exact cron fire time (used for the workflow ID) // context: the value passed via withContext(), or null if not set } ``` `applySchedules` is idempotent: re-running it on every startup always results in the declared set of schedules, with no duplicates. When `applySchedules` updates an existing schedule, it replaces the entire definition with the new declaration, so any optional setting left unset is cleared. For example, if a schedule was routed to a named queue and you re-apply it without `withQueueName(...)`, it reverts to the default scheduler queue. The schedule's status and last-fired time are preserved. #### WorkflowSchedule Parameters ```java new WorkflowSchedule(String scheduleName, String workflowName, String className, String cron) ``` - **scheduleName**: A unique name for this schedule (used for management operations). - **workflowName**: The name of the workflow method to invoke (as registered, or as set by `@Workflow(name=...)`). - **className**: The fully-qualified class name, or the short name set by [`@WorkflowClassName`](../reference/workflows-steps.md#workflowclassname). - **cron**: A [Spring 5.3+ CronExpression](https://docs.spring.io/spring-framework/docs/current/javadoc-api/org/springframework/scheduling/support/CronExpression.html). :::info https://www.spring-cron-generator.net/ is an online tool composing and understanding Spring cron expressions. ::: Common optional configuration via `with` methods: | Method | Description | |--------|-------------| | `withCronTimezone(ZoneId)` | Interpret the cron in this timezone (default: UTC). | | `withAutomaticBackfill(true)` | Retroactively start any firings missed while the app was down. | | `withQueueName(String)` | Enqueue executions on this queue instead of the default scheduler queue. | | `withStatus(ScheduleStatus.PAUSED)` | Create the schedule in a paused state. | | `withContext(Object)` | Attach a serializable context object passed to the workflow. | | `withApplicationName(String)` | The application that owns the schedule and runs its workflows, when several applications [share a system database](../../explanations/sharing-a-system-database.md). Defaults to the creating application. | :::info Scheduled runs always record their inputs with the application's configured serializer. A `serializationStrategy` declared on the scheduled workflow's `@Workflow` annotation is ignored for scheduled runs, whether they are fired by the cron, [backfilled](#backfill), or [triggered](#triggering-a-schedule-immediately). ::: ### Runtime Schedule Management Schedules can also be created, paused, resumed, and deleted at runtime: ```java // Create a schedule at runtime (throws if name already exists) dbos.createSchedule( new WorkflowSchedule("on-demand", "processReport", "com.example.ReportImpl", "0 0 * * * *")); // Pause and resume dbos.pauseSchedule("daily-report"); dbos.resumeSchedule("daily-report"); // Inspect Optional s = dbos.getSchedule("every-minute"); List active = dbos.listSchedules( List.of(ScheduleStatus.ACTIVE), null, null); // Delete dbos.deleteSchedule("on-demand"); ``` ### Backfill If your application was down for a period and you need to retroactively run the scheduled workflows that were missed, use `backfillSchedule`: ```java Instant start = Instant.parse("2025-01-01T00:00:00Z"); Instant end = Instant.parse("2025-01-02T00:00:00Z"); List> handles = dbos.backfillSchedule("every-minute", start, end); ``` DBOS uses the same deterministic workflow IDs as the live scheduler, so executions that already ran are skipped. Backfills (manual or automatic) compute missed executions using the schedule's **current** cron expression. If you update a schedule's cron expression and then backfill, the backfill generates one execution per tick of the new expression over the requested window—including times the old expression would never have matched. For example, changing a daily schedule to an hourly one and then backfilling yesterday enqueues 24 executions, not 1. You can also enable **automatic backfill** on a schedule so DBOS does this for you on every startup: ```java new WorkflowSchedule("every-minute", "everyMinute", "com.example.ExampleImpl", "0 * * * * *") .withAutomaticBackfill(true) ``` ### Triggering a Schedule Immediately To fire a scheduled workflow immediately outside its normal cadence: ```java WorkflowHandle handle = dbos.triggerSchedule("daily-report"); ``` --- ## Spring Boot Integration The `transact-spring-boot-starter` package integrates DBOS into a Spring Boot application with zero boilerplate: add the dependency, configure a datasource, annotate your beans, and DBOS launches automatically alongside the Spring context. :::tip The Java [Widget Store demo](https://github.com/dbos-inc/dbos-demo-apps/tree/main/java/widget-store) is a fully featured DBOS Spring Boot application that illustrates the features described in this documentation. ::: ### Adding the Dependency **Gradle** ```kotlin dependencies { implementation("dev.dbos:transact-spring-boot-starter:1.1.0") } ``` **Maven** ```xml dev.dbos transact-spring-boot-starter 1.1.0 ``` ### Configuration Instead of creating the [DBOSConfig](../reference/lifecycle.md#dbosconfig) for your app programmatically, the Spring Boot Starter package allows you to configure DBOS via Spring Boot [External Application Properties](https://docs.spring.io/spring-boot/reference/features/external-config.html#features.external-config.files). For example, you can set your application name and database connection in `application.yaml` (or `application.properties`): ```yml dbos: application: name: "my-app" # DBOS system database datasource: url: "jdbc:postgresql://localhost:5432/my_app_db" username: "postgres" password: "${PGPASSWORD}" ``` If `dbos.application.name` or `dbos.datasource` properties are not set, `spring.application.name` and `spring.datasource` are used as a fallback. ```yml spring: application: name: "my-app" datasource: url: "jdbc:postgresql://localhost:5432/my_app_db" username: "postgres" password: "${PGPASSWORD}" driver-class-name: "org.postgresql.Driver" ``` :::danger DBOS only supports PostgreSQL today. Attempting to use a non PostgreSQL database driver will throw an exception. ::: For a full list of DBOSConfig fields you can set via application properties, please see the [reference documentation](../reference/spring-boot-starter.md#configuration-properties). #### Programmatic Configuration Declare a `DBOSConfigCustomizer` bean to adjust the auto-configured `DBOSConfig` without replacing it: ```java @Bean public DBOSConfigCustomizer myCustomizer() { return config -> config .withAdminServer(true) .withAdminServerPort(8081) .withEnablePatching(true); } ``` Multiple `DBOSConfigCustomizer` beans are applied in `@Order` / `Ordered` order. To replace the config entirely, declare your own `@Bean DBOSConfig`. ### Defining Workflows and Steps In "vanilla" DBOS, you need to define an interface and a class in order use `@Workflow` and `@Step`. With Spring Boot Starter, you can use the `@Workflow` and `@Step` annotations on methods on any Spring-managed singleton bean (`@Service`, `@Component`, etc.). `transact-spring-boot-starter` uses [Spring AOP](https://docs.spring.io/spring-framework/reference/core/aop.html) instead of dynamic proxies, which do not require defining a separate interface. With Spring Boot Starter, you also don't need to manually register workflow and step proxies with [`dbos.registerProxy`](../reference/workflows-steps.md#registerproxy). The `DBOSWorkflowRegistrar` automatically scans all singleton beans after initialization and registers `@Workflow` methods automatically. Note, that you do however still need a proxy instance for invocating `@Workflow` and `@Step` methods on the same instance. This proxy reference can be self injected via an [`@Autowired`](https://docs.spring.io/spring-framework/reference/core/beans/annotation-config/autowired.html) setter. ```java @Service public class OrderService { private OrderService self; // self reference is injected automatically by Spring Boot Dependency Injection @Autowired @Lazy public void setSelf(OrderService self) { this.self = self; } @Workflow public String processOrder(String orderId) { // Invoke @Step via self reference to perform durable execution book keeping String result = self.chargeCard(orderId); self.sendConfirmation(orderId, result); return result; } @Step public String chargeCard(String orderId) { /* ... */ return "charged"; } @Step public void sendConfirmation(String orderId, String result) { /* ... */ } } ``` #### Multiple Beans of the Same Class When multiple beans of the same class exist, the `@Primary` one is registered under the default (empty) instance name. Additional beans of the same class are registered as [Workflow Class Instances](./workflow-classes.md) using their Spring bean name. To target a specific bean, call `startWorkflow` through that bean, or pass its bean name as the instance name to the `EnqueueOptions` constructor: `new EnqueueOptions(workflowName, className, beanName, QueueName.of(queue))`. ### Lifecycle `DBOSLifecycle` is a `SmartLifecycle` bean that calls `dbos.launch()` after all singletons are initialized and `dbos.shutdown()` when the context closes, so workflow beans are always registered before launch. Schedules and [queues](../reference/queues.md) are stored in the system database and can only be registered once DBOS is launched, so register them in an `ApplicationListener` handler, which Spring runs after `DBOSLifecycle` has started: ```java @Component public class QueueSetup implements ApplicationListener { @Autowired DBOS dbos; @Override public void onApplicationEvent(ContextRefreshedEvent event) { dbos.registerQueue("orders", QueueOptions.setWorkerConcurrency(5)); } } ``` ### Injecting `dbos` instance The auto-configured `DBOS` bean is available for injection anywhere in the application: ```java // Constructor injection @Service public class WidgetStoreService { private final DBOS dbos; private final WidgetStoreRepository repo; public WidgetStoreService(DBOS dbos, WidgetStoreRepository widgetStoreRepo) { this.dbos = dbos; this.repo = widgetStoreRepo; } } // @Autowired field injection @Service public class ScheduleSetup implements ApplicationListener { @Autowired DBOS dbos; @Autowired OrderService orderServiceProxy; @Override public void onApplicationEvent(ContextRefreshedEvent event) { dbos.applySchedules( new WorkflowSchedule("daily-orders", "processOrder", "com.example.OrderService", "0 0 9 * * *") ); } } ``` --- ## Transactional Steps Regular DBOS steps checkpoint their output _after_ the step body completes. If the application crashes after your database write but before the checkpoint is saved, the step runs again on recovery — potentially writing to the database twice. **Transactional step factories** solve this by committing the step output and your database work in the **same transaction**. On retry, DBOS finds the recorded output and returns it without re-executing, making the step exactly-once even for database writes. :::info Transactional step factories require the `tx_step_outputs` table in your application database. This table is created automatically at startup, but your database user must have `CREATE TABLE` privileges in the configured schema. If you use a restricted database role in production, grant the necessary privileges or create the table manually before deploying. ::: ### Choosing an approach | Situation | Recommendation | |-----------|----------------| | Plain JDBC / `DataSource` | [`JdbcStepFactory`](#jdbcstepfactory) from the `transact` module | | JDBI 3 | [`JdbiStepFactory`](#jdbistepfactory) from `transact-jdbi-step-factory` | | jOOQ | [`JooqStepFactory`](#jooqstepfactory) from `transact-jooq-step-factory` | | Spring Boot app | [`@TransactionalStep`](#transactionalstep) from `transact-spring-txstep-starter` | --- ### JdbcStepFactory `JdbcStepFactory` is included in the core `transact` module — no extra dependency needed. Construct it once (before `dbos.launch()`) and call `txStep` inside any `@Workflow` method: ```java JdbcStepFactory factory = new JdbcStepFactory(dbos, dataSource); ``` Inside a workflow, pass a lambda that receives an open `Connection`. Do **not** call `commit` or `close` on the connection — the factory manages the transaction. ```java class OrderWorkflowImpl implements OrderWorkflow { private final JdbcStepFactory factory; public OrderWorkflowImpl(DBOS dbos, DataSource dataSource) { this.factory = new JdbcStepFactory(dbos, dataSource); } @Workflow public String processOrder(String orderId) throws Exception { return factory.txStep(conn -> { try (var stmt = conn.prepareStatement( "INSERT INTO orders(id) VALUES (?)")) { stmt.setString(1, orderId); stmt.executeUpdate(); } return orderId; }, "insertOrder"); } } ``` Use the void overload when the step doesn't need to return a value: ```java factory.txStep(conn -> { conn.prepareStatement("DELETE FROM staging WHERE id = ?") .setString(1, id) .executeUpdate(); }, "deleteStaging"); ``` #### Isolation level To override the transaction isolation level, pass a `StepFactoryOptions` instead of a plain step name: ```java import dev.dbos.transact.txstep.IsolationLevel; import dev.dbos.transact.txstep.StepFactoryOptions; factory.txStep(conn -> { // ... return result; }, new StepFactoryOptions("insertOrder", IsolationLevel.SERIALIZABLE)); ``` Available levels: `DEFAULT` (pool default), `READ_COMMITTED`, `REPEATABLE_READ`, `SERIALIZABLE`, `READ_UNCOMMITTED`. Steps automatically retry on PostgreSQL serialization failures (`40001`) and deadlocks (`40P01`), so `SERIALIZABLE` isolation can be used safely with retries already configured. --- ### JdbiStepFactory Add the dependency: ```kotlin title="build.gradle.kts" implementation("dev.dbos:transact-jdbi-step-factory:") ``` ```xml title="pom.xml" dev.dbos transact-jdbi-step-factory VERSION ``` Construct with a `Jdbi` instance and call `inStep` (with return value) or `useStep` (void) inside workflows. The lambda receives an open `Handle` — do **not** call `commit` or `close` on it. ```java JdbiStepFactory factory = new JdbiStepFactory(dbos, jdbi); @Workflow public String processOrder(String orderId) { return factory.inStep(handle -> { handle.createUpdate("INSERT INTO orders(id) VALUES (:id)") .bind("id", orderId) .execute(); return orderId; }, "insertOrder"); } ``` Void variant: ```java factory.useStep(handle -> { handle.createUpdate("DELETE FROM staging WHERE id = :id") .bind("id", id) .execute(); }, "deleteStaging"); ``` #### Isolation level To override the transaction isolation level, pass a `JdbiStepOptions` instead of a plain step name: ```java import dev.dbos.transact.jdbi.JdbiStepOptions; import org.jdbi.v3.core.transaction.TransactionIsolationLevel; factory.inStep(handle -> { // ... return result; }, new JdbiStepOptions("insertOrder", TransactionIsolationLevel.SERIALIZABLE)); ``` Steps automatically retry on PostgreSQL serialization failures (`40001`) and deadlocks (`40P01`). --- ### JooqStepFactory Add the dependency: ```kotlin title="build.gradle.kts" implementation("dev.dbos:transact-jooq-step-factory:") ``` ```xml title="pom.xml" dev.dbos transact-jooq-step-factory VERSION ``` Construct with a `DSLContext` and call `txStepResult` (with return value) or `txStep` (void) inside workflows. The lambda receives a jOOQ `Configuration` with an open transaction — do **not** commit or close the connection. ```java JooqStepFactory factory = new JooqStepFactory(dbos, dslContext); @Workflow public String processOrder(String orderId) { return factory.txStepResult(trx -> { trx.dsl().execute("INSERT INTO orders(id) VALUES (?)", orderId); return orderId; }, "insertOrder"); } ``` Void variant: ```java factory.txStep(trx -> { trx.dsl().execute("DELETE FROM staging WHERE id = ?", id); }, "deleteStaging"); ``` --- ### `@TransactionalStep` If you are using Spring Boot, the `transact-spring-txstep-starter` module provides `@TransactionalStep` — an annotation that turns any Spring-managed method into a step factory step with no lambda wrapping required. #### Installation Add both the DBOS Spring Boot starter and this module: ```kotlin title="build.gradle.kts" implementation("dev.dbos:transact-spring-boot-starter:") implementation("dev.dbos:transact-spring-txstep-starter:") ``` ```xml title="pom.xml" dev.dbos transact-spring-boot-starter VERSION dev.dbos transact-spring-txstep-starter VERSION ``` #### Usage Annotate any Spring-managed method with `@TransactionalStep`. The method must be called through a Spring proxy — inject the bean into a `@Workflow`-annotated method in another Spring bean. When called from inside a `@Workflow` (and not already inside a step), the full step factory behaviour applies: the method runs in a `REQUIRES_NEW` transaction and the output is checkpointed atomically. When called from outside a workflow, or from inside any step (including another `@TransactionalStep`), it behaves like `@Transactional` — the transaction runs normally with `PROPAGATION_REQUIRED` but no DBOS checkpoint is recorded. This makes `@TransactionalStep` methods safe to call from any context. Steps automatically retry on PostgreSQL serialization failures (`40001`) and deadlocks (`40P01`) when called from within a workflow. To override the transaction isolation level: ```java @TransactionalStep(isolationLevel = Isolation.SERIALIZABLE) public Order saveOrder(Order order) { return orderRepository.save(order); } ``` `Isolation` is `org.springframework.transaction.annotation.Isolation`. ##### JdbcTemplate No extra dependencies needed. Spring Boot auto-configures `DataSourceTransactionManager` and `JdbcTemplate`. ```java @Service public class OrderStepService { @Autowired JdbcTemplate jdbc; @TransactionalStep public Order saveOrder(Order order) { jdbc.update("INSERT INTO orders(id, item, qty) VALUES (?, ?, ?)", order.id(), order.item(), order.qty()); return order; } } @Service public class OrderWorkflowService { @Autowired OrderStepService steps; @Workflow public Order processOrder(Order order) { return steps.saveOrder(order); } } ``` ##### JDBI Add `jdbi3-spring` so JDBI's `SpringTransactionHandler` reuses the active Spring transaction: ```kotlin implementation("org.jdbi:jdbi3-spring:") ``` Configure JDBI as a Spring bean: ```java @Configuration public class JdbiConfig { @Bean public Jdbi jdbi(DataSource dataSource) throws Exception { var factory = new JdbiFactoryBean(dataSource); factory.afterPropertiesSet(); return factory.getObject(); } } ``` Then annotate your step method as usual: ```java @Service public class OrderStepService { @Autowired Jdbi jdbi; @TransactionalStep public Order saveOrder(Order order) { jdbi.withHandle(h -> h.execute("INSERT INTO orders(id, item, qty) VALUES (?, ?, ?)", order.id(), order.item(), order.qty())); return order; } } ``` ##### jOOQ Spring Boot auto-configures `DSLContext` with `SpringTransactionProvider` when you add `spring-boot-starter-jooq`: ```kotlin implementation("org.springframework.boot:spring-boot-starter-jooq") ``` Set `spring.jooq.sql-dialect=POSTGRES` in your `application.properties`, then inject `DSLContext` directly: ```java @Service public class OrderStepService { @Autowired DSLContext dsl; @TransactionalStep public Order saveOrder(Order order) { dsl.execute("INSERT INTO orders(id, item, qty) VALUES (?, ?, ?)", order.id(), order.item(), order.qty()); return order; } } ``` ##### JPA / Hibernate Spring Boot auto-configures `JpaTransactionManager` when `spring-boot-starter-data-jpa` is present: ```kotlin implementation("org.springframework.boot:spring-boot-starter-data-jpa") ``` ```java @Service public class OrderStepService { @Autowired OrderRepository repo; // Spring Data JPA repository @TransactionalStep public Order saveOrder(Order order) { return repo.save(order); } } ``` #### Configuration | Property | Default | Description | |----------|---------|-------------| | `dbos.txstep.schema` | DBOS system schema | PostgreSQL schema for the `tx_step_outputs` table. Defaults to `DBOSConfig.databaseSchema` if unspecified. | The `tx_step_outputs` table is created lazily on startup — only if at least one `@TransactionalStep` method is found in the Spring context. Applications that never use the annotation incur no database contact. #### How it works 1. `TransactionalStepAspect` intercepts every `@TransactionalStep` call and delegates to `TransactionalStepFactory`. 2. The factory calls `DBOS.runStep()`, which checks `tx_step_outputs` for a prior result. If one exists, it is returned immediately (idempotent replay). 3. Otherwise, a `REQUIRES_NEW` Spring transaction is started. The method body runs, and the result is written to `tx_step_outputs` using `DataSourceUtils.getConnection()` — the same connection the transaction holds. 4. The transaction commits, making the user's write and the step output record atomic. If the method throws, the transaction rolls back and the error is recorded separately so retries can replay it. --- ## Steps(Tutorials) When using DBOS workflows, you should call any method that performs complex operations or accesses external APIs or services as a _step_. If a workflow is interrupted, upon restart it automatically resumes execution from the **last completed step**. A step can return any serializable value and may throw checked or unchecked exceptions. DBOS provides two ways to declare steps: [`runStep`](../reference/workflows-steps.md#runstep) for inline lambdas, and the [`@Step`](../reference/workflows-steps.md#step) annotation for named methods. ### runStep Use [`runStep`](../reference/workflows-steps.md#runstep) to run an inline lambda as a checkpointed step directly inside a workflow method. Here's a simple example: ```java class ExampleImpl implements Example { private final DBOS dbos; public ExampleImpl(DBOS dbos) { this.dbos = dbos; } @Workflow public int workflowFunction(int n) { int randomNumber = dbos.runStep( () -> ThreadLocalRandom.current().nextInt(n), // generate a random number as a checkpointed step "generateRandomNumber" // A name for the step ); return randomNumber; } } ``` ### `@Step` Annotation Use the [`@Step`](../reference/workflows-steps.md#step) annotation to declare a method as a step. Annotated steps must be called through a DBOS proxy — calling them directly on the implementation bypasses DBOS and the call will not be checkpointed. You can define steps and workflows on the same interface. In that case, the workflow must call step methods through a proxy reference (often called `self`) rather than `this`: ```java interface Example { int workflowFunction(int n); int generateRandomNumber(int n); } class ExampleImpl implements Example { private Example self; // proxy reference for calling @Step methods public void setSelf(Example self) { this.self = self; } @Workflow public int workflowFunction(int n) { return self.generateRandomNumber(n); // must call through proxy, not this } @Step public int generateRandomNumber(int n) { return ThreadLocalRandom.current().nextInt(n); } } // Setup: ExampleImpl impl = new ExampleImpl(); Example proxy = dbos.registerProxy(Example.class, impl); impl.setSelf(proxy); ``` You can also put steps on a separate interface, which is useful when multiple workflows share the same set of steps: ```java interface StepService { int generateRandomNumber(int n); } class StepServiceImpl implements StepService { @Step public int generateRandomNumber(int n) { return ThreadLocalRandom.current().nextInt(n); } } class ExampleImpl implements Example { private final StepService steps; public ExampleImpl(StepService steps) { this.steps = steps; } @Workflow public int workflowFunction(int n) { return steps.generateRandomNumber(n); } } // Setup: StepService stepsProxy = dbos.registerProxy(StepService.class, new StepServiceImpl()); Example proxy = dbos.registerProxy(Example.class, new ExampleImpl(stepsProxy)); ``` ### When to Make Something a Step You should make a method a step if you're using it in a DBOS workflow and it performs a [**nondeterministic**](./workflow-tutorial.md#determinism) operation. A nondeterministic operation is one that may return different outputs given the same inputs. Common nondeterministic operations include: - Accessing an external API or service, like serving a file from [AWS S3](https://aws.amazon.com/s3/), calling an external API like [Stripe](https://stripe.com/), or accessing an external data store like [Elasticsearch](https://www.elastic.co/elasticsearch/). - Accessing files on disk. - Generating a random number. - Getting the current time. You **cannot** call, start, or enqueue workflows from within steps. These operations should be performed from workflow methods. You can call one step from another step, but the called step becomes part of the calling step's execution rather than functioning as a separate step. ### Configurable Retries You can optionally configure a step to automatically retry any error a set number of times with exponential backoff. This is useful for automatically handling transient failures, like making requests to unreliable APIs. #### With `runStep` Retries are configurable through `StepOptions` passed to `runStep`. Available options include: - [`withMaxAttempts`](../reference/workflows-steps.md#runstep) - Maximum number of times this step is automatically retried on failure. - [`withRetryInterval`](../reference/workflows-steps.md#runstep) - Initial delay between retries in seconds. - [`withBackoffRate`](../reference/workflows-steps.md#runstep) - Exponential backoff multiplier between retries. For example, let's write a step that fetches a website, and configure it to retry failures (such as if the site to be fetched is temporarily down) up to 10 times: ```java class ExampleImpl implements Example { private final DBOS dbos; public ExampleImpl(DBOS dbos) { this.dbos = dbos; } private String fetchStep(String url) throws Exception { HttpClient client = HttpClient.newHttpClient(); HttpRequest request = HttpRequest.newBuilder() .uri(URI.create(url)) .build(); HttpResponse response = client.send( request, HttpResponse.BodyHandlers.ofString() ); return response.body(); } @Workflow public String fetchWorkflow(String inputURL) throws Exception { return dbos.runStep( () -> fetchStep(inputURL), new StepOptions("fetchStep") .withMaxAttempts(10) .withRetryInterval(Duration.ofMillis(500)) .withBackoffRate(2.0) ); } } ``` #### With `@Step` The same retry options are available as annotation parameters: ```java interface Example { String fetchWorkflow(String inputURL) throws Exception; String fetchStep(String url) throws Exception; } class ExampleImpl implements Example { private Example self; public void setSelf(Example self) { this.self = self; } @Workflow public String fetchWorkflow(String inputURL) throws Exception { return self.fetchStep(inputURL); } @Step( maxAttempts = 10, intervalSeconds = 0.5, backOffRate = 2.0 ) public String fetchStep(String url) throws Exception { HttpClient client = HttpClient.newHttpClient(); HttpRequest request = HttpRequest.newBuilder() .uri(URI.create(url)) .build(); HttpResponse response = client.send( request, HttpResponse.BodyHandlers.ofString() ); return response.body(); } } ``` If a step exhausts all retry attempts, it throws an exception to the calling workflow. ### Selective Retry with `shouldRetry` By default, every exception triggers a retry (up to `maxAttempts`). Use `shouldRetry` to short-circuit retries for exceptions you know are non-transient — for example, validation errors or 4xx HTTP responses — without wasting attempts on them. #### With `runStep` Pass a `Predicate` via `StepOptions.withShouldRetry`: ```java @Workflow public String fetchWorkflow(String url) throws Exception { return dbos.runStep( () -> fetchStep(url), new StepOptions("fetchStep") .withMaxAttempts(5) .withShouldRetry(e -> { // retry on network errors; fail fast on HTTP 4xx if (e instanceof HttpClientResponseException ex) { return ex.getStatus().getCode() >= 500; } return true; }) ); } ``` #### With `@Step` Implement `StepShouldRetry` as a public class with a public no-arg constructor and reference it from the annotation: ```java public class RetryOnServerError implements StepShouldRetry { @Override public boolean shouldRetry(Throwable e) { if (e instanceof HttpClientResponseException ex) { return ex.getStatus().getCode() >= 500; } return true; } } interface Example { String fetchStep(String url) throws Exception; } class ExampleImpl implements Example { @Step(maxAttempts = 5, shouldRetry = RetryOnServerError.class) public String fetchStep(String url) throws Exception { // ... } } ``` ### Step Factories Regular steps checkpoint their output _after_ the step body completes. If your step writes to a database and the application crashes before the checkpoint is saved, the step runs again on recovery — potentially writing twice. If you need your database write and the DBOS checkpoint to be committed in the **same transaction**, use a step factory: - **Plain JDBC / JDBI / jOOQ**: use [`JdbcStepFactory`, `JdbiStepFactory`, or `JooqStepFactory`](../tutorials/step-factory-tutorial.md). - **Spring Boot**: annotate the method with [`@TransactionalStep`](../tutorials/step-factory-tutorial.md#transactionalstep) from `transact-spring-txstep-starter`. See the [Step Factory tutorial](../tutorials/step-factory-tutorial.md) for full details. --- ## Testing Workflows Because `DBOS` is a regular Java object injected into your workflow classes — not a global static — it can be mocked with any standard Java mocking library such as [Mockito](https://site.mockito.org/). This lets you test your workflow logic in complete isolation, without a PostgreSQL database. ### Unit Testing The key insight is that `dbos.runStep(lambda, name)` takes a lambda that is only executed by the real DBOS runtime. When `DBOS` is mocked, the lambda is never called — you control what the mock returns. This lets you test workflow branching logic (payment succeeds, inventory is low, etc.) by stubbing step outcomes directly. #### Setup Construct your workflow class directly, passing a mock `DBOS`: ```java import static org.mockito.Mockito.*; class CheckoutWorkflowTest { // Mockito struggles with generic ThrowingSupplier/ThrowingRunnable types, // so define typed helpers to avoid unchecked-cast warnings. private static ThrowingRunnable anyRunnable() { return ArgumentMatchers.any(); } private static ThrowingSupplier anySupplier() { return ArgumentMatchers.any(); } private DBOS mockDBOS; private WidgetStoreRepository mockRepo; private WidgetStoreService mockSelf; private WidgetStoreService service; @BeforeEach void setUp() { mockDBOS = mock(DBOS.class); mockRepo = mock(WidgetStoreRepository.class); mockSelf = mock(WidgetStoreService.class); service = new WidgetStoreService(mockDBOS, mockRepo); service.setSelf(mockSelf); // inject the mocked self-proxy } } ``` `mockRepo` and `mockSelf` are mocked separately. Because the workflow body calls them only via `dbos.runStep(() -> repo.someMethod(), "name")`, the lambdas are never executed against the mocks — but you can verify that the steps were invoked in the right order with the right names. #### Writing Tests Stub step return values and verify the workflow's sequence of DBOS calls: ```java @Test void checkoutWorkflow_paymentSuccessful_paysAndDispatchesOrder() throws Exception { int orderId = 42; when(mockDBOS.runStep(anySupplier(), eq("createOrder"))).thenReturn(orderId); when(mockDBOS.recv(eq(PAYMENT_STATUS), any())).thenReturn(Optional.of("paid")); service.checkoutWorkflow(); InOrder inOrder = Mockito.inOrder(mockDBOS); inOrder.verify(mockDBOS).runStep(anyRunnable(), eq("subtractInventory")); inOrder.verify(mockDBOS).runStep(anySupplier(), eq("createOrder")); inOrder.verify(mockDBOS).setEvent(eq(PAYMENT_ID), any()); inOrder.verify(mockDBOS).recv(eq(PAYMENT_STATUS), any()); inOrder.verify(mockDBOS).runStep(anyRunnable(), eq("markOrderPaid")); inOrder.verify(mockDBOS).startWorkflow(anyRunnable()); inOrder.verify(mockDBOS).setEvent(eq(ORDER_ID), eq(String.valueOf(orderId))); } @Test void checkoutWorkflow_insufficientInventory_setsNullPaymentIdAndReturns() throws Exception { doThrow(new RuntimeException("Insufficient Inventory")) .when(mockDBOS) .runStep(anyRunnable(), eq("subtractInventory")); service.checkoutWorkflow(); verify(mockDBOS).setEvent(eq(PAYMENT_ID), eq(null)); verify(mockDBOS, never()).runStep(anySupplier(), eq("createOrder")); } ``` :::tip Add an `@AfterEach` assertion that `mockRepo` and `mockSelf` had no direct interactions. This guards against workflow code accidentally calling them outside a `runStep` wrapper: ```java @AfterEach void verifyNoDirectCalls() { verifyNoInteractions(mockRepo, mockSelf); } ``` ::: #### Spring Boot With the Spring Boot starter, `DBOS` is injected by Spring and the workflow class is a `@Service`. You can test it the same way — construct the service directly with mocks rather than loading the Spring context: ```java @BeforeEach void setUp() { mockDBOS = mock(DBOS.class); service = new MyWorkflowService(mockDBOS, otherDependencies...); } ``` This is faster than `@SpringBootTest` and keeps the test free of container overhead. ### Integration Testing For tests that exercise real durable execution — recovery, exactly-once steps, queues — you need a live PostgreSQL database and a real `DBOS` instance. `DBOS` implements `AutoCloseable`, so use try-with-resources to guarantee shutdown: ```java @Test void workflow_resumesAfterInterruption() throws Exception { DBOSConfig config = DBOSConfig.defaults("test-app") .withAppVersion("0.1.0") .withDatabaseUrl(System.getenv("DBOS_TEST_JDBC_URL")) .withDbUser(System.getenv("PGUSER")) .withDbPassword(System.getenv("PGPASSWORD")); try (var dbos = new DBOS(config)) { var workflows = dbos.registerProxy(MyWorkflows.class, new MyWorkflowsImpl(dbos)); dbos.launch(); // exercise real durable behaviour workflows.myWorkflow("input"); } } ``` For test isolation, use a fresh database per test run. The easiest approach is [Testcontainers](https://testcontainers.com/) with the PostgreSQL module, which spins up a throwaway PostgreSQL container: ```java @Testcontainers class MyWorkflowIntegrationTest { @Container static final PostgreSQLContainer postgres = new PostgreSQLContainer<>("postgres:latest"); @Test void myWorkflow_completesSuccessfully() throws Exception { DBOSConfig config = DBOSConfig.defaults("test-app") .withAppVersion("0.1.0") .withDatabaseUrl(postgres.getJdbcUrl()) .withDbUser(postgres.getUsername()) .withDbPassword(postgres.getPassword()); try (var dbos = new DBOS(config)) { var workflows = dbos.registerProxy(MyWorkflows.class, new MyWorkflowsImpl(dbos)); dbos.launch(); workflows.myWorkflow("input"); } } } ``` --- ## Upgrading Workflow Code(Tutorials) A challenge encountered when operating long-running durable workflows in production is **how to deploy breaking changes without disrupting in-progress workflows.** A breaking change to a workflow is one that changes which steps are run, or the order in which the steps are run. If a breaking change was made to a workflow and that workflow is replayed by the recovery system, the checkpoints created by the previous version of the code may not match the steps called by the workflow in the new version of the code, causing recovery to fail. DBOS supports two strategies for safely upgrading workflow code: **patching** and **versioning**. ### Patching In patching, the result of a call to [`dbos.patch()`](../reference/methods.md#patch) is used to conditionally execute the new code. `dbos.patch()` returns `true` for new calls (those executing after the breaking change) and `false` for old calls (those that executed before the breaking change). Therefore, if `dbos.patch()` returns `true`, the workflow should follow the new code path, otherwise it must follow the prior codepath. To use patching, you must enable it in the configuration: ```java var config = DBOSConfig.defaults("my-app-name").withEnablePatching(); ``` For example, let's say our original workflow is: ```java @Workflow public int workflow() { dbos.runStep(() -> foo(), "foo"); dbos.runStep(() -> bar(), "bar"); } ``` We want to replace the call to `foo()` with a call to `baz()`. This is a breaking change because it changes what steps run. We can make this breaking change safely using a patch: ```java @Workflow public int workflow() { if (dbos.patch("use-baz")) { dbos.runStep(() -> baz(), "baz"); } else { dbos.runStep(() -> foo(), "foo"); } dbos.runStep(() -> bar(), "bar"); } ``` Now, new workflows will run `baz()`, while old workflows will reexecute `foo()`. #### Deprecating and Removing Patches Patches add complexity and runtime overhead; fortunately they don't need to stay in your code forever. Once all workflows that started before you deployed the patch are complete, you can safely remove patches from your code. :::tip You can use the [list workflows APIs](./workflow-management.md#listing-workflows) to see what workflows are still active. ::: First, you must deprecate the patch with [`dbos.deprecatePatch()`](../reference/methods.md#deprecatepatch) `dbos.deprecatePatch` must be used for a transition period prior to fully removing the patch, as it allows coexistence with any ongoing workflows that used `dbos.patch()`. For example, here's how to deprecate the patch above: ```java @Workflow public int workflow() { if (dbos.deprecatePatch("use-baz")) { dbos.runStep(() -> baz(), "baz"); } dbos.runStep(() -> bar(), "bar"); } ``` Then, when all workflows that started before you deprecated the patch are complete, you can remove the patch entirely: ```java @Workflow public int workflow() { dbos.runStep(() -> baz(), "baz"); dbos.runStep(() -> bar(), "bar"); } ``` If any mistakes happen during the process (a breaking change is not patched, or a patch is deprecated or removed prematurely), the workflow will throw a `DBOSUnexpectedStepError` pointing to the step where the problem occurred. #### How Patching Works Under the hood, when you call `dbos.patch()` from a workflow, it attempts to insert a "patch marker" at its current point in your workflow history (this is a new row in the `operation_outputs` table in your database). If it successfully inserts the patch marker or if the patch marker is already present, then the workflow should take the patch codepath. If there is already a record present in this point in your workflow history and it is not a patch marker, then the workflow must be old (it already continued past this point with old code), and `dbos.patch()` returns `false`. When you deprecate a patch with `dbos.deprecatePatch()`, new workflows no longer insert patch markers into their workflow history. However, if a workflow contains the patch marker in its history, it continues past that patch marker, safely ignoring it. Once all workflows with patch markers are complete, the patch may be safely removed. ### Versioning When using versioning, DBOS **versions** applications and workflows, and only continues workflow execution with the same application version that started the workflow. All workflows are tagged with the application version on which they started. We recommend setting the application version explicitly through configuration, and changing it whenever you deploy changed workflow code: ```java var config = DBOSConfig.defaults("my-app-name").withAppVersion("1.0.0"); ``` If you don't set a version, DBOS computes one from a hash of the DBOS SDK version, the application name, and the name, signature, and bytecode of each registered workflow method. The computed version is a fallback: it does not change when you change steps or other code your workflows call, and it changes whenever you upgrade the DBOS SDK or rename the application, even if your workflows did not change. When DBOS tries to recover workflows, it only recovers workflows whose version matches the current application version. This prevents recovery of workflows that depend on different code. When using versioning, we recommend **blue-green** code upgrades: - When deploying a new version of your code, launch new processes running your new code version, but retain some processes running your old code version. - Direct new traffic to your new processes while your old processes "drain" and complete all workflows of the old code version. - Then, once all workflows of the old version are complete (you can use [`dbos.listWorkflows`](../reference/methods.md#listworkflows) to check), you can retire the old code version. #### Application Version Management DBOS provides methods to inspect and manage registered application versions at runtime: ```java // List all versions seen by this database, newest first List versions = dbos.listApplicationVersions(); // Get the currently promoted "latest" version VersionInfo latest = dbos.getLatestApplicationVersion(); // Promote a version (useful during blue-green cutover) dbos.setLatestApplicationVersion("2.0.0"); ``` `VersionInfo` is a record with the following fields: - **versionId**: A generated unique ID for the version record. - **versionName**: The human-readable version string (e.g., `"2.0.0"`). - **versionTimestamp**: When this version was promoted. - **createdAt**: When the version record was first inserted. See [`listApplicationVersions`](../reference/methods.md#listapplicationversions), [`getLatestApplicationVersion`](../reference/methods.md#getlatestapplicationversion), and [`setLatestApplicationVersion`](../reference/methods.md#setlatestapplicationversion) for full API details. --- ## Workflows on Class Instances DBOS supports registering multiple instances of the same workflow implementation class under different names. This is useful when the same workflow logic should run against different configurations — for example, different API endpoints, tenant-specific credentials, or database connections — without duplicating code. ### The Problem Suppose you have a workflow that processes data from a remote service: ```java interface DataProcessor { void process(String jobId); } class DataProcessorImpl implements DataProcessor { private final String serviceUrl; public DataProcessorImpl(String serviceUrl) { this.serviceUrl = serviceUrl; } @Workflow public void process(String jobId) { dbos.runStep(() -> fetchAndStore(serviceUrl, jobId), "fetch-" + jobId); } } ``` If you need to run this workflow against two different service URLs, you need two instances — but both share the same `@Workflow` method. DBOS must know which instance to use when recovering an interrupted workflow. ### Registering Named Instances Give each instance a unique name using the three-argument `registerProxy` overload: ```java DBOS dbos = new DBOS(config); DataProcessor processorA = dbos.registerProxy( DataProcessor.class, new DataProcessorImpl("https://service-a.example.com"), "service-a" ); DataProcessor processorB = dbos.registerProxy( DataProcessor.class, new DataProcessorImpl("https://service-b.example.com"), "service-b" ); dbos.launch(); // Each proxy routes workflows to its own instance processorA.process("job-1"); processorB.process("job-2"); ``` The instance name is stored alongside the workflow record in the database. When DBOS recovers an interrupted workflow, it uses the stored name to find the correct instance and resume execution on it. :::warning All named instances must be registered before `dbos.launch()`. If DBOS tries to recover a workflow for an instance that hasn't been registered, recovery will fail. ::: ### Enqueueing to a Named Instance When enqueueing a workflow via `DBOSClient` from external code, pass the instance name to the `EnqueueOptions` constructor, after the class name, to target a specific instance: ```java var client = new DBOSClient(dbUrl, dbUser, dbPassword); var options = new EnqueueOptions( "process", "com.example.DataProcessorImpl", "service-a", QueueName.of("my-queue")); client.enqueueWorkflow(options, new Object[]{"job-3"}); ``` ### Using `@WorkflowClassName` for Stable Names If your class name might change across refactors, annotate the implementation class with [`@WorkflowClassName`](../reference/workflows-steps.md#workflowclassname) to give it a stable portable name: ```java @WorkflowClassName("data-processor") class DataProcessorImpl implements DataProcessor { ... } ``` Then reference that stable name in `WorkflowSchedule`, `EnqueueOptions`, or any other place that takes a class name string. ### Filtering by Instance Name Use `withInstanceName` on `ListWorkflowsInput` to list workflows that ran on a specific instance: ```java var workflows = dbos.listWorkflows( new ListWorkflowsInput().withInstanceName("service-a") ); ``` --- ## Communicating with Workflows(Tutorials) DBOS provides a few different ways to communicate with your workflows. You can: - [Send messages to workflows](#workflow-messaging-and-notifications) - [Publish events from workflows for clients to read](#workflow-events) ### Workflow Messaging and Notifications You can send messages to a specific workflow. This is useful for signaling a workflow or sending notifications to it while it's running. ##### send ```java void send(String destinationId, Object message, String topic, String idempotencyKey) ``` You can call `dbos.send()` to send a message to a workflow. Messages can optionally be associated with a topic and are queued on the receiver per topic. You can also call [`send`](../reference/client.md#send) from outside of your DBOS application with the [DBOS Client](../reference/client.md) or with the ['dbos.send_message' PL/pgSQL function](../../explanations/system-tables.md#dbossend_message) ##### recv ```java Optional recv(String topic, Duration timeout) ``` Workflows can call `dbos.recv()` to receive messages sent to them, optionally for a particular topic. Each call to `recv()` waits for and consumes the next message to arrive in the queue for the specified topic, returning `Optional.empty()` if the wait times out. If the topic is not specified, this method only receives messages sent without a topic. ##### sendBulk ```java void sendBulk(List messages) ``` You can call `dbos.sendBulk()` to send multiple messages to workflows in a single batch. Each `SendMessage` in the list specifies its own destination, message, topic, and optional idempotency key — messages need not share the same destination workflow. ```java dbos.sendBulk(List.of( new SendMessage(orderWorkflowId, "confirmed", ORDER_STATUS), new SendMessage(inventoryWorkflowId, order, RESERVE_TOPIC), new SendMessage(notificationWorkflowId, customerId, NOTIFY_TOPIC))); ``` You can also call [`sendBulk`](../reference/client.md#sendbulk) from outside your DBOS application with the [DBOS Client](../reference/client.md). ##### Messages Example Messages are especially useful for sending notifications to a workflow. For example, in a payments system, after redirecting customers to a payments page, the checkout workflow must wait for a notification that the user has paid. To wait for this notification, the payments workflow uses `recv()`, executing failure-handling code if the notification doesn't arrive in time: ```java interface Checkout { void checkoutWorkflow(); } class CheckoutImpl implements Checkout { private static final String PAYMENT_STATUS = "payment_status"; @Workflow public void checkoutWorkflow() { // Validate the order, redirect the customer to a payments page, // then wait for a notification. Optional paymentStatus = dbos.recv(PAYMENT_STATUS, Duration.ofSeconds(60)); if (paymentStatus.isPresent() && paymentStatus.get().equals("paid")) { // Handle a successful payment. } else { // Handle a failed payment or timeout. } } } ``` An endpoint waits for the payment processor to send the notification, then uses `send()` to forward it to the workflow: ```java app.post("/payment_webhook/{workflow_id}/{payment_status}", ctx -> { String workflowId = ctx.pathParam("workflow_id"); String paymentStatus = ctx.pathParam("payment_status"); // Send the payment status to the checkout workflow. dbos.send(workflowId, paymentStatus, PAYMENT_STATUS, null); ctx.result("Payment status sent"); }); ``` ##### Reliability Guarantees All messages are persisted to the database, so if `send` completes successfully, the destination workflow is guaranteed to be able to `recv` it. If you're sending a message from a workflow, DBOS guarantees exactly-once delivery. If you're sending a message from normal Java code, you can use a unique idempotency key to guarantee exactly-once delivery. ### Workflow Events Workflows can publish _events_, which are key-value pairs associated with the workflow. They are useful for publishing information about the status of a workflow or to send a result to clients while the workflow is running. ##### setEvent ```java void setEvent(String key, Object value, SerializationStrategy serialization) ``` Any workflow can call [`dbos.setEvent`](../reference/methods.md#setevent) to publish a key-value pair, or update its value if it has already been published. ##### getEvent ```java Optional getEvent(String workflowId, String key, Duration timeout) ``` You can call [`dbos.getEvent`](../reference/methods.md#getevent) to retrieve the value published by a particular workflow identity for a particular key. If the event does not yet exist, this call waits for it to be published, returning `Optional.empty()` if the wait times out. You can also call [`getEvent`](../reference/client.md#getevent) from outside of your DBOS application with [DBOS Client](../reference/client.md). ##### Events Example Events are especially useful for writing interactive workflows that communicate information to their caller. For example, in a checkout system, after validating an order, the checkout workflow needs to send the customer a unique payment ID. To communicate the payment ID to the customer, it uses events. The payments workflow emits the payment ID using `setEvent()`: ```java interface Checkout { void checkoutWorkflow(); } class CheckoutImpl implements Checkout { private static final String PAYMENT_ID = "payment_id"; @Workflow public void checkoutWorkflow() { // ... validation logic String paymentId = generatePaymentId(); dbos.setEvent(PAYMENT_ID, paymentId); // ... continue processing } } ``` The handler that originally started the workflow uses `getEvent()` to await this payment ID, then returns it: ```java app.post("/checkout/{idempotency_key}", ctx -> { String idempotencyKey = ctx.pathParam("idempotency_key"); // Idempotently start the checkout workflow in the background. WorkflowHandle handle = dbos.startWorkflow( () -> checkoutProxy.checkoutWorkflow(), new StartWorkflowOptions().withWorkflowId(idempotencyKey) ); // Wait for the checkout workflow to send a payment ID, then return it. Optional paymentId = dbos.getEvent(handle.workflowId(), PAYMENT_ID, Duration.ofSeconds(60)); if (paymentId.isEmpty()) { ctx.status(404); ctx.result("Checkout failed to start"); } else { ctx.result(paymentId.get()); } }); ``` ##### Reliability Guarantees All events are persisted to the database, so the latest version of an event is always retrievable. Additionally, if `getEvent` is called in a workflow, the retrieved value is persisted in the database so workflow recovery can use that value, even if the event is later updated. --- ## Workflow Management(3) You can view and manage your durable workflow executions via the [DBOS Console](../../conductor/workflow-management.md) or programmatically. ### Listing Workflows You can list your application's workflows programmatically via [`dbos.listWorkflows`](../reference/methods.md#listworkflows) or using the [`DBOSClient`](../reference/client.md#listworkflows). You can also view a searchable and expandable list of your application's workflows from its page on the [DBOS Console](../../conductor/workflow-management.md). ### Listing Workflow Steps You can list the steps of a workflow programmatically via [`dbos.listWorkflowSteps`](../reference/methods.md#listworkflowsteps) or using the [`DBOSClient`](../reference/client.md#listworkflowsteps). You can also visualize a workflow's execution as a trace timeline (showing the workflow, its steps, and its child workflows and their steps) from its page on the [DBOS Console](../../conductor/workflow-management.md). For example, here is the trace of a workflow that processes multiple tasks concurrently by enqueuing child workflows: ### Cancelling Workflows You can cancel the execution of a workflow from the web UI, programmatically via [`dbos.cancelWorkflow`](../reference/methods.md#cancelworkflow), or using the [`DBOSClient`](../reference/client.md#cancelworkflow). If the workflow is currently executing, cancelling it preempts its execution (interrupting it at the beginning of its next step). If the workflow is enqueued, cancelling removes it from the queue. To also cancel all descendant workflows spawned by the cancelled workflow, pass `cancelChildren = true`: ```java dbos.cancelWorkflow(workflowId, true); // or for multiple workflows: dbos.cancelWorkflows(workflowIds, true); ``` When a workflow is cancelled while it is waiting inside [`recv`](../reference/methods.md#recv) or [`getEvent`](../reference/methods.md#getevent), a `DBOSWorkflowCancelledException` is thrown to abort it immediately rather than waiting for the timeout to expire. ### Resuming Workflows You can resume a workflow from its last completed step from the web UI, programmatically via [`dbos.resumeWorkflow`](../reference/methods.md#resumeworkflow), or using the [`DBOSClient`](../reference/client.md#resumeworkflow). You can use this to resume workflows that are cancelled or that have exceeded their maximum recovery attempts. You can also use this to start an enqueued workflow immediately, bypassing its queue. ### Forking Workflows You can start a new execution of a workflow by **forking** it from a specific step. When you fork a workflow, DBOS generates a new workflow with a new workflow ID, copies to that workflow the original workflow's inputs and all its steps up to the selected step, then begins executing the new workflow from the selected step. Forking a workflow is useful for recovering from outages in downstream services (by forking from the step that failed after the outage is resolved) or for "patching" workflows that failed due to a bug in a previous application version (by forking from the bugged step to an application version on which the bug is fixed). You can fork a workflow programmatically using [`dbos.forkWorkflow`](../reference/methods.md#forkworkflow) or using the [`DBOSClient`](../reference/client.md#forkworkflow). You can also fork a workflow from a step from the web UI by clicking on that step in the workflow's graph visualization: ### Workflow Attributes You can attach custom metadata to any workflow as a `Map`. Attributes are stored as JSON and are searchable. **At creation:** ```java dbos.startWorkflow(() -> proxy.processOrder(orderId), new StartWorkflowOptions() .withAttributes(Map.of("customerId", "cust-123", "region", "us-west"))); ``` **After creation (from anywhere, including inside a workflow):** ```java dbos.updateWorkflowAttributes(workflowId, Map.of("status", "awaiting-payment", "invoiceId", "inv-456")); ``` When called from within a workflow, `updateWorkflowAttributes` is recorded as a step and executes exactly once even under recovery. **Filtering by attributes:** ```java // Find all workflows for customer cust-123 List results = dbos.listWorkflows( new ListWorkflowsInput() .withAttributes(Map.of("customerId", "cust-123"))); ``` The filter uses PostgreSQL's `@>` containment operator — it matches any workflow whose `attributes` map contains all the specified key-value pairs. Pass a subset of the attributes to match; extra attributes on the workflow are ignored. --- ## Workflows(Tutorials) Workflows provide **durable execution** so you can write programs that are **resilient to any failure**. Workflows are comprised of [steps](./step-tutorial.md), which wrap ordinary Java methods. If a workflow is interrupted for any reason (e.g., an executor restarts or crashes), when your program restarts the workflow automatically resumes execution from the last completed step. To write a workflow, annotate a method with [`@Workflow`](../reference/workflows-steps.md#workflow). All workflow methods must be registered before DBOS is launched. A workflow method can have any parameters and return type (including void), as long as they are serializable. Here's an example of a workflow: ```java interface Example { public String workflow(); } class ExampleImpl implements Example { private final DBOS dbos; public ExampleImpl(DBOS dbos) { this.dbos = dbos; } private void stepOne() { System.out.println("Step one completed!"); } private void stepTwo() { System.out.println("Step two completed!"); } @Override @Workflow public String workflow() { dbos.runStep(() -> stepOne(), "stepOne"); dbos.runStep(() -> stepTwo(), "stepTwo"); return "success"; } } public class App { public static void main(String[] args) throws Exception { // Configure and create a DBOS instance DBOSConfig config = ... DBOS dbos = new DBOS(config); // Register the workflow, creating a proxy object Example proxy = dbos.registerProxy(Example.class, new ExampleImpl(dbos)); // Launch DBOS after registering all workflows dbos.launch(); // Call the registered workflow through the proxy String result = proxy.workflow(); System.out.println("Workflow result: " + result); } } ``` ### Starting Workflows In The Background One common use-case for workflows is building reliable background tasks that keep running even when your program is interrupted, restarted, or crashes. You can use [`startWorkflow`](../reference/workflows-steps.md#startworkflow) to start a workflow in the background. When you start a workflow this way, it returns a [workflow handle](../reference/workflows-steps.md#workflowhandle), from which you can access information about the workflow or wait for it to complete and retrieve its result. Here's an example: ```java public void runWorkflowExample(DBOS dbos, Example proxy) throws Exception { // Start the background task WorkflowHandle handle = dbos.startWorkflow(() -> proxy.workflow()); // Wait for the background task to complete and retrieve its result String result = handle.getResult(); System.out.println("Workflow result: " + result); } ``` After starting a workflow in the background, you can use [`retrieveWorkflow`](../reference/methods.md#retrieveworkflow) to retrieve a workflow's handle from its ID. You can also retrieve a workflow's handle from outside of your DBOS application with [`DBOSClient.retrieveWorkflow`](../reference/client.md#retrieveworkflow). If you need to run many workflows in the background and manage their concurrency or flow control, use [queues](./queue-tutorial.md). ### Workflow IDs and Idempotency Every time you execute a workflow, that execution is assigned a unique ID, by default a [UUID](https://en.wikipedia.org/wiki/Universally_unique_identifier). You can access this ID from the [`DBOS.workflowId`](../reference/methods.md#workflowid) method. Workflow IDs are useful for communicating with workflows and developing interactive workflows. You can set the workflow ID of a workflow using `withWorkflowId` method of `WorkflowOptions` or `StartWorkflowOptions`. Workflow IDs are **globally unique** within your application. An assigned workflow ID acts as an idempotency key: if a workflow is called multiple times with the same ID, it executes only once. This is useful if your operations have side effects like making a payment or sending an email. For example: ```java public void directInvocationExample(DBOS dbos, Example proxy) throws Exception { String myID = "unique-workflow-id-123"; WorkflowOptions options = new WorkflowOptions().withWorkflowId(myID); try (var _ctx = new options.setContext()) { var result = proxy.workflow(); System.out.println("Result: " + result); } } public void startWorkflowExample(DBOS dbos, Example proxy) throws Exception { String myID = "unique-workflow-id-123"; WorkflowHandle handle = dbos.startWorkflow( () -> proxy.exampleWorkflow(), new StartWorkflowOptions().withWorkflowId(myID) ); String result = handle.getResult(); System.out.println("Result: " + result); } ``` ### Determinism Workflows are in most respects normal Java methods. They can have loops, branches, conditionals, and so on. However, a workflow method must be **deterministic**: if called multiple times with the same inputs, it should invoke the same steps with the same inputs in the same order (given the same return values from those steps). If you need to perform a non-deterministic operation like accessing the database, calling a third-party API, generating a random number, or getting the local time, you shouldn't do it directly in a workflow method. Instead, you should do all non-deterministic operations in [steps](./step-tutorial.md). :::warning Java's threading and concurrency APIs are non-deterministic. You should use them only inside steps. ::: For example, **don't do this**: ```java @Workflow public String workflow() { // Random number generation is not deterministic! // This workflow is not idempotent! int randomChoice = new Random().nextInt(2); if (randomChoice == 0) { return dbos.runStep(() -> stepOne(), "stepOne"); } else { return dbos.runStep(() -> stepTwo(), "stepTwo"); } } ``` Instead, do this: ```java private int generateChoice() { return new Random().nextInt(2); } @Workflow public String workflow() { // this workflow is idempotent because the random number generation // is inside a step so it only gets executed once per workflow ID int randomChoice = dbos.runStep(() -> generateChoice(), "generateChoice"); if (randomChoice == 0) { return dbos.runStep(() -> stepOne(), "stepOne"); } else { return dbos.runStep(() -> stepTwo(), "stepTwo"); } } ``` ### Workflow Timeouts You can set a timeout for a workflow using [`withTimeout`](../reference/workflows-steps.md#startworkflow) in `WorkflowOptions` and `StartWorkflowOptions`. When the timeout expires, the workflow and all its children (by default) are cancelled. Cancelling a workflow sets its status to CANCELLED and preempts its execution at the beginning of its next step. You can detach a child workflow from its parent's timeout by starting it with a custom timeout using `withTimeout`. Timeouts are **start-to-completion**: if a workflow is [enqueued](./queue-tutorial.md), the timeout does not begin until the workflow is dequeued and starts execution. Also, timeouts are durable: they are stored in the database and persist across restarts, so workflows can have very long timeouts. ```java // set timeout for direct invocation var options = new WorkflowOptions().withTimeout(Duration.ofHours(12)); try (var _ctx = options.setContext()) { proxy.workflow(); } // set timeout with start workflow var handle = dbos.startWorkflow( () -> proxy.workflow(), new StartWorkflowOptions().withTimeout(Duration.ofHours(12)) ); ``` ### Durable Sleep You can use [`sleep`](../reference/methods.md#sleep) to put your workflow to sleep for any period of time. This sleep is **durable**—DBOS saves the wakeup time in the database so that even if the workflow is interrupted and restarted multiple times while sleeping, it still wakes up on schedule. Sleeping is useful for scheduling work to run in the future (even days, weeks, or months from now). For example: ```java public String runTask(String task) { // Execute the task... return "task completed"; } @Workflow public String exampleWorkflow(Duration sleepTime, String task) { // Sleep for the specified duration dbos.sleep(sleepTime); // Execute the task after sleeping String result = dbos.runStep( () -> runTask(task), "runTask" ); return result; } ``` ### Debouncing Workflows You can debounce workflows to delay their execution until some time has passed since the workflow was last called. This is useful for preventing wasted work when a workflow may be triggered multiple times in quick succession. For example, if a user is editing an input field, you can debounce their changes to execute a processing workflow only after they haven't edited the field for some time: ```java @Workflow public String processInput(String userInput) { ... } var debouncer = dbos.debouncer() .withDebounceTimeout(Duration.ofMinutes(5)); // Each time a user submits a new input, debounce the processInput workflow. // The workflow will wait until 60 seconds after the user stops submitting new inputs, // then process the last input submitted. void onUserInputSubmit(String userId, String userInput) { debouncer.debounce( userId, Duration.ofSeconds(60), () -> svc.processInput(userInput)); } ``` See the [debouncing reference](../reference/methods.md#debouncing) for more details. ### Workflow Guarantees Workflows provide the following reliability guarantees. These guarantees assume that the application and database may crash and go offline at any point in time, but are always restarted and return online. 1. Workflows always run to completion. If a DBOS process is interrupted while executing a workflow and restarts, it resumes the workflow from the last completed step. 2. [Steps](./step-tutorial.md) are tried _at least once_ but are never re-executed after they complete. If a failure occurs inside a step, the step may be retried, but once a step has completed (returned a value or thrown an exception to the calling workflow), it will never be re-executed. If an exception is thrown from a workflow, the workflow **terminates**—DBOS records the exception, sets the workflow status to `ERROR`, and **does not recover the workflow**. This is because uncaught exceptions are assumed to be nonrecoverable. If your workflow performs operations that may transiently fail (for example, sending HTTP requests to unreliable services), those should be performed in [steps with configured retries](./step-tutorial.md#configurable-retries). DBOS provides [tooling](../reference/methods.md#workflow-management-methods) to help you identify failed workflows and examine the specific uncaught exceptions. --- ## Upgrading(Java) ### Upgrading a Running Application For every DBOS upgrade, migrate the system database before anything that needs the new schema runs. With `withMigrate(true)` (the default), `dbos.launch()` migrates the system database. If you run with `withMigrate(false)`, run [`dbosctl sysdb migrate`](../conductor/reference/dbosctl.md#dbosctl-sysdb-migrate) before deploying the upgrade. [`DBOSClient`](./reference/client.md) never migrates, so upgrade clients only after an upgraded application has launched or you have run `dbosctl sysdb migrate`. An application or client that needs a newer schema than the system database has throws `IllegalStateException` at launch or construction. ### Upgrading to v1.1 Most code written against DBOS Transact Java 1.0 compiles and runs unchanged on 1.1; the exceptions are listed below. 1.1 deprecates in-memory queues and a few other APIs, which will be removed in 2.0, and validates some inputs that 1.0 accepted. This section covers what might require a change to your code or configuration, and what to replace deprecated APIs with. For new features, see the [release notes](https://github.com/dbos-inc/dbos-transact-java/releases). #### 1.1 Is Required Before 1.2 If your application servers run 1.0, upgrade all of them to 1.1 before any server runs 1.2 or later, and don't run 1.0 servers against the same system database as servers on 1.2 or later. 1.2 changes how two things are stored in the system database, and 1.1 is the release that understands both the old and the new formats: - **Debouncing.** 1.2 debounces by delaying the workflow itself on its queue, instead of through a separate debouncer workflow. 1.1 still debounces through a debouncer workflow, but when a workflow debounced by 1.2 already holds the key, 1.1 extends that workflow's delay and replaces its arguments, as 1.2 does. 1.0 doesn't recognize such a workflow: a 1.0 server debouncing the same key can wait on a workflow that never answers it and then run the work a second time. - **Workflow inputs and outputs.** 1.2 writes them to new tables. 1.1 reads both the new tables and the old columns, but 1.0 reads only the old columns, so it can't recover or return the result of a workflow started by 1.2. #### Changes That May Require Action ##### Java CLI Removed The Java `dbos` CLI (the `transact-cli` module and its native binaries) has been removed. Use [`dbosctl`](../conductor/reference/dbosctl.md#system-database-commands) instead: | Before (1.0) | After (1.1) | |---|---| | `dbos migrate` | [`dbosctl sysdb migrate`](../conductor/reference/dbosctl.md#dbosctl-sysdb-migrate) | | `dbos reset` | [`dbosctl sysdb reset`](../conductor/reference/dbosctl.md#dbosctl-sysdb-reset) | ##### Application Names If you use [Conductor](../conductor/overview.md) or DBOS Cloud, `dbos.launch()` now throws `IllegalArgumentException` if your application name doesn't follow the [naming rule](./reference/lifecycle.md#dbosconfig): 3–256 lowercase letters, numbers, dashes, and underscores. Self-hosted applications only log a warning. To fix a name, change the name passed to `DBOSConfig.defaults(...)` or `withAppName`. In Spring Boot, set `dbos.application.name`; without it, DBOS uses `spring.application.name`, which often contains uppercase letters or dots. Rows created before 1.1 aren't owned by any application, so changing the name as part of this upgrade doesn't require transferring ownership. ##### Stricter Validation These inputs used to be accepted and now throw `IllegalArgumentException`: - A queue configuration, from `registerQueue` or `updateQueue`, whose `workerConcurrency` exceeds `concurrency`, or whose rate limit sets only one of `max` and `period`. See the [queues reference](./reference/queues.md#queueoptions) for all the rules. A queue already stored with an invalid configuration keeps running. - A negative priority, from `withPriority` on `StartWorkflowOptions` or `EnqueueOptions`, or on a debouncer. - A priority on a debouncer that has no queue. Iterating `readStream` for a workflow ID that doesn't exist now throws `DBOSNonExistentWorkflowException` instead of ending as an empty stream. ##### Workflows Enqueued Without a Version A workflow enqueued by [`DBOSClient`](./reference/client.md) without `withAppVersion` is now dequeued only by executors running the latest application version; in 1.0, any executor dequeued it. During a blue-green upgrade, such workflows therefore run on the new executors once they launch. If executors on an older version must run a workflow, set `withAppVersion` when enqueuing it. ##### Event Waits During a Rolling Upgrade from 1.0 1.1 sends event notifications from the application instead of from database triggers, and its migration removes the triggers. While 1.0 executors are still running, events they set don't wake `getEvent` calls in 1.1 processes, which then see the event only at their next re-check, up to a minute later. If your 1.1 servers wait on events from workflows still running on 1.0 executors, for example to answer a request, set `withUseListenNotify(false)` on them until the last 1.0 executor has stopped; `getEvent` then checks every second. ##### Several Applications Sharing a System Database Every workflow, queue, schedule, and application version is now owned by the application that created it, and listing operations return only the calling application's objects, plus those no application owns (such as everything created before 1.1). If several applications share one system database, give each [`DBOSClient`](./reference/client.md#named-and-unnamed-clients) the `applicationName` it acts for; a client without one sees every application's rows. Rows created before 1.1 are owned by no application, and every application treats them as its own: any application may dequeue an unowned workflow, and every application polls unowned queues and fires unowned schedules. Before adding a second application to a system database, transfer the unowned rows to the application that created them: ```shell dbosctl sysdb rename-application --to my-app --adopt-unclaimed-rows ``` [`DBOSClient.renameApplication`](./reference/client.md#renameapplication) does the same from code. See [Unowned Rows](../explanations/sharing-a-system-database.md#unowned-rows). ##### Constructors `DBOSConfig` and `ListWorkflowsInput` gained fields, so calls to their all-arguments constructors no longer compile. Neither of these types is designed to be constructed via its all-arguments constructor. Build a base `DBOSConfig` with `DBOSConfig.defaults(...)` or `DBOSConfig.defaultsFromEnv(...)` and customize with its `with` methods. Build a base `ListWorkflowsInput` with its default constructor and customize with its `with` methods. `WorkflowStatus`, `StepInfo`, and `VersionInfo` also gained fields. DBOS returns these types rather than applications building them, so this should only affect test code that constructs them, for example to stub `listWorkflows` or `listWorkflowSteps` in a mock. The `DBOSSystemDatabaseException` constructor takes a `SQLException` instead of a `Throwable`. Code that constructs it must be updated and recompiled. #### Database-Backed Queues In-memory queues, declared with `new Queue(...)` and registered with `dbos.registerQueue(Queue)` before launch, are deprecated. They exist only in the process that declares them: Conductor can't list them, and no other executor can see or poll them. Instead, register queues in the system database with `dbos.registerQueue(String, QueueOptions)` after `dbos.launch()`, and enqueue on them by name. **Before:** ```java Queue queue = new Queue("example-queue").withWorkerConcurrency(5); dbos.registerQueue(queue); dbos.launch(); dbos.startWorkflow(() -> proxy.processTask(task), new StartWorkflowOptions().withQueue(queue)); ``` **After:** ```java dbos.launch(); dbos.registerQueue("example-queue", QueueOptions.setWorkerConcurrency(5)); dbos.startWorkflow(() -> proxy.processTask(task), new StartWorkflowOptions(QueueName.of("example-queue"))); ``` When migrating your queues, note that: - Register every queue your application previously declared in memory. Starting a workflow on a queue that isn't registered throws. - Pass the queue name as a `QueueName`. `new StartWorkflowOptions(String)` takes a *workflow ID*, so `new StartWorkflowOptions("example-queue")` starts an unqueued workflow. - Replace `dbos.getQueue(name)` with `dbos.findQueue(name)`. `getQueue` reads only in-memory queues. - By default, `registerQueue` overwrites an existing queue's configuration only if this executor runs the latest application version. Pass a [`QueueConflictResolution`](./reference/queues.md#queueconflictresolution) to change that. #### Moving Off Legacy Partitioned Queues The `partitionQueue` option (`QueueOptions.setPartitionQueue`, and `Queue.withPartitioningEnabled` for in-memory queues) is deprecated. Under it, the queue-wide `concurrency`, `workerConcurrency`, and rate limit applied to each partition. Replace them with the per-partition limits, `partitionConcurrency`, `partitionWorkerConcurrency`, and `partitionRateLimit`; setting any of them partitions the queue. **Before:** ```java dbos.registerQueue("partitioned-queue", QueueOptions.setPartitionQueue(true).andConcurrency(1)); ``` **After:** ```java dbos.registerQueue("partitioned-queue", QueueOptions.setPartitionConcurrency(1)); ``` `dbos.updateQueue` can't change the limits of a queue registered with `partitionQueue`. To move a queue across, re-register it with `dbos.registerQueue`, which replaces its whole configuration. :::warning Partitioning a queue that wasn't partitioned before strands the workflows already on it. Workflows enqueued without a partition key are never dequeued from a partitioned queue. Drain the queue before giving it its first per-partition limit, or re-enqueue its workflows with a partition key. Workflows already stranded this way can be moved to a queue that is not partitioned with [`dbos.resumeWorkflow(workflowId, queueName)`](./reference/methods.md#resumeworkflow). ::: #### Deprecations The following APIs are deprecated in 1.1 and will be removed in 2.0. | Deprecated | Replacement | |---|---| | `new Queue(...)` constructors and `Queue.withName`, `withConcurrency`, `withWorkerConcurrency`, `withPriorityEnabled`, `withPartitioningEnabled`, `withRateLimit` (all overloads), `withPollingInterval` | `dbos.registerQueue(String, QueueOptions)` after launch | | `Queue.partitioningEnabled()` | `Queue.isPartitioned()`, or `Queue.isLegacyPartitioned()` to detect a queue partitioned with the deprecated flag | | `dbos.registerQueue(Queue)`, `dbos.registerQueues(Queue...)` | `dbos.registerQueue(String, QueueOptions)` after launch | | `dbos.getQueue(String)` | `dbos.findQueue(String)` | | `Queue`-typed overloads: `new StartWorkflowOptions(Queue)`, `StartWorkflowOptions.withQueue(Queue)`, `ForkOptions.withQueue(Queue)`, `Debouncer.withQueue(Queue)`, `DebouncerClient.withQueue(Queue)`, `DBOSConfig.withListenQueue(Queue)`, `DBOSConfig.withListenQueues(Queue...)` | The `QueueName` or `String` overloads | | `QueueOptions.setPriorityEnabled`, `withPriorityEnabled`, `andPriorityEnabled`, the `QueueOptions.priorityEnabled()` accessor, and `Queue.priorityEnabled()` | None. Every queue dequeues in priority order; set a priority on the workflow instead. | | `QueueOptions.setPartitionQueue`, `withPartitionQueue`, `andPartitionQueue`, and the `partitionQueue()` accessor | `setPartitionConcurrency`, `setPartitionWorkerConcurrency`, `setPartitionRateLimit` (and their `and`/`with` forms) | | The seven-argument `QueueOptions` constructor (without per-partition limits) | The static `QueueOptions.set...` factories | | `DBOSClient.EnqueueOptions`, and the `DBOSClient.enqueueWorkflow` / `enqueuePortableWorkflow` overloads that take it | The top-level [`dev.dbos.transact.EnqueueOptions`](./reference/client.md#enqueueoptions), whose constructors take the workflow name, optional class and instance names, and a `QueueName`. For portable enqueue, set `withSerialization(SerializationStrategy.PORTABLE)` and call `enqueueWorkflow(options, positionalArgs, namedArgs)`. | | `Debouncer.withDeduplicationId`, `DebouncerClient.withDeduplicationId` | None. From the next release the debouncer sets the deduplication ID itself and ignores this setting. | | `ExternalState`, `DBOSIntegration.getExternalState`, `DBOSIntegration.upsertExternalState` (the `event_dispatch_kv` API) | Store integration state in your own table. A shared system database migration will drop the `event_dispatch_kv` table sometime after Java 2.0. | | `DBOSSystemDatabaseException.databaseException()` | `getCause()` | --- ### Upgrading to v1.0 #### Breaking Changes ##### ExternalState.updateTime: BigDecimal → Instant The `updateTime` field on the `ExternalState` record changed from `BigDecimal` to `java.time.Instant`, and the `withUpdateTime` builder method updated accordingly. This only affects custom plugin authors who directly construct or pattern-match on `ExternalState`. ```java // Before ExternalState state = new ExternalState(service, wfName, key, value, new BigDecimal(System.currentTimeMillis()), BigInteger.ZERO); // After ExternalState state = new ExternalState(service, wfName, key, value, Instant.now(), BigInteger.ZERO); ``` ##### ForkOptions.timeout: Timeout → @Nullable Duration The `timeout` field on `ForkOptions` changed from the `Timeout` sealed interface to `@Nullable Duration`. The `withTimeout(Timeout)` and `withNoTimeout()` methods were removed. ```java // Before new ForkOptions().withTimeout(Timeout.none()); new ForkOptions().withNoTimeout(); // After new ForkOptions().withTimeout((Duration) null); // no timeout new ForkOptions().withTimeout(Duration.ofMinutes(5)); // explicit timeout ``` `Timeout.of(...)` usage can be replaced directly with the equivalent `Duration`: ```java // Before new ForkOptions().withTimeout(Timeout.of(Duration.ofMinutes(5))); // After new ForkOptions().withTimeout(Duration.ofMinutes(5)); ``` Note: `Timeout` itself is still used by `StartWorkflowOptions` and `WorkflowOptions` — only `ForkOptions` changed. ##### CLI: postgres and workflow subcommands removed The `dbos postgres` and `dbos workflow` subcommand groups have been removed from the CLI. The CLI now only supports `dbos migrate` and `dbos reset`. Use the [`DBOSClient`](./reference/client.md) API or the [DBOS Console](../conductor/workflow-management.md) to manage workflows programmatically. Additionally, the CLI now ships as a pre-compiled native binary (via GraalVM AOT compilation) for Linux, macOS, and Windows. Download the appropriate binary from the GitHub Releases page — no JVM required. :::note The Java CLI was removed in v1.1 in favor of [`dbosctl`](../conductor/reference/dbosctl.md). See [Java CLI Removed](#java-cli-removed). ::: ##### Jackson upgraded to 3.1.x The Jackson library dependency was upgraded from 2.x to 3.1.x. Jackson 3.x has breaking API changes relative to 2.x (notably, `JsonRuntimeException` was removed). If your application has an explicit dependency on Jackson 2.x, you will need to upgrade it to 3.x. Jackson 3.x follows the same JSON format as 2.x, so no data migration is required. ##### Cancellation while waiting in recv() / getEvent() `recv()` and `getEvent()` now throw `DBOSWorkflowCancelledException` when the calling workflow is cancelled while they are waiting for a message or event. Previously the behavior was inconsistent; this aligns Java with the TypeScript and Python behavior. If your code catches `Exception` or `RuntimeException` around `recv`/`getEvent`, no change is needed. If you were relying on the previous behavior where cancellation during a wait would leave the workflow in a non-cancelled state, update accordingly. --- ### Upgrading to v0.9 #### Breaking Changes ##### Queue.withRateLimit(int, double) The `withRateLimit(int limit, double period)` overload on the `Queue` record has been removed. Replace it with the new `withRateLimit(int limit, long period, TimeUnit unit)` overload: ```java // Before new Queue("example-queue").withRateLimit(100, 60.0); // After new Queue("example-queue").withRateLimit(100, 60, TimeUnit.SECONDS); ``` ##### DBOS.registerWorkflow The static `DBOS.registerWorkflow` method has been removed. Call `DBOSIntegration.registerWorkflow` instead, which is accessible via `dbos.integration()`: ```java // Before DBOS.registerWorkflow(name, target, method); // After dbos.integration().registerWorkflow(name, target, method); ``` ##### DBOSIntegration.registerWorkflow `DBOSIntegration.registerWorkflow` now returns a `RegisteredWorkflow` instead of `void`. Code that discards the return value continues to compile without changes. Code that stores the result must update its declared type from `void` to `RegisteredWorkflow`. #### New Features ##### Step Factories Two new mechanisms let you commit a step's database write and its DBOS checkpoint atomically, so a crash between the two can never leave them out of sync. **Step factories** (non-Spring): construct a factory for your database library and pass it to `DBOS.runStep`: ```java // Plain JDBC var factory = new JdbcStepFactory(dataSource); dbos.runStep(factory, conn -> { conn.prepareStatement("INSERT INTO orders ...").executeUpdate(); return orderId; }); ``` `JdbcStepFactory`, `JdbiStepFactory`, and `JooqStepFactory` are available for JDBC, JDBI 3, and jOOQ respectively. **`@TransactionalStep`** (Spring Boot): annotate any Spring-managed method with `@TransactionalStep` and the `transact-spring-txstep-starter` module handles the rest — no factory wiring required: ```java @TransactionalStep public Order createOrder(OrderRequest request) { // Spring transaction + DBOS checkpoint committed atomically return orderRepository.save(new Order(request)); } ``` See the [Step Factories tutorial](./tutorials/step-factory-tutorial.md) for setup instructions and per-library examples. ##### sendBulk `DBOS.sendBulk` and `DBOSClient.sendBulk` let you send multiple workflow messages in a single batch: ```java dbos.sendBulk(List.of( new SendMessage(workflowIdA, "hello", "topic"), new SendMessage(workflowIdB, "world", "topic"))); ``` Each message in the batch is delivered independently; messages need not share the same destination. `DBOSClient.sendBulk` accepts an optional `SendOptions` for serialization and fork-delivery control. ##### Debouncer A new `Debouncer` class lets you coalesce repeated workflow invocations on the same key into a single execution that uses the most recently supplied arguments. The workflow fires after `debouncePeriod` of inactivity, or after an absolute `debounceTimeout` cap regardless of ongoing calls: ```java var debouncer = dbos.debouncer() .withDebounceTimeout(Duration.ofMinutes(5)); WorkflowHandle handle = debouncer.debounce( userId, Duration.ofSeconds(60), () -> svc.processInput(userInput)); ``` `DebouncerClient` provides the same capability from external code without a running DBOS executor. See the [Debouncing reference](./reference/methods.md#debouncing) for the full API. ##### Dynamic queue management Queues can now be registered, updated, and deleted at runtime without restarting your application. Queue configuration is persisted to the system database and survives restarts. Register a queue after `dbos.launch()` using `dbos.registerQueue`: ```java dbos.registerQueue("pipeline-queue", QueueOptions.setConcurrency(10).andRateLimit(100, Duration.ofSeconds(60))); ``` Update configuration at runtime with `dbos.updateQueue` — only the fields you specify are changed: ```java dbos.updateQueue("pipeline-queue", QueueOptions.setConcurrency(20)); ``` Additional methods — `dbos.findQueue`, `dbos.listQueues`, and `dbos.deleteQueue` — let you inspect and manage queues at runtime. [`DBOSClient`](./reference/client.md#queue-management-methods) exposes the same operations for managing queues from outside your application. See [Queues & Concurrency](./tutorials/queue-tutorial.md) for a full guide and [Queues reference](./reference/queues.md) for the complete API. ##### CockroachDB support CockroachDB is now a supported system database backend alongside PostgreSQL. No configuration changes are required; DBOS auto-detects CockroachDB and adjusts its behaviour accordingly (for example, `useListenNotify` is automatically set to `false`). #### Deprecations ##### Admin server The built-in admin server is deprecated. It still ships, and will be removed in a future release. Use [DBOS Conductor](https://docs.dbos.dev/conductor) instead. The related configuration APIs — `withAdminServer()`, `disableAdminServer()`, `enableAdminServer()`, and `withAdminServerPort()` on `DBOSConfig`, and the `dbos.admin-server.*` properties in Spring Boot — are deprecated alongside it. --- ### Upgrading to v0.8 DBOS Transact Java v0.8 contains several breaking changes. These changes were made to improve the developer experience as well as how DBOS Transact Java integrates into the larger Java ecosystem. This document explains how to update your existing DBOS Java app to v0.8. :::info Although we cannot guarantee that v0.8 will be the final release with breaking changes, our intention is to minimize or eliminate further breaking changes in DBOS Transact Java after v0.8. ::: #### DBOS Instance API `DBOS` is now an instance class instead of a static utility class. While static methods were easier to access, they are harder to test and mock. Furthermore, a `DBOS` instance API fits better into Dependency Injection based Java frameworks like [Spring](https://spring.io/). Prior to v0.8, you would configure DBOS via the static `configure` method: ```java DBOSConfig dbosConfig = DBOSConfig.defaultsFromEnv("my-app") .withAppVersion("0.1.0"); DBOS.configure(dbosConfig); ``` Now, you pass the DBOSConfig instance to directly to the DBOS constructor: ```java DBOSConfig dbosConfig = DBOSConfig.defaultsFromEnv("my-app") .withAppVersion("0.1.0"); DBOS dbos = new DBOS(dbosConfig); ``` `DBOS` implements the [AutoClosable](https://docs.oracle.com/javase/8/docs/api/java/lang/AutoCloseable.html) interface. This allows `DBOS` to work with the [try-with-resources](https://docs.oracle.com/javase/tutorial/essential/exceptions/tryResourceClose.html) statement or with JUnit's [@AutoClose annotation](https://docs.junit.org/6.0.3/api/org.junit.jupiter.api/org/junit/jupiter/api/AutoClose.html). ```java var dbosConfig = DBOSConfig.defaultsFromEnv("my-app") .withAppVersion("0.1.0"); try (var dbos = new DBOS(dbosConfig)) { Example proxy = dbos.registerProxy(Example.class, new ExampleImpl(dbos)); dbos.launch(); proxy.workflow(); } ``` With this change, registering DBOS proxies becomes slightly more involved. Previously, you could register a proxy object anytime prior to calling `DBOS.launch()`. Now, you must construct the `DBOS` instance before calling `registerProxy`. Like prior releases, all proxies must be registered before calling `DBOS.launch()`. :::info `DBOS.registerWorkflows()` is now named [`DBOS.registerProxy`](./reference/workflows-steps.md#registerproxy). ::: For [plugin](./reference/plugins.md) developers, the `DBOS` instance is now provided as a parameter on `dbosLaunched`. For more information, see the [Lifecycle Listeners documentation](./reference/plugins.md#lifecycle-listeners) #### DBOS API Changes Beyond the overarching change from static to instance methods, there were assorted other minor breaking changes to the DBOS API: * `DBOS.registerWorkflows` was renamed to `DBOS.registerProxy` * `DBOS.registerQueue` previously returned the `Queue` object, now it is void return * Several methods with `@Nullable` return types have been changed to return `Optional` * `DBOS.getWorkflowStatus()` * `DBOS.recv()` * `DBOS.getEvent()` * Several methods used by [plugins](./reference/plugins.md) were moved from `DBOS` to `DBOSIntegration`. The `DBOSIntegration` instance can be accessed via `DBOS.integration()`. * `DBOSIntegration.registerLifecycleListener` * `DBOSIntegration.getRegisteredWorkflows` * `DBOSIntegration.getRegisteredWorkflowsInstances` * `DBOSIntegration.startRegisteredWorkflow` * `DBOSIntegration.getExternalState` * `DBOSIntegration.upsertExternalState` :::info `DBOSIntegration.startRegisteredWorkflow` was previously named `DBOS.startWorkflow`. There are several `DBOS.startWorkflow` overloads, only the one with a `RegisteredWorkflow` parameter was renamed and moved. ::: #### Strongly Typed Fields In a variety of places across the public API surface area, fields have been changed to more semantically relevant types. For example, previously we represented both timeout and deadline as `Long` with semantic information encoded in the field names - i.e. `timeoutMs` and `deadlineEpochMs`. Now, we use the more semantically aligned types of [`Duration`](https://docs.oracle.com/javase/8/docs/api/java/time/Duration.html) for timeout and [`Instant`](https://docs.oracle.com/javase/8/docs/api/java/time/Instant.html) for deadline. With the change in type, we also simplified the field names to drop the semantic type information. WorkflowStatus * `status` field was previously a `String`, now it's a `WorkflowState` enum value * `name` field was renamed `workflowName` * `createdAt` and `updatedAt` changed their type from `Long` to `Instant` * `timeoutMs` was renamed `timeout` and the type changed from `Long` to `Duration` * `deadlineEpochMs` was renamed `deadline` and the type changed from `Long` to `Instant` * `startedAtEpochMs` was renamed `startedAt` and the type changed from `Long` to `Instant` StepInfo * `startedAtEpochMs` was renamed `startedAt` and the type was changed from `Long` to `Instant` * `completedAtEpochMs` was renamed `completedAt` and the type was changed from `Long` to `Instant` StepOptions * `intervalSeconds` was renamed `retryInterval` and the type was changed from `Double` to `Duration` ListWorkflowsInput * `withStartTime` and `withEndTime` changed parameter type from `OffsetDateTime` to `Instant` * `withStatuses` was renamed `withStatus` and the parameter type changed from `List` to `List` * `withWorkflowId` was renamed `withWorkflowIds` #### @Step / StepOptions Changes In previous versions, both `@Step` and `StepOptions` had a boolean `retriesAllowed` field. This field has been removed. Additionally, the default `maxAttempts` field of both `@Step` and `StepOptions` has been changed to 1. To enable retries, simply set `maxAttempts` to a value greater than `. :::info Note, as covered above, the `StepOptions.intervalSeconds` field was renamed to `retryInterval` and the type was changed to `Duration`. Annotations cannot use reference types like `Duration` so `@Step` still has an `intervalSeconds` field of type `double`. ::: #### @Scheduled Removed The `@Scheduled` annotation has been removed. For durable scheduled code in your app, you can use the new [Schedule Management Methods](./reference/methods.md#schedule-management-methods). #### DBOSClient Changes `DBOSClient.EnqueueOptions` changed the order of the parameters for the three string constructor. Previously, the className parameter was first and the workflowName was second. For consistency with other DBOS APIs, these two parameters have swapped position. Across the code base, when specifying the workflow name, class name, instance name of a registered workflow, we have the parameters in that order. Additionally, similar to DBOS changes detailed above, `DBOSClient.getWorkflowStatus` and `DBOS.getEvent` now return `Optional` instead of a `@Nullable` value; --- ## Production Checklist This page describes best practices you should follow when operating a DBOS application in production. ### Managing Postgres DBOS is entirely built on Postgres. Here are some recommendations for configuring a Postgres database to best work with DBOS. **Use any Postgres** - DBOS is compatible with any Postgres database, including standard self-hosted Postgres, RDS, Aurora, Google Cloud SQL, Azure PostgreSQL, Supabase, Neon, Planetscale, TimescaleDB, AlloyDB, PgDog, etc. **If using a connection pooler, use it in session mode** - Connect your DBOS applications to your Postgres database either directly or using a connection pooler in session mode. Do not use a connection pooler in transaction mode as some Postgres features that DBOS uses (e.g., LISTEN/NOTIFY) are not compatible with it. [This page](https://www.pgbouncer.org/features.html) documents the differences. **Configure a retention policy** - You should limit how much workflow history DBOS keeps in your system database. If you use Conductor, you can configure a [retention policy](../conductor/retention.md) for your application from the DBOS Console. **Manage the DBOS schema** - DBOS creates tables for its internal state in its [system database](../explanations/system-tables.md). By default, a DBOS application automatically creates these on startup. However, in production environments, a DBOS application may not run with sufficient privilege to create databases or tables. In that case, the [`dbosctl sysdb migrate`](../conductor/reference/dbosctl.md#dbosctl-sysdb-migrate) command can be run with a privileged user to create all DBOS system tables or migrate them to the latest version. Then, a DBOS application can run with lower privilege (requiring only access to the DBOS tables in the system database). If your database is managed by a DBA, `dbosctl sysdb migrate` can also print the SQL for them to apply instead of running it itself. ### Scalability You can easily scale a DBOS application by adding more servers to it, so the scalability of DBOS is fundamentally determined by the database it is connected to. We recommend taking these steps in Postgres to guarantee the scalability of your application. **Database Sizing** - For a typical workload, processing 1000 actions (steps or workflows) per second (2 billion actions per month) with DBOS requires 4 Postgres vCPUs. This is a conservative estimate that leaves headroom for unexpected bursts or spikes. When sizing your database for DBOS, we recommend using that number (scaled to your actual workload size) as a starting point then measuring usage in practice. **Monitor Connection Usage** - You can check the maximum number of connections your Postgres database can accept by running `SHOW max_connections;`. Typically, a Postgres database can support 100 connections per gigabyte of memory. You should make sure you have enough connections to support all your DBOS application servers. You can configure the maximum number of connections a DBOS application can make through the `sys_db_pool_size`/`systemDatabasePoolSize` configuration parameter. We do not recommend setting this to less than 5. **Monitor Database Usage** - A DBOS workflow requires two-three database writes (one at the beginning to checkpoint its input, one at the end to checkpoint its outcome, and optionally one to dequeue it if it was enqueued) plus one additional write per step (to checkpoint the step's outcome). In [benchmarks](https://www.dbos.dev/blog/benchmarking-workflow-execution-scalability-on-postgres), a DBOS application using a single Postgres database can sustain a throughput of >40K workflows or steps per second. In practice, if your expected load exceeds 1K workflows or steps per second, you should perform load tests to verify your Postgres database can handle the load or if you need a larger instance. If it approaches or exceeds 40K workflows or steps per second, we recommend sharding workflows across multiple Postgres servers. ### Availability To maximize availability of your DBOS application, we recommend using a highly available Postgres database. Most cloud Postgres providers provide multi-AZ replication with automatic failover, so your database can seamlessly fail over to a backup if anything goes wrong. Note that there is nothing DBOS-specific about this—we recommend following industry best practices for maximizing Postgres availability. If your Postgres database does become unavailable, all DBOS applications connected to it will pause workflow execution until they reconnect. When your database becomes available again, they will seamlessly resume. If you use [DBOS Conductor](../conductor/overview.md), note that it is out-of-band and off your application's workflow execution path, so its availability does not affect the availability of your applications. If your connection to Conductor is interrupted, your applications will continue operating normally. All Conductor features (recovery, observability, workflow management) will automatically resume once connectivity is restored. ### Upgrading the DBOS Library We recommend regularly upgrading the DBOS library to its latest version to take advantage of new features. All implementations of the DBOS library follow strict semantic versioning. Minor version upgrades do not introduce breaking changes. Major version upgrades may introduce breaking changes, but these are always documented in the release notes. New library versions are always announced on GitHub ([Python](https://github.com/dbos-inc/dbos-transact-py/releases), [TypeScript](https://github.com/dbos-inc/dbos-transact-ts/releases), [Go](https://github.com/dbos-inc/dbos-transact-golang/releases), [Java](https://github.com/dbos-inc/dbos-transact-java/releases)) and on the [community Discord](https://discord.com/invite/jsmC6pXGgX). --- ## Deploying With Google Cloud Run ## Deploying a DBOS App on Google Cloud Run This guide covers deploying a DBOS application to [Google Cloud Run](https://cloud.google.com/run) with a [Cloud SQL for PostgreSQL](https://cloud.google.com/sql/docs/postgres) database. It includes best practices for security, availability, and scalability. This guide assumes [DBOS Conductor](../conductor/overview.md) is hosted separately. ### Choosing a Cloud Run Execution Mode Cloud Run offers three execution modes, each mapping differently to DBOS workloads: - **Service** handles HTTP requests and auto-scales based on traffic and CPU usage. Best for synchronous workflows. - **Worker Pool** runs always-on instances with no HTTP listener. Best for queue-heavy applications that need all DBOS background services online at all times. - **Job** runs a container to completion and exits. Useful for periodic batch work with no always-on requirement. #### Service A [Cloud Run service](https://cloud.google.com/run/docs/overview/what-is-cloud-run#services) listens for HTTP requests and scales automatically based on traffic and CPU usage. DBOS runs background services—the scheduler, queue runner, recovery service, and Conductor connection—that operate independently of HTTP requests. These require CPU at all times, so you should use [instance-based billing](https://docs.cloud.google.com/run/docs/configuring/billing-settings) and set `--min-instances=1` to keep one instance always on. This is similar to the requirements for using [sidecars on Cloud Run](https://docs.cloud.google.com/run/docs/deploying#sidecars). :::caution Database connection exhaustion In Service mode, use a connection pooler like [PgBouncer](https://www.pgbouncer.org/) in front of your Cloud SQL instance. Cloud Run can scale to hundreds of instances under load, which may exhaust your database's maximum connections. PgBouncer must run in **session mode**—DBOS uses LISTEN/NOTIFY, which is [incompatible with transaction mode](https://www.pgbouncer.org/features.html). ::: #### Worker Pool A [Cloud Run worker pool](https://cloud.google.com/run/docs/overview/what-is-cloud-run#workers) runs always-on containers without an HTTP listener. Because instances never scale to zero, all DBOS background services stay online. Worker pools suit DBOS applications that rely heavily on queues. Every instance actively dequeues and processes workflows, and the pool can be resized via the [Cloud Run REST API](https://docs.cloud.google.com/run/docs/reference/rest). Worker pools don't auto-scale, but you can implement an **external scaler** from within the pool. Use a DBOS [scheduled workflow](../golang/tutorials/scheduled-workflows.md) that periodically checks queue length with [`ListWorkflows`](../golang/reference/methods.md#listworkflows) and calls the [Cloud Run Admin API](https://cloud.google.com/run/docs/reference/rest/v2/projects.locations.workerPools) to resize the pool based on load. This works _even from within the pool_ because DBOS guarantees only one process runs a scheduled function at a time, even across multiple instances. This prevents a thundering herd of conflicting resize requests. See [Scaling a worker pool](#scaling-a-worker-pool) below for a full walkthrough. #### Job A [Cloud Run job](https://cloud.google.com/run/docs/overview/what-is-cloud-run#jobs) runs a container to completion and exits without listening for HTTP requests. Because DBOS has a [built-in scheduler](../golang/tutorials/scheduled-workflows.md), you typically don't need Cloud Run Jobs. However, Jobs suit applications that consist entirely of periodic work with no always-on requirement—the job starts, runs workflows to completion, and shuts down, so you only pay for the time it runs. ### Deploying to Cloud Run Deploying a DBOS application to Cloud Run is no different from deploying any other containerized application. You need a Dockerfile, a database, and the standard Cloud Run deployment commands. The one DBOS-specific detail is the **database connection string**: it must be provided in `key=value` format (e.g., `user=postgres password=secret database=myappdb host=/cloudsql/...`). On Cloud Run, use the `--add-cloudsql-instances` flag to mount the [Cloud SQL Auth Proxy](https://cloud.google.com/sql/docs/postgres/connect-run) Unix socket, then pass the socket path as the `host` parameter. This gives your app a private, encrypted path to the database with no public IP. :::tip Schema migration By default, DBOS creates its [system tables](../explanations/system-tables.md) on startup. If your Cloud Run service account doesn't have DDL privileges, run [`dbosctl sysdb migrate`](../conductor/reference/dbosctl.md#dbosctl-sysdb-migrate) with a privileged user before deploying. :::
Walkthrough: deploying a DBOS app This walkthrough deploys a sample DBOS Go application ([source code](https://github.com/dbos-inc/dbos-demo-apps/tree/main/golang/cloudrun)) to Cloud Run with a Cloud SQL PostgreSQL database. It covers project setup, infrastructure, and deployment in both **Service** and **Worker Pool** modes. --- **Google Cloud project setup** You need a Google Cloud project with billing enabled and the required APIs turned on. Install the [Google Cloud SDK](https://cloud.google.com/sdk/docs/install-sdk), then: ```bash gcloud auth login gcloud projects create [YOUR_PROJECT_ID] gcloud config set project [YOUR_PROJECT_ID] gcloud beta billing projects link [YOUR_PROJECT_ID] \ --billing-account=[YOUR_BILLING_ACCOUNT_ID] gcloud services enable \ run.googleapis.com \ sqladmin.googleapis.com \ compute.googleapis.com \ servicenetworking.googleapis.com \ secretmanager.googleapis.com \ artifactregistry.googleapis.com \ cloudbuild.googleapis.com ``` --- **VPC networking for Cloud SQL** Create a VPC with a subnet for Cloud Run, allocate an IP range for VPC peering, and establish the peering connection: ```bash # Create VPC gcloud compute networks create main-vpc --subnet-mode=custom # Create subnet for Cloud Run gcloud compute networks subnets create run-subnet \ --network=main-vpc \ --region=us-central1 \ --range=10.0.0.0/24 # Allocate IP range for Cloud SQL peering gcloud compute addresses create google-managed-services-default \ --global \ --purpose=VPC_PEERING \ --prefix-length=16 \ --description="Peering for Google Cloud SQL" \ --network=main-vpc # Establish VPC peering gcloud services vpc-peerings connect \ --service=servicenetworking.googleapis.com \ --ranges=google-managed-services-default \ --network=main-vpc ``` --- **Cloud SQL PostgreSQL instance** Create a private-IP-only PostgreSQL instance and an application database: ```bash # Create the Cloud SQL instance (private IP only) gcloud sql instances create my-postgres-instance \ --database-version=POSTGRES_17 \ --tier=db-perf-optimized-N-2 \ --region=us-central1 \ --root-password="[YOUR_STRONG_PASSWORD]" \ --network=projects/[YOUR_PROJECT_ID]/global/networks/main-vpc \ --no-assign-ip # Create the application database gcloud sql databases create myappdb --instance=my-postgres-instance ``` Store the database password in Secret Manager: ```bash echo -n "[YOUR_STRONG_PASSWORD]" | gcloud secrets create db-password \ --data-file=- \ --replication-policy="automatic" ``` Store the [DBOS Conductor](../conductor/overview.md) API key: ```bash echo -n "[YOUR_CONDUCTOR_API_KEY]" | gcloud secrets create conductor-api-key \ --data-file=- \ --replication-policy="automatic" ``` :::note For production, consider creating a dedicated database user instead of using the `postgres` superuser. Grant it only the permissions your application needs. ::: --- **IAM service account and permissions** Create a service account for Cloud Run and grant it access to the secrets and Cloud SQL: ```bash # Create service account gcloud iam service-accounts create run-identity \ --display-name="Cloud Run Service Account" # Grant access to the database password secret gcloud secrets add-iam-policy-binding db-password \ --member="serviceAccount:run-identity@[YOUR_PROJECT_ID].iam.gserviceaccount.com" \ --role="roles/secretmanager.secretAccessor" # Grant access to the Conductor API key secret (if using Conductor) gcloud secrets add-iam-policy-binding conductor-api-key \ --member="serviceAccount:run-identity@[YOUR_PROJECT_ID].iam.gserviceaccount.com" \ --role="roles/secretmanager.secretAccessor" # Grant Cloud SQL client role gcloud projects add-iam-policy-binding [YOUR_PROJECT_ID] \ --member="serviceAccount:run-identity@[YOUR_PROJECT_ID].iam.gserviceaccount.com" \ --role="roles/cloudsql.client" ``` When deploying with `--source`, Cloud Build runs under the project's default Compute Engine service account, not `run-identity`. This account needs the `cloudbuild.builds.builder` role to build and push container images: ```bash # [YOUR_PROJECT_NUMBER] is the numeric project number (not the project ID) # Find it in the Google Cloud Console under project settings gcloud projects add-iam-policy-binding [YOUR_PROJECT_ID] \ --member="serviceAccount:[YOUR_PROJECT_NUMBER]-compute@developer.gserviceaccount.com" \ --role="roles/cloudbuild.builds.builder" ``` --- **Sample Dockerfile** Multi-stage build with a distroless runtime image: ```dockerfile title="Dockerfile" # --- Build Stage --- FROM golang:1.24 as builder WORKDIR /app COPY go.mod go.sum ./ RUN go mod download && go mod tidy COPY . . RUN CGO_ENABLED=0 GOOS=linux go build -o main . # --- Run Stage --- FROM gcr.io/distroless/static-debian12 COPY --from=builder /app/main / EXPOSE 8080 CMD ["/main"] ``` You can test the build locally against a local PostgreSQL instance before deploying: ```bash # Build the image docker build -t dbos-go-starter-image . # Run with a local Postgres docker run --rm -p 8080:8080 \ -e DB_USER=postgres \ -e DB_PASSWORD="your_local_db_password" \ -e DB_NAME=myappdb \ -e INSTANCE_UNIX_SOCKET=host.docker.internal \ -e DBOS_CONDUCTOR_KEY="your_conductor_api_key" \ dbos-go-starter-image ``` Then hit `http://localhost:8080/workflow/1` to start a DBOS workflow. --- **Deploy** Deploy from source—Cloud Build automatically builds your container and pushes it to Artifact Registry. **Service** ```bash gcloud run deploy my-app \ --source . \ --region us-central1 \ --no-cpu-throttling \ --min-instances=1 \ --service-account run-identity@[YOUR_PROJECT_ID].iam.gserviceaccount.com \ --network main-vpc \ --subnet run-subnet \ --vpc-egress private-ranges-only \ --add-cloudsql-instances [YOUR_PROJECT_ID]:us-central1:my-postgres-instance \ --set-secrets DB_PASSWORD=db-password:latest,DBOS_CONDUCTOR_KEY=conductor-api-key:latest \ --set-env-vars DB_USER=postgres,DB_NAME=myappdb,INSTANCE_UNIX_SOCKET=/cloudsql/[YOUR_PROJECT_ID]:us-central1:my-postgres-instance \ --allow-unauthenticated ``` Key flags: - **`--no-cpu-throttling`** Enables [instance-based billing](https://docs.cloud.google.com/run/docs/configuring/billing-settings) so DBOS background services keep running between requests. - **`--min-instances=1`** Keeps one instance always on so background services never stop. - **`--set-secrets`** Injects secrets from Secret Manager as environment variables. - **`--add-cloudsql-instances`** Mounts the Cloud SQL Auth Proxy socket, letting the app connect via `INSTANCE_UNIX_SOCKET`. - **`--source .`** Builds your Dockerfile remotely via [Cloud Build](http://cloud.google.com/build). - **`--allow-unauthenticated`** Makes the service publicly accessible. **Worker Pool** ```bash gcloud beta run worker-pools deploy my-app \ --source . \ --region us-central1 \ --instances=1 \ --service-account run-identity@[YOUR_PROJECT_ID].iam.gserviceaccount.com \ --network main-vpc \ --subnet run-subnet \ --vpc-egress private-ranges-only \ --add-cloudsql-instances [YOUR_PROJECT_ID]:us-central1:my-postgres-instance \ --set-secrets DB_PASSWORD=db-password:latest,DBOS_CONDUCTOR_KEY=conductor-api-key:latest \ --set-env-vars DB_USER=postgres,DB_NAME=myappdb,INSTANCE_UNIX_SOCKET=/cloudsql/[YOUR_PROJECT_ID]:us-central1:my-postgres-instance,GCP_PROJECT_ID=[YOUR_PROJECT_ID],GCP_REGION=us-central1,WORKER_POOL_NAME=my-app ``` Key flags: - **`--set-secrets`** Injects secrets from Secret Manager as environment variables. - **`--add-cloudsql-instances`** Mounts the Cloud SQL Auth Proxy socket, letting the app connect via `INSTANCE_UNIX_SOCKET`. - **`--source .`** Builds your Dockerfile remotely via [Cloud Build](https://cloud.google.com/build). - **`--instances=1`** Initial always-on instance count. The [scaling workflow](#scaling-a-worker-pool) adjusts this based on queue depth. - **`GCP_PROJECT_ID`**, **`GCP_REGION`**, **`WORKER_POOL_NAME`** Used by the scaling workflow to call the Cloud Run API. --- **Build logs** During deployment, `gcloud` streams Cloud Build output to your terminal. You can also view logs in the [Cloud Build console](https://console.cloud.google.com/cloud-build/builds) or with: ```bash gcloud builds list --limit=5 --region=us-central1 gcloud builds log [BUILD_ID] --region=us-central1 ``` **Service URL (service mode only)** On successful deployment, `gcloud` prints the service URL: ``` Service URL: https://my-app-XXXXXXXXXX.us-central1.run.app ``` Retrieve it later with: ```bash gcloud run services describe my-app --region us-central1 --format='value(status.url)' ``` **Application logs** For a **service**: ```bash gcloud logging read \ 'resource.type=cloud_run_revision AND resource.labels.service_name=my-app' \ --limit 100 --format='text' ``` For a **worker pool**: ```bash gcloud logging read \ 'resource.type=cloud_run_worker_pool AND resource.labels.worker_pool_name=my-app' \ --limit 100 --format='text' ``` Or view logs in the [Cloud Run console](https://console.cloud.google.com/run) under the **Logs** tab. **Test the deployment (service mode only)** Start a DBOS workflow: ```bash curl -X GET https://my-app-XXXXXXXXXX.us-central1.run.app/workflow/1 ``` This runs the three-step `ExampleWorkflow` with task ID `1`. Each step takes 5 seconds. Poll progress with: ```bash curl -X GET https://my-app-XXXXXXXXXX.us-central1.run.app/last_step/1 ``` Returns `1`, `2`, or `3` depending on how many steps have completed.
### Scaling a Worker Pool Worker pools don't auto-scale, but you can build an **external scaler** inside the pool using a DBOS [scheduled workflow](../golang/tutorials/scheduled-workflows.md). DBOS guarantees only one instance runs a scheduled function at a time—even across a multi-instance pool—preventing a thundering herd of conflicting resize requests. The complete implementation is in the [cloud-run demo app](https://github.com/dbos-inc/dbos-demo-apps/tree/main/golang/cloudrun). :::tip IAM permissions The worker pool's service account needs permission to manage Cloud Run resources and to act as itself when creating new revisions.
IAM commands ```bash # Grant Cloud Run admin role gcloud projects add-iam-policy-binding [YOUR_PROJECT_ID] \ --member="serviceAccount:run-identity@[YOUR_PROJECT_ID].iam.gserviceaccount.com" \ --role="roles/run.admin" # Grant actAs permission on the service account itself gcloud iam service-accounts add-iam-policy-binding \ run-identity@[YOUR_PROJECT_ID].iam.gserviceaccount.com \ --member="serviceAccount:run-identity@[YOUR_PROJECT_ID].iam.gserviceaccount.com" \ --role="roles/iam.serviceAccountUser" ```
::: The scheduled workflow periodically checks the queue depth and resizes the pool to match by calling the [Cloud Run Admin API](https://cloud.google.com/run/docs/reference/rest/v2/projects.locations.workerPools). It authenticates with a short-lived access token from the [GCE metadata server](https://cloud.google.com/compute/docs/metadata/overview), reads the current instance count with a `GET`, and updates it with a `PATCH`. Here's an example in Go (the same approach works in any DBOS-supported language):
Scaling workflow snippet ```go title="main.go" func ScalingWorkflow(ctx dbos.DBOSContext, input dbos.ScheduledWorkflowInput) (any, error) { // 1. Read queue length by listing all enqueued/pending workflows workflows, err := dbos.ListWorkflows(ctx, dbos.WithQueuesOnly(), dbos.WithQueueName(taskQueue.Name)) if err != nil { return "", fmt.Errorf("failed to list workflows: %w", err) } qlen := len(workflows) // 2. Get current instance count from the Cloud Run Admin API currentInstances, err := dbos.RunAsStep(ctx, func(stepCtx context.Context) (int, error) { return getWorkerPoolInstances(stepCtx) }) if err != nil { return "", fmt.Errorf("failed to get current instances: %w", err) } // 3. Compute desired instances: ceil(queue_depth / worker_concurrency) desiredInstances := int(math.Ceil(float64(qlen) / float64(WORKER_CONCURRENCY))) if desiredInstances < 1 { desiredInstances = 1 } // 4. Resize the pool if needed if desiredInstances != currentInstances { _, err := dbos.RunAsStep(ctx, func(stepCtx context.Context) (string, error) { return setWorkerPoolInstances(stepCtx, desiredInstances) }) if err != nil { return "", fmt.Errorf("failed to set instances: %w", err) } } return fmt.Sprintf("qlen=%d, instances=%d", qlen, desiredInstances), nil } ```
### Upgrading Workflow Code Deploying new code to Cloud Run creates a new **revision**. By default, Cloud Run routes all traffic to the latest revision immediately. Understanding how revisions interact with [upgrading DBOS code](../golang/tutorials/upgrading-workflows.md) is key to safely deploying changes without disrupting in-progress workflows. DBOS supports two strategies for deploying breaking changes: **versioning** and **patching**. Each maps differently to Cloud Run's revision model depending on whether you run a Service or a Worker Pool. #### Cloud Run revisions Every `gcloud run deploy` or `gcloud beta run worker-pools deploy` creates a new revision (e.g., `my-app-00001-abc`). Cloud Run injects the revision name into every container as the `K_REVISION` environment variable. For **services**, you can [split traffic](https://cloud.google.com/run/docs/rollouts-rollbacks-traffic-migration) between revisions, enabling blue-green or canary deployments. By default, 100% of traffic goes to the latest revision. For **worker pools**, a new deployment replaces all running instances. Old instances are shut down regardless of what they were processing. #### Service mode ##### Versioning Set [`ApplicationVersion`](../golang/reference/dbos-context.md#newcontext) to `K_REVISION` so each Cloud Run revision gets a distinct DBOS version. Workflows started on a revision are tagged with that revision's version. A DBOS process only recovers workflows matching its own version, so old workflows won't be replayed with new code. To drain old workflows, keep the previous revision active (with a share of traffic or `--min-instances=1`) until all its workflows complete. You can check with [`ListWorkflows`](../golang/reference/methods.md#listworkflows). ##### Patching With [patching](../golang/tutorials/upgrading-workflows.md#patching), fix the application version to a constant and enable patching in the [DBOS configuration](../golang/reference/dbos-context.md#newcontext). Since all revisions share the same version, new containers automatically recover in-progress workflows from previous deployments. Cloud Run routes traffic to the latest revision by default, so new requests go to the new code while the patching logic in your workflow handles the transition for recovered workflows. #### Worker pool mode When you deploy a new worker pool revision, Cloud Run replaces all running instances. Old instances shut down, and new instances start with the new code. ##### Versioning If `ApplicationVersion` is set to `K_REVISION`, the new instances have a different version than workflows started by the old instances. Those in-progress workflows won't be automatically recovered because the version doesn't match. To migrate them, [fork](../golang/tutorials/workflow-management.md#forking-workflows) the old workflows to the new version using [`ForkWorkflow`](../golang/reference/methods.md#forkworkflow) with the new `ApplicationVersion`. The new workers will then execute the forked workflows. You can automate this as part of a post-deployment step or a startup routine that lists old-version workflows and forks them. ##### Patching With a fixed application version and patching enabled, the new worker pool instances automatically recover workflows from the previous deployment. [Conductor](../conductor/overview.md) detects that the old instances went down and that new instances with the same version are available, triggering recovery without any manual intervention. #### Advanced scenarios More complex deployment strategies are possible. You can combine versioning and patching—for example, using versioning for major changes and patching for hotfixes within a version. In Service mode, you can use Cloud Run [revision tags](https://cloud.google.com/run/docs/rollouts-rollbacks-traffic-migration#tags) to route a subset of traffic to a tagged revision, letting you test new workflow code in production before shifting all traffic. --- ## Deploying With Kubernetes This guide covers deploying a DBOS application on Kubernetes. It walks through DBOS-specific deployment concepts, then provides a full walkthrough covering infrastructure, secrets, database migrations, application deployment, and autoscaling. The Kubernetes manifests are portable to any conformant cluster. --- ### Deployment DBOS is a library — it does not require any sidecar, operator, or external service besides PostgreSQL. Pods are stateless and interchangeable and should use a standard [Deployment](https://kubernetes.io/docs/concepts/workloads/controllers/deployment/). ### Configuration DBOS configuration contains sensitive values: the [system database URL](../explanations/system-tables.md) and, if using [Conductor](../conductor/overview.md), an API key. Store these as [Kubernetes Secrets](https://kubernetes.io/docs/concepts/configuration/secret/) and inject them via [`secretKeyRef`](https://kubernetes.io/docs/concepts/configuration/secret/#using-secrets-as-environment-variables). For Git-safe storage, encrypt with [Sealed Secrets](https://github.com/bitnami-labs/sealed-secrets), [SOPS](https://github.com/getsops/sops), or a cloud-native secrets manager. :::info Connecting to DBOS Conductor If you use [DBOS-hosted Conductor](https://console.dbos.dev/), no `DBOS_CONDUCTOR_URL` is needed. The SDK connects automatically. If you [self-host Conductor](../conductor/self-hosting/hosting-conductor.md), set `DBOS_CONDUCTOR_URL` in your application's environment. When Conductor is in a different cluster, use `wss://` so the WebSocket connection is encrypted. In the same cluster, use `ws://`, as Conductor requires TLS termination at the ingress layer. ::: ### Database Privilege Separation DBOS applications store workflow state in [system tables](../explanations/system-tables.md). These tables must be created before the application can start. Run [`dbosctl sysdb migrate`](../conductor/reference/dbosctl.md#dbosctl-sysdb-migrate) with an **admin** role that can create schema and grant permissions, and run the application with a **restricted** role that can only read/write data. Use the `--app-role` flag to grant the necessary schema permissions to the restricted role. `dbosctl sysdb migrate` works well as a Kubernetes [Job](https://kubernetes.io/docs/concepts/workloads/controllers/job/) that you compose into your CI/CD pipeline. It is a single static binary carrying the migrations, so the Job needs no SDK toolchain and no copy of your application. ### Availability In addition to [general tips](./checklist.md) for running a DBOS-enabled app in production: **Readiness probe** — have the probe wait until DBOS is launched before Kubernetes routes traffic to pods. **Resource limits** — DBOS doesn't add significant CPU or memory overhead, but all DBOS SDKs run background tasks; setting more than 1000m CPU can significantly improve the performance of a busy application. **Replicas** — configure more than one replica. Each replica starts an independent DBOS worker that can process scheduled workflows and handle tasks from DBOS queues. Each replica should have a unique executor ID (which is automatically assigned when using [DBOS Conductor](../conductor/overview.md)) ### Upgrading Workflow Code DBOS workflows can run for weeks or years while the underlying code evolves. Two patterns support this: 1. **Application versioning** — DBOS SDKs store a version number alongside each workflow record. You should create a separate Deployment per active version. Point the Service selector at the latest version only, such that new HTTP requests creating DBOS workflows go exclusively to the new Deployment. Old deployments will stay alive and keep executing pending workflows. Once workflows for an old version complete, delete its Deployment. 2. **Workflow patching** — Keep a single Deployment. Add conditional logic (patches) that detect which code path a recovering workflow should take. See the [workflow patching guide](../python/tutorials/upgrading-workflows.md) for details. ### Scaling with KEDA [KEDA](https://keda.sh/) scales application pods based on external metrics. A simple pattern for scaling based on DBOS queue depth. When using [DBOS Conductor](../conductor/overview.md), you can install an [autoscaling policy](../conductor/autoscaling.md#attaching-a-policy-with-the-api) for your application and configure a KEDA [ScaledObject](https://keda.sh/docs/latest/concepts/scaling-deployments/) to size your application based on the [policy recommendation](../conductor/autoscaling.md#reading-the-desired-executor-count). --- ### Walkthrough (AWS EKS) **EKS (AWS)** This walkthrough deploys a sample DBOS Go application on EKS with RDS PostgreSQL, Sealed Secrets, database migrations, and KEDA autoscaling.
Set environment variables Set these variables before proceeding — replace the placeholder values with your own: ```bash # Your AWS account ID (12-digit number) AWS_ACCOUNT_ID=123456789012 # AWS region for all resources AWS_REGION=us-west-2 # PostgreSQL admin password (used for the RDS master user) POSTGRES_PASSWORD='choose-a-secure-password' # Password for the restricted application database role APP_ROLE_PASSWORD='choose-another-secure-password' # Conductor API key (from the Console after registering your app) CONDUCTOR_API_KEY='your-api-key' # Conductor URL # DBOS-hosted Conductor: wss://cloud.dbos.dev/conductor/v1alpha1 # Self-hosted (same cluster): ws://conductor.dbos.svc.cluster.local:8090 # Self-hosted (external): wss://your-conductor-hostname/conductor/ CONDUCTOR_URL='wss://cloud.dbos.dev/conductor/v1alpha1' ```
#### Infrastructure
CLI tools required on your workstation | Tool | Purpose | Install | |------|---------|---------| | **AWS CLI** | AWS account access | [Install guide](https://docs.aws.amazon.com/cli/latest/userguide/getting-started-install.html) | | **eksctl** | Create and manage EKS clusters | [Install guide](https://eksctl.io/installation/) | | **kubectl** | Interact with Kubernetes | Included with eksctl, or [install separately](https://kubernetes.io/docs/tasks/tools/) | | **Helm** | Install cluster add-ons (KEDA, Sealed Secrets) | [Install guide](https://helm.sh/docs/intro/install/) | | **kubeseal** | Encrypt Kubernetes secrets | [Install guide](https://github.com/bitnami-labs/sealed-secrets#kubeseal) | | **Go** | Build the DBOS application | [Install guide](https://go.dev/doc/install) | | **Docker** | Build application container images | [Install guide](https://docs.docker.com/get-docker/) | Verify your AWS credentials are configured: ```bash aws sts get-caller-identity ```
**DBOS Conductor** This walkthrough connects the application to [DBOS Conductor](../conductor/overview.md) for workflow recovery and observability. You can use either [DBOS-hosted Conductor](https://console.dbos.dev/) or a [self-hosted Conductor](../conductor/self-hosting/hosting-conductor-with-kubernetes.md). You'll need the **Conductor URL** and an **API key** — both are available from the Console after [registering your application](../conductor/overview.md#connecting-to-conductor). **Create an EKS Cluster** Create a managed EKS cluster with two nodes. This takes approximately 15 minutes.
Create EKS cluster ```bash eksctl create cluster \ --name dbos-app-cluster \ --region $AWS_REGION \ --version 1.31 \ --nodegroup-name default \ --node-type t3.medium \ --nodes 2 \ --managed ``` `eksctl` automatically: - Creates a VPC with public and private subnets - Configures the [Amazon VPC CNI](https://docs.aws.amazon.com/eks/latest/userguide/managing-vpc-cni.html) - Sets up your `~/.kube/config` to point at the new cluster Once complete, verify the cluster is ready: ```bash kubectl get nodes ``` You should see two nodes in `Ready` status.
**Create a Namespace** All resources in this walkthrough are deployed to a dedicated `dbos` namespace: ```bash kubectl create namespace dbos ``` **Provision an RDS PostgreSQL Instance** Your DBOS application needs a PostgreSQL database for its [system tables](../explanations/system-tables.md).
RDS provisioning commands Find the VPC and private subnets that `eksctl` created: ```bash # Get the VPC ID VPC_ID=$(aws ec2 describe-vpcs \ --filters "Name=tag:alpha.eksctl.io/cluster-name,Values=dbos-app-cluster" \ --query "Vpcs[0].VpcId" --output text --region $AWS_REGION) echo "VPC: $VPC_ID" # Get the private subnets (array for bash/zsh compatibility) PRIVATE_SUBNETS=($(aws ec2 describe-subnets \ --filters "Name=vpc-id,Values=$VPC_ID" \ "Name=tag:aws:cloudformation:logical-id,Values=SubnetPrivate*" \ --query "Subnets[*].SubnetId" --output text --region $AWS_REGION)) echo "Private subnets: ${PRIVATE_SUBNETS[@]}" ``` Create a DB subnet group from the private subnets: ```bash aws rds create-db-subnet-group \ --db-subnet-group-name dbos-app-db \ --db-subnet-group-description "DBOS application RDS subnets" \ --subnet-ids "${PRIVATE_SUBNETS[@]}" \ --region $AWS_REGION ``` Create a security group that allows PostgreSQL access from the EKS nodes: ```bash # Get the EKS cluster security group EKS_SG=$(aws ec2 describe-security-groups \ --filters "Name=vpc-id,Values=$VPC_ID" \ "Name=tag:aws:eks:cluster-name,Values=dbos-app-cluster" \ --query "SecurityGroups[0].GroupId" \ --output text --region $AWS_REGION) echo "EKS SG: $EKS_SG" # Create a security group for RDS RDS_SG=$(aws ec2 create-security-group \ --group-name dbos-app-rds \ --description "Allow PostgreSQL from EKS nodes" \ --vpc-id $VPC_ID \ --query "GroupId" --output text --region $AWS_REGION) echo "RDS SG: $RDS_SG" # Allow inbound PostgreSQL from EKS nodes aws ec2 authorize-security-group-ingress \ --group-id $RDS_SG \ --protocol tcp --port 5432 \ --source-group $EKS_SG \ --region $AWS_REGION ``` Create the RDS instance: ```bash aws rds create-db-instance \ --db-instance-identifier dbos-app-pg \ --db-instance-class db.t4g.micro \ --engine postgres \ --engine-version 16 \ --master-username postgres \ --master-user-password "$POSTGRES_PASSWORD" \ --allocated-storage 20 \ --db-subnet-group-name dbos-app-db \ --vpc-security-group-ids $RDS_SG \ --no-publicly-accessible \ --region $AWS_REGION ``` Wait for the instance to become available (this takes a few minutes): ```bash aws rds wait db-instance-available \ --db-instance-identifier dbos-app-pg \ --region $AWS_REGION ``` Get the RDS endpoint: ```bash RDS_ENDPOINT=$(aws rds describe-db-instances \ --db-instance-identifier dbos-app-pg \ --query "DBInstances[0].Endpoint.Address" \ --output text --region $AWS_REGION) echo "RDS endpoint: $RDS_ENDPOINT" ```
Create the database and application role from a pod inside the cluster (since the RDS instance is not publicly accessible):
Create database and role ```bash kubectl run pg-setup --restart=Never \ --namespace dbos \ --image=postgres:16 \ --env="PGPASSWORD=$POSTGRES_PASSWORD" \ --command -- bash -c " psql -h $RDS_ENDPOINT -U postgres -c 'CREATE DATABASE dbos_app;' psql -h $RDS_ENDPOINT -U postgres -c \"CREATE ROLE dbos_app_role WITH LOGIN PASSWORD '$APP_ROLE_PASSWORD';\" " # Wait for the pod to finish, then clean up sleep 15 && kubectl logs pg-setup -n dbos && kubectl delete pod pg-setup -n dbos ``` This creates: - `dbos_app` — the application's system database for workflow state - `dbos_app_role` — a restricted role the application uses at runtime (granted permissions by `dbosctl sysdb migrate`)
**Install Cluster Add-ons**
Helm installs (Sealed Secrets, KEDA) **Sealed Secrets** — encrypt secrets for safe Git storage: ```bash helm repo add sealed-secrets https://bitnami-labs.github.io/sealed-secrets helm install sealed-secrets sealed-secrets/sealed-secrets \ --namespace kube-system ``` **KEDA** — event-driven autoscaling: ```bash helm repo add kedacore https://kedacore.github.io/charts helm repo update helm install keda kedacore/keda \ --namespace keda --create-namespace ``` Verify both add-ons are running: ```bash # Sealed Secrets controller kubectl get pods -n kube-system -l app.kubernetes.io/name=sealed-secrets # KEDA kubectl get pods -n keda ```
**Create ECR Repositories** We push two container images to Amazon ECR — one for the application and one for the migration job: ```bash aws ecr create-repository --repository-name dbos-app --region $AWS_REGION aws ecr create-repository --repository-name dbos-migrate --region $AWS_REGION ``` Note the repository URIs from the output (e.g., `${AWS_ACCOUNT_ID}.dkr.ecr.${AWS_REGION}.amazonaws.com/dbos-app`). The commands below use `$AWS_ACCOUNT_ID` and `$AWS_REGION`, which you set earlier. Authenticate Docker with ECR (tokens expire after 12 hours): ```bash aws ecr get-login-password --region $AWS_REGION | \ docker login --username AWS --password-stdin \ ${AWS_ACCOUNT_ID}.dkr.ecr.${AWS_REGION}.amazonaws.com ``` #### Secrets Several components need sensitive credentials. We use [Bitnami Sealed Secrets](https://github.com/bitnami-labs/sealed-secrets): create a regular Secret, encrypt it with `kubeseal`, and apply the encrypted `SealedSecret` to the cluster. The controller decrypts it in-cluster into a standard Kubernetes Secret that pods can reference. The encrypted form is safe to commit to Git. **Secrets Inventory** | Secret | Keys | Used by | |--------|------|---------| | `postgres-admin` | `password`, `database-url` | Migration Job — admin access to create/update DBOS system tables | | `dbos-app-db` | `database-url` | DBOS Application — restricted access to `dbos_app` database | | `conductor-api-key` | `api-key` | DBOS Application — authenticates with Conductor | **Create and Seal Secrets**
kubeseal commands for all 3 secrets Create each secret, pipe it through `kubeseal`, and save the encrypted form: ```bash # 1. PostgreSQL admin credentials (used by the migration Job) kubectl create secret generic postgres-admin \ --namespace dbos \ --from-literal=password="$POSTGRES_PASSWORD" \ --from-literal=database-url="postgresql://postgres:${POSTGRES_PASSWORD}@${RDS_ENDPOINT}:5432/dbos_app?sslmode=require" \ --dry-run=client -o yaml | \ kubeseal --controller-name=sealed-secrets --controller-namespace=kube-system --format yaml \ > sealed-postgres-admin.yaml # 2. DBOS application database credentials (restricted role) kubectl create secret generic dbos-app-db \ --namespace dbos \ --from-literal=database-url="postgresql://dbos_app_role:${APP_ROLE_PASSWORD}@${RDS_ENDPOINT}:5432/dbos_app?sslmode=require" \ --dry-run=client -o yaml | \ kubeseal --controller-name=sealed-secrets --controller-namespace=kube-system --format yaml \ > sealed-dbos-app-db.yaml # 3. Conductor API key kubectl create secret generic conductor-api-key \ --namespace dbos \ --from-literal=api-key="$CONDUCTOR_API_KEY" \ --dry-run=client -o yaml | \ kubeseal --controller-name=sealed-secrets --controller-namespace=kube-system --format yaml \ > sealed-conductor-api-key.yaml ```
**Apply and Verify** ```bash kubectl apply -f sealed-postgres-admin.yaml kubectl apply -f sealed-dbos-app-db.yaml kubectl apply -f sealed-conductor-api-key.yaml ``` Verify the controller has decrypted them into regular Kubernetes Secrets: ```bash kubectl get secrets -n dbos ``` You should see all three secrets with type `Opaque`: ``` NAME TYPE DATA AGE conductor-api-key Opaque 1 10s dbos-app-db Opaque 1 10s postgres-admin Opaque 2 10s ``` #### Database Migrations DBOS applications store workflow state in [system tables](../explanations/system-tables.md). These tables must be created before the application can start. We use a separate Kubernetes Job that runs [`dbosctl sysdb migrate`](../conductor/reference/dbosctl.md#dbosctl-sysdb-migrate) with **admin** credentials, then the application itself runs with a **restricted** role that can only read/write data — not modify schema. This separation follows the principle of least privilege: the application never holds the keys to alter its own schema.
Migration image and build The migration image contains only `dbosctl` — it doesn't include your application code, or a toolchain for the language that code is written in. `dbosctl` ships as a statically linked release binary carrying the migrations, so the image is a download rather than a build: ```dockerfile title="Dockerfile.migrate" FROM alpine:latest # Pin this to the dbosctl release you have tested; leave it empty for the latest. ARG DBOSCTL_VERSION="" RUN apk --no-cache add ca-certificates curl RUN curl -sSfL https://raw.githubusercontent.com/dbos-inc/dbos-ctl/main/install.sh \ | VERSION="${DBOSCTL_VERSION}" BIN_DIR=/usr/local/bin sh ENTRYPOINT ["dbosctl"] ``` Pin `DBOSCTL_VERSION` for a pipeline you want to be reproducible: an unpinned build takes whatever the latest release is on the day it runs, which is not what you tested. ```bash # Set your ECR repository URI ECR_MIGRATE=${AWS_ACCOUNT_ID}.dkr.ecr.${AWS_REGION}.amazonaws.com/dbos-migrate # Build for linux/amd64 docker build --platform linux/amd64 \ -t ${ECR_MIGRATE}:latest \ -f Dockerfile.migrate . # Push to ECR docker push ${ECR_MIGRATE}:latest ```
The Job runs `dbosctl sysdb migrate --app-role dbos_app_role`, which: 1. Creates the DBOS system tables in the `dbos_app` database (if they don't exist) 2. Applies any pending schema migrations 3. Grants the necessary permissions to `dbos_app_role` so the application can read and write workflow state
manifests/migrate-job.yaml ```yaml apiVersion: batch/v1 kind: Job metadata: name: dbos-migrate namespace: dbos spec: backoffLimit: 3 template: spec: restartPolicy: OnFailure containers: - name: migrate image: ${AWS_ACCOUNT_ID}.dkr.ecr.${AWS_REGION}.amazonaws.com/dbos-migrate:latest args: - "sysdb" - "migrate" - "--app-role" - "dbos_app_role" env: - name: DBOS_SYSTEM_DATABASE_URL valueFrom: secretKeyRef: name: postgres-admin key: database-url ``` The `postgres-admin` secret contains the admin connection string, which has the privileges needed to create tables and grant permissions. The `--app-role` flag tells `dbosctl sysdb migrate` to grant the specified role access to the system tables.
**Run the Migration** ```bash kubectl apply -f manifests/migrate-job.yaml ``` Wait for the job to complete: ```bash kubectl get jobs -n dbos ``` ``` NAME STATUS COMPLETIONS DURATION AGE dbos-migrate Complete 1/1 8s 30s ``` Check the logs to confirm the migration succeeded: ```bash kubectl logs -n dbos job/dbos-migrate ``` **Re-running Migrations** When a new `dbosctl` release adds migrations — usually alongside a DBOS SDK version that expects them — rebuild the migration image against that release and re-run the Job. Re-running a Job built from the same pinned release is harmless but does nothing: `dbosctl sysdb migrate` skips migrations that are already recorded. Since Kubernetes Job names must be unique, delete the old job first: ```bash kubectl delete job dbos-migrate -n dbos kubectl apply -f manifests/migrate-job.yaml ``` In a CI/CD pipeline, you would typically give each migration job a unique name (e.g., `dbos-migrate-v2`) or use a Helm hook with `hook-delete-policy: before-hook-creation`. #### Application Deployment Build and push the application image to ECR, then apply the application manifest.
Dockerfile, manifest, and ECR push ```dockerfile title="Dockerfile" FROM golang:1.25-alpine AS builder WORKDIR /app COPY go.mod go.sum ./ RUN go mod download COPY . . RUN CGO_ENABLED=0 GOOS=linux go build -o dbos-app . FROM alpine:latest RUN apk --no-cache add ca-certificates WORKDIR /app COPY --from=builder /app/dbos-app . EXPOSE 8080 CMD ["./dbos-app"] ``` ```bash # Set your ECR repository URI ECR_REPO=${AWS_ACCOUNT_ID}.dkr.ecr.${AWS_REGION}.amazonaws.com/dbos-app # Build for linux/amd64 (EKS nodes run Linux) docker build --platform linux/amd64 -t ${ECR_REPO}:latest . # Push to ECR docker push ${ECR_REPO}:latest ``` ```yaml title="manifests/dbos-app.yaml" apiVersion: apps/v1 kind: Deployment metadata: name: dbos-app namespace: dbos spec: replicas: 1 selector: matchLabels: app: dbos-app template: metadata: labels: app: dbos-app spec: containers: - name: dbos-app image: ${AWS_ACCOUNT_ID}.dkr.ecr.${AWS_REGION}.amazonaws.com/dbos-app:latest env: - name: DBOS_SYSTEM_DATABASE_URL valueFrom: secretKeyRef: name: dbos-app-db key: database-url - name: DBOS_CONDUCTOR_URL value: "${CONDUCTOR_URL}" - name: DBOS_CONDUCTOR_KEY valueFrom: secretKeyRef: name: conductor-api-key key: api-key ports: - containerPort: 8080 readinessProbe: httpGet: path: /healthz port: 8080 initialDelaySeconds: 5 periodSeconds: 10 livenessProbe: httpGet: path: /healthz port: 8080 initialDelaySeconds: 10 periodSeconds: 30 resources: requests: cpu: 500m memory: 256Mi limits: cpu: "2" memory: 512Mi --- apiVersion: v1 kind: Service metadata: name: dbos-app namespace: dbos spec: selector: app: dbos-app ports: - port: 8080 targetPort: 8080 ``` Replace `${CONDUCTOR_URL}` with the value you set earlier: - **DBOS-hosted Conductor**: `wss://cloud.dbos.dev/conductor/v1alpha1` - **Self-hosted (same cluster)**: `ws://conductor.dbos.svc.cluster.local:8090` - **Self-hosted (external)**: `wss:///conductor/` The database URL and API key are pulled from the Sealed Secrets created in the [Secrets](#secrets) section.
Trusting a self-signed TLS certificate If your self-hosted Conductor uses a **CA-signed certificate** (e.g., from [cert-manager](https://cert-manager.io/) with Let's Encrypt), no extra configuration is needed — the system CA bundle already trusts it. If Conductor uses a **self-signed certificate**, the application must explicitly trust it. Store the certificate as a Kubernetes Secret, mount it into an init container that appends it to the system CA bundle, and point the app container at the extended bundle via `SSL_CERT_FILE`: 1. Create a TLS secret from your certificate files: ```bash kubectl create secret tls dbos-tls \ --cert=tls.crt --key=tls.key --namespace dbos ``` 2. Add an init container and volumes to the Deployment spec: ```yaml initContainers: - name: setup-certs image: alpine:latest command: ["sh", "-c", "cp /etc/ssl/certs/ca-certificates.crt /certs/ca-certificates.crt && cat /tls/tls.crt >> /certs/ca-certificates.crt"] volumeMounts: - { name: tls-cert, mountPath: /tls, readOnly: true } - { name: ca-certs, mountPath: /certs } # ... app container with: # env: # - name: SSL_CERT_FILE # value: "/certs/ca-certificates.crt" # volumeMounts: # - { name: ca-certs, mountPath: /certs, readOnly: true } volumes: - name: tls-cert secret: secretName: dbos-tls items: - { key: tls.crt, path: tls.crt } - name: ca-certs emptyDir: {} ```
:::tip When using DBOS-hosted Conductor, you don't need to set `DBOS_CONDUCTOR_URL` in the manifest. ::: **Deploy the Application** ```bash kubectl apply -f manifests/dbos-app.yaml ``` Verify the pod is running: ```bash kubectl get pods -n dbos -l app=dbos-app kubectl logs -n dbos -l app=dbos-app --tail=20 ``` You should see Conductor connection messages in the logs. You can also verify the connection in the Console UI. To access the application locally, use port-forwarding: ```bash kubectl port-forward svc/dbos-app -n dbos 8080:8080 ``` Then in another terminal: ```bash curl http://localhost:8080/healthz ``` #### Scaling with KEDA [KEDA](https://keda.sh/) (Kubernetes Event-Driven Autoscaling) scales your application pods based on external metrics. In this section, KEDA polls the application's queue-depth endpoint and adjusts the replica count so that your application deployment has enough capacity to absorb load. **How the Metrics Endpoint Works** The sample application exposes `GET /metrics/:queueName`, which returns the number of workflows currently waiting on a queue: ```bash curl http://dbos-app.dbos.svc.cluster.local:8080/metrics/taskQueue # {"queue_length": 7} ``` KEDA uses the `metrics-api` trigger to poll this endpoint and extract `queue_length`. It computes the desired replica count as: ``` desiredReplicas = ceil(queue_length / targetValue) ``` With `targetValue: "2"` (matching the queue's `WithWorkerConcurrency(2)`), each pod handles two concurrent workflows. For example, if 7 workflows are queued, KEDA scales to `ceil(7 / 2) = 4` pods. **ScaledObject Manifest**
manifests/keda-scaledobject.yaml ```yaml apiVersion: keda.sh/v1alpha1 kind: ScaledObject metadata: name: dbos-app-scaledobject namespace: dbos spec: scaleTargetRef: name: dbos-app pollingInterval: 15 cooldownPeriod: 60 minReplicaCount: 1 maxReplicaCount: 10 triggers: - type: metrics-api metadata: url: "http://dbos-app.dbos.svc.cluster.local:8080/metrics/taskQueue" valueLocation: "queue_length" targetValue: "2" ``` | Field | Value | Description | |-------|-------|-------------| | `scaleTargetRef.name` | `dbos-app` | The Deployment to scale | | `pollingInterval` | `15` | Seconds between metric checks (default 30) | | `cooldownPeriod` | `60` | Seconds to wait after the last trigger activation before scaling down (default 300) | | `minReplicaCount` | `1` | Minimum replicas — must be ≥1 because the `metrics-api` trigger polls the app itself | | `maxReplicaCount` | `10` | Upper bound for the replica count | | `url` | `http://dbos-app...` | In-cluster URL to the app's queue metrics endpoint | | `valueLocation` | `queue_length` | JSON field to extract from the response | | `targetValue` | `"2"` | Desired metric value per replica — matches `WithWorkerConcurrency(2)` |
**Apply and Verify** ```bash kubectl apply -f manifests/keda-scaledobject.yaml ``` Verify the ScaledObject is ready: ```bash kubectl get scaledobject -n dbos ``` ``` NAME SCALETARGETKIND SCALETARGETNAME MIN MAX TRIGGERS AUTHENTICATION READY ACTIVE FALLBACK AGE dbos-app-scaledobject apps/v1.Deployment dbos-app 1 10 metrics-api True False False 10s ``` KEDA auto-creates a Horizontal Pod Autoscaler (HPA) under the hood: ```bash kubectl get hpa -n dbos ``` ``` NAME REFERENCE TARGETS MINPODS MAXPODS REPLICAS AGE keda-hpa-dbos-app-scaledobject Deployment/dbos-app 0/2 (avg) 1 10 1 10s ``` **Test Autoscaling** With port-forwarding active (`kubectl port-forward svc/dbos-app -n dbos 8080:8080`), enqueue several long-running workflows to build up queue depth. Each call enqueues a workflow that sleeps for 60 seconds: ```bash for i in $(seq 1 10); do curl -s http://localhost:8080/enqueue/60 done ``` Watch the pods scale up (in another terminal): ```bash kubectl get pods -n dbos -l app=dbos-app -w ``` You should see new pods appear as KEDA detects the growing queue: ``` NAME READY STATUS RESTARTS AGE dbos-app-xxxxxxxxx-aaaaa 1/1 Running 0 5m dbos-app-xxxxxxxxx-bbbbb 1/1 Running 0 15s dbos-app-xxxxxxxxx-ccccc 1/1 Running 0 15s dbos-app-xxxxxxxxx-ddddd 1/1 Running 0 15s dbos-app-xxxxxxxxx-eeeee 1/1 Running 0 15s ``` Check the current queue depth: ```bash curl -s http://localhost:8080/metrics/taskQueue ``` As workflows complete and the queue drains, the metric drops. After the `cooldownPeriod` (60 seconds of no trigger activation), KEDA scales back down to `minReplicaCount` (1). #### Cleanup To tear down all AWS resources when done: ```bash # Delete the EKS cluster (includes VPC, security groups, and node group) eksctl delete cluster --name dbos-app-cluster --region $AWS_REGION # Delete the RDS instance aws rds delete-db-instance --db-instance-identifier dbos-app-pg \ --skip-final-snapshot --region $AWS_REGION # Delete ECR repositories aws ecr delete-repository --repository-name dbos-app --force --region $AWS_REGION aws ecr delete-repository --repository-name dbos-migrate --force --region $AWS_REGION # Delete the RDS security group RDS_SG=$(aws ec2 describe-security-groups \ --filters "Name=group-name,Values=dbos-app-rds" \ --query "SecurityGroups[0].GroupId" --output text --region $AWS_REGION) aws ec2 delete-security-group --group-id $RDS_SG --region $AWS_REGION # Delete the DB subnet group aws rds delete-db-subnet-group --db-subnet-group-name dbos-app-db --region $AWS_REGION ``` --- ## Workflow Recovery When the execution of a durable workflow is interrupted (for example, if its executor is restarted, interrupted, or crashes), another executor must recover the workflow and resume its execution. To prevent duplicate work, it is important to detect interruptions promptly and to recover each workflow only once. This guide describes how to manage workflow recovery in a production environment. ### Managing Recovery #### Recovery On A Single Server If hosting an application on a single server without Conductor, each time you restart your application's process, DBOS recovers all workflows that were executing before the restart (all `PENDING` workflows). #### Recovery in a Distributed Setting When self-hosting in a distributed setting without Conductor, it is important to manage workflow recovery so that when an executor crashes, restarts, or is shut down, its workflows are recovered. You should assign each executor running a DBOS application an executor ID through DBOS configuration. Each workflow is tagged with the ID of the executor that started it. When an application with an executor ID restarts, it only recovers pending workflows assigned to that executor ID and owned by that application, so applications [sharing a system database](../explanations/sharing-a-system-database.md) never recover each other's workflows. #### Recovery With Conductor If your application is connected to [DBOS Conductor](../conductor/overview.md), workflow recovery is automatic: when Conductor detects that an executor is unhealthy, it signals another executor to recover its workflows. See [Distributed Recovery](../conductor/distributed-recovery.md) for details. --- ## Human-in-the-Loop This example shows how to use DBOS to add **human-in-the-loop** to your AI agent. Production AI agents often need to wait for human approval before performing critical tasks. However, because there are real people involved, approval doesn't always happen instantly, and agents need to be able to reliably wait hours or days for human intervention, then seamlessly resume when it arrives. This application demonstrates how to build reliable human-in-the-loop with durable workflows. We'll see how to build agents that can wait hours or days for human input to arrive (surviving process restarts). We'll also see how to use workflow introspection to monitor active agents and create an "inbox" of workflows that need approval. All source code is [available on GitHub](https://github.com/dbos-inc/dbos-demo-apps/tree/main/python/agent-inbox). ### Durable Agents First, we'll write the scaffold of a durable and observable agent with human-in-the-loop. This agent needs human approval for a key step. To get the approval, it calls [`DBOS.recv`](../tutorials/workflow-communication.md#workflow-messaging-and-notifications). This **durably waits** for a configurable length of time (potentially hours or days) for a notification to arrive, automatically recovering from the transient failures and process restarts that will inevitably occur during a long wait. To allow us to create an "inbox" of agents pending approval, the workflow also publishes its status using [`DBOS.set_event`](../tutorials/workflow-communication.md#workflow-events). We'll see later how this helps us build observability endpoints to list all active, completed, or failed agents. ```python @DBOS.workflow() def durable_agent(request: AgentStartRequest): # Set an agent status the frontend can query agent_status: AgentStatus = AgentStatus( name=request.name, task=request.task, status="working", created_at=datetime.now().isoformat(), question=f"Should I proceed with task: {request.task}?", ) DBOS.set_event(AGENT_STATUS, agent_status) print("Starting agent:", agent_status) # Do some work... # Upon reaching the step that needs approval, update status # to `pending_approval` and await an approval notification. agent_status.status = "pending_approval" DBOS.set_event(AGENT_STATUS, agent_status) approval: Optional[HumanResponseRequest] = DBOS.recv(timeout_seconds=3600) # If approved, continue execution. Otherwise, raise an exception # and terminate the agent. if approval is None: # If approval times out, treat it as a denial agent_status.status = "denied" DBOS.set_event(AGENT_STATUS, agent_status) print("Agent timed out:", agent_status) raise Exception("Agent timed out awaiting approval") elif approval.response == "deny": agent_status.status = "denied" DBOS.set_event(AGENT_STATUS, agent_status) print("Agent denied:", agent_status) raise Exception("Agent denied approval") else: agent_status.status = "working" print("Agent approved:", agent_status) DBOS.set_event(AGENT_STATUS, agent_status) # Do some more work... return "Agent successful" ``` ### Calling and Notifying the Agent Here's the endpoint to notify an agent it was approved or denied. It uses [`DBOS.send`](../tutorials/workflow-communication.md#workflow-messaging-and-notifications) to send a message to an agent awaiting approval, waking it up. ```python @app.post("/agents/{agent_id}/respond") def respond_to_agent(agent_id: str, response: HumanResponseRequest): # Notify an agent it has been approved or denied DBOS.send(agent_id, response) return {"ok": True} ``` Here's the endpoint to start a new agent. It starts the agent as a durable background task using `DBOS.start_workflow`. ```python @app.post("/agents") def start_agent(request: AgentStartRequest): # Start a durable agent in the background DBOS.start_workflow(durable_agent, request) return {"ok": True} ``` ### Building an Agent Inbox To build an "agent inbox", we need to be able to see which agents are pending approval. We can do this with the DBOS workflow introspection API. We list all active agents with [`DBOS.list_workflows`](../reference/contexts.md#list_workflows), then retrieve the status of each. We return only agents that currently have the `pending_approval` status. ```python @app.get("/agents/waiting", response_model=list[AgentStatus]) async def list_waiting_agents(): # List all active agents and retrieve their statuses agent_workflows = await DBOS.list_workflows_async( status="PENDING", name=durable_agent.__qualname__ ) statuses: list[AgentStatus] = await asyncio.gather( *[DBOS.get_event_async(w.workflow_id, AGENT_STATUS) for w in agent_workflows] ) for s, w in zip(statuses, agent_workflows): s.agent_id = w.workflow_id # Only return active agents that are currently awaiting human approval return [status for status in statuses if status.status == "pending_approval"] ``` We also build endpoints to list successful and failed agents: ```python @app.get("/agents/approved", response_model=list[AgentStatus]) async def list_approved_agents(): # List all successful agents and retrieve their statuses agent_workflows = await DBOS.list_workflows_async( status="SUCCESS", name=durable_agent.__qualname__ ) statuses = await asyncio.gather( *[DBOS.get_event_async(w.workflow_id, AGENT_STATUS) for w in agent_workflows] ) return list(statuses) @app.get("/agents/denied", response_model=list[AgentStatus]) async def list_denied_agents(): # List all failed agents and retrieve their statuses agent_workflows = await DBOS.list_workflows_async( status="ERROR", name=durable_agent.__qualname__ ) statuses = await asyncio.gather( *[DBOS.get_event_async(w.workflow_id, AGENT_STATUS) for w in agent_workflows] ) return list(statuses) ``` ### Try it Yourself! Clone and enter the [dbos-demo-apps](https://github.com/dbos-inc/dbos-demo-apps) repository: ```shell git clone https://github.com/dbos-inc/dbos-demo-apps.git cd dbos-demo-apps/python/agent-inbox ``` Then follow the instructions in the [README](https://github.com/dbos-inc/dbos-demo-apps/tree/main/python/agent-inbox) to run the app. --- ## Reliable Customer Service Agent In this example, you'll learn how to build a reliable AI-powered customer service agent with DBOS and [LangGraph](https://langchain-ai.github.io/langgraph/). This example demonstrates how **DBOS makes it easy to connect your AI agent to your existing production systems**, especially when integrating **human decision-making** into automated processes. You can chat with this LLM-powered AI agent to check the status of your purchase order, or request a refund for your order. Even if the agent is interrupted during refund processing, upon restart it automatically recovers, finishes processing the refund, then proceeds to the next step in its workflow. Try running this agent and pressing the `Crash System` button at any time. You can see that when it restarts, it resumes its pending refund processing. All source code is [available on GitHub](https://github.com/dbos-inc/dbos-demo-apps/tree/main/python/reliable-refunds-langchain). ### Overview This customer service AI agent allows users to chat and check the status of their purchase order or request a refund. If the order exceeds a certain cost threshold, the refund request will be automatically escalated to a customer service admin via email for manual review. Based on the admin's decision (approval or rejection), the agent will either process the refund or decline the request accordingly. Let's zoom in to the refund process. The refund process is asynchronous, meaning the user can continue chatting (or leaving and coming back in a few days) with the agent for other tasks while a background process handles the refund workflow. This ensures that the chatbot remains responsive and is not blocked by the potentially long manual review process, which could take hours or even days. The architecture diagram of the refund processing workflow: ![refund workflow](assets/langgraph-agent-workflow.png) There are two main challenges in implementing this process within an AI agent: 1. **Asynchronous Processing**: The approval process may take days, so the workflow must be invoked asynchronously in the background. This ensures that the chatbot can respond to user input quickly and continue handling other interactions without being blocked. 2. **Workflow Reliability**: The workflow must be durable and fault tolerant. If the agent is interrupted during refund processing (e.g., server crashes, network connectivity issues), it should automatically recover upon restart, complete the refund, and seamlessly proceed to the next step. Traditional solutions typically require setting up a **job queue** and separate **queue consumers** to process tasks asynchronously, along with an **external orchestrator** like AWS Step Functions to coordinate multiple subprocesses, guaranteeing the workflow runs to completion. DBOS provides a simpler solution - [durable execution through a database-backed library](https://www.dbos.dev/blog/what-is-lightweight-durable-execution), so you can control durable execution more simply and entirely within your application code. There’s no need to run and stitch together external orchestration services. In the following sections of this tutorial, we'll walk you through how we built a reliable customer service agent using DBOS + LangGraph. ### Writing an AI-Powered Refund Agent Now let's build this agent step-by-step! #### Import and Initialize the App Let's start off with imports and initializing DBOS. We'll also set up FastAPI to serve HTTP requests. ```python showLineNumbers from dbos import DBOS, DBOSConfig, SQLAlchemyDatasource from fastapi import FastAPI database_url = os.environ.get("DBOS_DATABASE_URL") if database_url is None: raise Exception("DBOS_DATABASE_URL not set") app = FastAPI() config: DBOSConfig = { "name": "reliable-refunds-langchain", "system_database_url": database_url, "application_version": "0.1.0", "conductor_key": os.environ.get("CONDUCTOR_KEY"), } DBOS(config=config) ds = SQLAlchemyDatasource.create(database_url) APPROVAL_TIMEOUT_SEC = 60 * 60 * 24 * 7 # One week timeout for manual review sg_api_key = os.environ.get("SENDGRID_API_KEY") assert sg_api_key, "Error: SENDGRID_API_KEY is not set" from_email = os.environ.get("SENDGRID_FROM_EMAIL") assert from_email, "Error: SENDGRID_FROM_EMAIL is not set" admin_email = os.environ.get("ADMIN_EMAIL", None) assert admin_email, "Error: ADMIN_EMAIL is not set" callback_domain = os.environ.get("DBOS_APP_HOSTNAME", "http://localhost:8000") ``` #### Defining Tools for the Agent One great feature of DBOS is that it provides durable execution as a library, allowing seamless integration with popular AI frameworks like LangGraph. To use the DBOS decorated functions as tools for this agent, you simply decorate the function with `@tool` and provide a docstring so that the LLM can correctly identify when to invoke it. This agent has two tools: 1. `tool_get_purchase_by_id`: invokes a database transaction function to retrieve order status. 2. `process_refund`: a workflow process the refund request. ```python showLineNumbers from langchain_core.tools import tool # This tool lets the agent look up the details of an order given its ID. @ds.transaction() def get_purchase_by_id(order_id: int) -> Optional[Purchase]: DBOS.logger.info(f"Looking up purchase by order_id {order_id}") query = purchases.select().where(purchases.c.order_id == order_id) result = ds.sql_session().execute(query) row = result.first() return Purchase.from_row(row) if row is not None else None # Define a wrapper function to make the output JSON serializable. @tool def tool_get_purchase_by_id(order_id: int) -> str: """Look up a purchase by its order id.""" return asdict(get_purchase_by_id(order_id)) # This tool processes a refund for an order. If the order exceeds a cost threshold, # it escalates to manual review. @tool @DBOS.workflow() def process_refund(order_id: int): """Process a refund for an order given an order ID.""" purchase = get_purchase_by_id(order_id) if purchase is None: DBOS.logger.error(f"Refunding invalid order {order_id}") return "We're unable to process your refund. Please check your input and try again." DBOS.logger.info(f"Processing refund for purchase {purchase}") if purchase.price > 1000: update_purchase_status(purchase.order_id, OrderStatus.PENDING_REFUND.value) DBOS.start_workflow(approval_workflow, purchase) return f"Because order_id {purchase.order_id} exceeds our cost threshold, your refund request must undergo manual review. Please check your order status later." else: update_purchase_status(purchase.order_id, OrderStatus.REFUNDED) return f"Your refund for order_id {purchase.order_id} has been approved." ``` We decorate the `tool_get_purchase_by_id` and `process_refund` functions with LangChain's `@tool` decorator, so that they can be recognized and used by LLMs. We decorate `process_refund` as a DBOS workflow. This way, if the agent's workflow is interrupted while processing a refund, when it restarts, it will resume from the last completed step. DBOS guarantees that once the agent's workflow starts, you will always get a refund, and never be refunded twice! #### Asynchronous Human-in-the-Loop Workflow If an order exceeds a certain cost threshold, the refund request will be escalated for manual review. In this case, the `process_refund` workflow starts a child workflow called `approval_workflow` which contains the following step: - An email is sent to an admin for manual review. - The approval workflow **pauses** until a human decision is made. - When the admin clicks approve or reject, it sends an HTTP request to the `/approval/{workflow_id}/{status}` endpoint, which then notifies the pending workflow about the decision. - Based on the response, the workflow either proceeds with the refund or rejects the request. ```python showLineNumbers # This workflow manages manual review. It sends an email to a reviewer, then waits up to a week # for the reviewer to approve or deny the refund request. @DBOS.workflow() def approval_workflow(purchase: Purchase): send_email(purchase) status = DBOS.recv(timeout_seconds=APPROVAL_TIMEOUT_SEC) if status == "approve": DBOS.logger.info("Refund approved :)") update_purchase_status(purchase.order_id, OrderStatus.REFUNDED) return "Approved" else: DBOS.logger.info("Refund rejected :/") update_purchase_status(purchase.order_id, OrderStatus.REFUND_REJECTED) return "Rejected" @app.get("/approval/{workflow_id}/{status}") def approval_endpoint(workflow_id: str, status: str): DBOS.send(workflow_id, status) msg = Template(Path(os.path.join(html_dir, "confirm.html")).read_text()).substitute( result=status, wfid=workflow_id, ) return HTMLResponse(msg) # This function sends an email to a manual reviewer. The email contains links that send notifications # to the approval workflow to approve or deny a refund. @DBOS.step() def send_email(purchase: Purchase): content = f"{callback_domain}/approval/{DBOS.workflow_id}" msg = Template(Path(os.path.join(html_dir, "email.html")).read_text()).substitute( purchaseid=purchase.order_id, purchaseitem=purchase.item, orderdate=purchase.order_date, price=purchase.price, content=content, datetime=time.strftime("%Y-%m-%d %H:%M:%S %Z"), ) message = Mail( from_email=from_email, to_emails=admin_email, subject="Refund Validation", html_content=msg, ) email_client = SendGridAPIClient(sg_api_key) email_client.send(message) DBOS.logger.info(f"Message sent from {from_email} to {admin_email}") # This function updates the status of a purchase. @ds.transaction() def update_purchase_status(order_id: int, status: OrderStatus): query = ( purchases.update() .where(purchases.c.order_id == order_id) .values(order_status=status) ) ds.sql_session().execute(query) ``` The `process_refund` tool uses `DBOS.start_workflow` to execute the approval workflow asynchronously and returns back to the chatbot as soon as the workflow is started, so the chatbot is not blocked by the potentially long review period. #### Setting Up LangGraph Once we define all the tools we need, let's set up LangGraph. We'll use OpenAI's `gpt-5.4-mini` model to answer each chat message. We'll configure LangGraph to store message history in Postgres so it persists across app restarts. ```python showLineNumbers def create_agent(): llm = ChatOpenAI(model="gpt-5.4-mini") tools = [tool_get_purchase_by_id, process_refund] llm_with_tools = llm.bind_tools(tools, parallel_tool_calls=False) prompt = ChatPromptTemplate.from_messages( [ ( "system", "You are a helpful refund agent. You always speak in fluent, natural, conversational language. You can look up order status and process refunds.", ), MessagesPlaceholder(variable_name="messages"), ] ) # This is our refund agent. It follows these instructions to process refunds. # It uses two tools: one to look up order status, one to actually process refunds. agent = prompt | llm_with_tools # Create a state machine using the graph builder graph_builder = StateGraph(State) def chatbot(state: State): return {"messages": [agent.invoke(state["messages"])]} graph_builder.add_node("chatbot", chatbot) tool_node = ToolNode(tools=tools) graph_builder.add_node("tools", tool_node) graph_builder.add_conditional_edges( "chatbot", tools_condition, ) # Any time a tool is called, we return to the chatbot to decide the next step graph_builder.add_edge("tools", "chatbot") graph_builder.add_edge(START, "chatbot") # Create a checkpointer LangChain can use to store message history in Postgres. connection_string = ( make_url(database_url) .set(drivername="postgres") .render_as_string(hide_password=False) ) pool = ConnectionPool(connection_string) checkpointer = PostgresSaver(pool) graph = graph_builder.compile(checkpointer=checkpointer) return graph class ChatSchema(BaseModel): message: str # Currently supports only one chat thread chat_config = {"configurable": {"thread_id": "1"}} compiled_agent = create_agent() ``` Each time a user inputs a message, the agent traverses the DAG until it reaches the "end" node, then responds to the user. The agent diagram (generated by LangGraph) looks simple because DBOS handles all the complex workflows in the "tools" node. ![LangGraph diagram](assets/langgraph-agent-architect.png) #### Handling Chats Now, let's chat! We define two endpoints. The first endpoint handles each new incoming user input. It invokes the agent DAG with the user's message, and filter the response messages for the frontend. ```python showLineNumbers @app.post("/chat") def chat_workflow(chat: ChatSchema): # Invoke the agent DAG with the user's message events = compiled_agent.stream( {"messages": [HumanMessage(chat.message)]}, config=chat_config, stream_mode="values", ) # Filter the response messages for the frontend response_messages = [] for event in events: if "messages" in event: latest_msg = event["messages"][-1] if isinstance(latest_msg, AIMessage) and latest_msg.content: response_messages.append( {"isUser": False, "content": latest_msg.content} ) return response_messages ``` Next, let's add a history endpoint that retrieves the current chat thread from the database. This function is called when we start/refresh the chatbot page so it can display your chat history. We read the current thread's messages from the agent's checkpointed state and parse them for the frontend. ```python showLineNumbers @app.get("/history") def history_endpoint(): # Retrieve the messages from the chat history and parse them for the frontend state = compiled_agent.get_state(chat_config) messages = state.values.get("messages", []) if state else [] message_list = [] for message in messages: if isinstance(message, HumanMessage) and message.content: message_list.append({"isUser": True, "content": message.content}) elif isinstance(message, AIMessage) and message.content: message_list.append({"isUser": False, "content": message.content}) return message_list ``` Additionally, let's serve the app's frontend from an HTML file using FastAPI. In production, we recommend using DBOS primarily for the backend, with your frontend deployed elsewhere. ```python showLineNumbers @app.get("/") def frontend(): with open(os.path.join("html", "app.html")) as file: html = file.read() return HTMLResponse(html) ``` Finally, launch DBOS and start the FastAPI server with Uvicorn: ```python showLineNumbers if __name__ == "__main__": DBOS.launch() uvicorn.run(app, host="0.0.0.0", port=8000) ``` ### Try it Yourself! Clone and enter the [dbos-demo-apps](https://github.com/dbos-inc/dbos-demo-apps) repository: ```shell git clone https://github.com/dbos-inc/dbos-demo-apps.git cd dbos-demo-apps/python/reliable-refunds-langchain ``` Then follow the instructions in the [README](https://github.com/dbos-inc/dbos-demo-apps/tree/main/python/reliable-refunds-langchain) to run the app. --- ## CI/CD Slackbot In this example, we use DBOS and [Bolt](https://slack.dev/bolt-python) to build a CI/CD Slackbot that helps you trigger application deployments and track deployment pipeline progress directly from Slack. It listens to commands such as `/deploy` and `/check_status` in a Slack channel and posts status updates back to the channel. This app demonstrates how DBOS enables: * **Asynchronous background workflows** for each deployment. The Slackbot can't synchronously run an entire deployment pipeline (Slackbots have a [3 second](https://api.slack.com/apis/events-api#retries) time limit), so it instead durably enqueues the pipeline to run in the background. * **Concurrency control using queues**. In this example, the deployment pipeline utilizes a queue with a concurrency limit of 1, so only one deployment can run at a time. * **Real-time status updates** via DBOS events, allowing you to monitor the progress of your deployment pipelines. * **Durable execution** making sure an application deployment resumes from where it left off even if the Slackbot crashes or restarts. All source code is [available on GitHub](https://github.com/dbos-inc/dbos-demo-apps/tree/main/python/deploy-tracker-slackbot). ![Deploy Tracker Slackbot](./assets/deploy-bot.png) ### Main Deployment Workflow The core of this Slackbot is the main deployment workflow. It contains three steps: build, test, and deploy. When the workflow starts, it notifies the Slack channel and updates the status to "started". Then, after each step finishes, it sends a Slack message to notify users about its progress. It also uses `DBOS.set_event` to set a custom status object that users can query. After all steps are complete, it sends a final Slack message to the channel. ```python @DBOS.workflow() def deploy_tracker_workflow(user_id: str, channel_id: str): post_slack_message( message=f"Deployment workflow {DBOS.workflow_id} started.", channel=channel_id ) # Use DBOS's built-in event tracking to mark deployment status DBOS.set_event("deploy_status", "started") # Additional steps for deployment can be added here duration = build_step() post_slack_message( message=f"[Workflow {DBOS.workflow_id}] Step *Build 🚧* completed in {duration:.2f} seconds.", channel=channel_id, ) DBOS.set_event("deploy_status", "build_completed") duration = test_step() post_slack_message( message=f"[Workflow {DBOS.workflow_id}] Step *Test 🧪* completed in {duration:.2f} seconds.", channel=channel_id, ) DBOS.set_event("deploy_status", "tests_completed") duration = deploy_step() post_slack_message( message=f"[Workflow {DBOS.workflow_id}] Step *Deploy 🚀* completed in {duration:.2f} seconds.", channel=channel_id, ) DBOS.set_event("deploy_status", "deployed") post_slack_message( message=f"Deployment workflow {DBOS.workflow_id} completed successfully. <@{user_id}> :tada:", channel=channel_id, ) ``` ### Concurrency-Limited Queue To ensure only one deployment workflow runs at a time, we register a queue with `global_concurrency=1`. This is helpful when the deployment pipeline consumes significant resources and you don't want to exhaust your resource pool when multiple users trigger deployments. ```python # Create a queue with concurrency of 1 so only one deployment workflow runs at a time DBOS.register_queue("deploy-tracker-queue", global_concurrency=1) ``` Because `register_queue` writes the configuration to the system database, it must be called after `DBOS.launch()`. ### Handle Slash Commands in Slack We need to define handlers to handle each Slash Command from Slack. Here, we implement two commands: `/deploy` and `/check_status`. #### Handle `/deploy` This command enqueues a deployment workflow and acknowledges the user as soon as the workflow is confirmed to be enqueued (due to the 3-second limit). It also lists workflows in the queue and provides information on how many workflows are currently queued. ```python @app.command("/deploy") def handle_deploy_command(ack, say, command, logger): try: # Acknowledge the command within 3 seconds, after confirming the workflow has enqueued user_id = command["user_id"] channel_id = command["channel_id"] handle = DBOS.enqueue_workflow( "deploy-tracker-queue", deploy_tracker_workflow, user_id, channel_id ) ack({"response_type": "in_channel"}) # Check the queue size, and inform the user if there are pending workflows queued_workflows = DBOS.list_queued_workflows( queue_name="deploy-tracker-queue", load_input=False ) queue_size = len(queued_workflows) say( f"Deployment workflow has been enqueued by <@{user_id}>, workflow ID {handle.get_workflow_id()}. Currently {queue_size} workflow(s) in the queue." ) # Wait for workflow to complete handle.get_result() except Exception as e: logger.error(f"Error handling slash command: {e}") say(f"An error occurred while processing your deployment request: {e}") ``` #### Handle `/check_status` This endpoint queries the workflow status as well as the custom "deploy_status" event set by the deployment workflow. Users can query the status of their deployment pipelines anytime. ```python @app.command("/check_status") def handle_check_status_command(ack, body, say, logger): try: # Acknowledge the command within 3 seconds ack({"response_type": "in_channel"}) workflow_id = body["text"].strip() # Check the workflow status and latest deployment status handle = DBOS.retrieve_workflow(workflow_id, existing_workflow=False) status = handle.get_status().status deploy_status = DBOS.get_event(workflow_id, "deploy_status", 10) say( f"Deployment workflow {workflow_id} is currently *{status}*. Latest deployment status: *{deploy_status}*." ) except Exception as e: logger.error(f"Error handling slash command: {e}") say(f"An error occurred while processing your check status request: {e}") ``` ### Try it Yourself! Clone and enter the [dbos-demo-apps](https://github.com/dbos-inc/dbos-demo-apps) repository: ```shell git clone https://github.com/dbos-inc/dbos-demo-apps.git cd dbos-demo-apps/python/deploy-tracker-slackbot ``` Then follow the instructions in the [README](https://github.com/dbos-inc/dbos-demo-apps/tree/main/python/deploy-tracker-slackbot) to run the app. --- ## Document Ingestion Pipeline In this example, we'll use DBOS to build a **reliable and scalable data processing pipeline**. We'll show how DBOS can help you horizontally scale an application to process many items concurrently and seamlessly recover from failures. Specifically, we'll build a pipeline that indexes PDF documents for RAG, though you can use a similar design pattern to build almost any data pipeline. To show the pipeline works, we'll also build a simple chat agent that can accurately answer questions about the indexed documents. For example, after ingesting the last few years of Apple 10-K filings, the chat agent can accurately answer questions about Apple's financials: ![Document Detective UI](./assets/document_detective.png) All source code is [available on GitHub](https://github.com/dbos-inc/dbos-demo-apps/tree/main/python/document-detective). ### Import and Initialize the App Let's start with imports and initializing DBOS. ```python import os from tempfile import TemporaryDirectory from typing import List import requests import uvicorn from dbos import DBOS, DBOSConfig, WorkflowHandle from fastapi import FastAPI from fastapi.responses import HTMLResponse from llama_index.core import Document, Settings, StorageContext, VectorStoreIndex from llama_index.readers.file import PDFReader from llama_index.vector_stores.postgres import PGVectorStore from pydantic import BaseModel, HttpUrl from sqlalchemy import make_url database_url = os.environ.get("DBOS_SYSTEM_DATABASE_URL") if not database_url: raise Exception("DBOS_SYSTEM_DATABASE_URL not set") app = FastAPI() config: DBOSConfig = { "name": "document-detective", "application_version": "0.1.0", "system_database_url": database_url, "conductor_key": os.environ.get("CONDUCTOR_KEY"), } DBOS(config=config) ``` Next, let's initialize LlamaIndex to store and query the vector index we'll be constructing. ```python def configure_index(): Settings.chunk_size = 512 db = make_url(database_url) vector_store = PGVectorStore.from_params( database=db.database, host=db.host, password=db.password, port=db.port, user=db.username, perform_setup=False, # Set up during migration step ) storage_context = StorageContext.from_defaults(vector_store=vector_store) index = VectorStoreIndex([], storage_context=storage_context) chat_engine = index.as_chat_engine() return index, chat_engine index, chat_engine = configure_index() ``` ### Building a Durable Data Ingestion Pipeline Now, let's write the document ingestion pipeline. Because ingesting and indexing documents may take a long time, we need to build a pipeline that's both _concurrent_ and _reliable_. It needs to process multiple documents at once and it needs to be resilient to failures, so if the application is interrupted or restarted, or encounters an error, it can recover from where it left off instead of restarting from the beginning or losing some documents entirely. We'll build a concurrent, reliable data ingestion pipeline using DBOS workflows and queues. This workflow takes in a batch of document URLs, enqueues them all for indexing, and waits for all documents to finish being indexed. If it's ever interrupted or restarted, it recovers the indexing of each document from the last completed step, guaranteeing that every document is indexed and none are lost. ```python @DBOS.workflow() def index_documents(urls: List[HttpUrl]): handles: List[WorkflowHandle] = [] # Enqueue each document for indexing for url in urls: handle = DBOS.enqueue_workflow("indexing_queue", index_document, url) handles.append(handle) # Wait for all documents to finish indexing, count the total number of indexed pages indexed_pages = 0 for handle in handles: indexed_pages += handle.get_result() print(f"Indexed {len(urls)} documents totaling {indexed_pages} pages") ``` This workflow indexes an individual document. It calls two steps: `download_document` to download the document and parse it into pages, then `index_page` to add a parsed page to the vector index. ```python @DBOS.workflow() def index_document(document_url: HttpUrl) -> int: pages = download_document(document_url) for page in pages: index_page(page) return len(pages) ``` Here's the code for the steps in `index_document`: ```python @DBOS.step() def download_document(document_url: HttpUrl): # Download the document to a temporary file print(f"Downloading document from {document_url}") with TemporaryDirectory() as temp_dir: temp_file_path = os.path.join(temp_dir, "document.pdf") with open(temp_file_path, "wb") as temp_file: with requests.get(document_url, stream=True) as r: r.raise_for_status() for page in r.iter_content(chunk_size=8192): temp_file.write(page) # Parse the document into pages reader = PDFReader() pages = reader.load_data(temp_file_path) return pages @DBOS.step() def index_page(page: Document): # Insert a page into the vector index try: index.insert(page) except Exception as e: print("Error indexing page:", page, e) ``` Next, let's write the endpoint for indexing. It starts the indexing workflow in the background on a batch of documents. ```python class URLList(BaseModel): urls: List[HttpUrl] @app.post("/index") def index_endpoint(urls: URLList): DBOS.start_workflow(index_documents, urls.urls) ``` ### Chatting With Your Data Now, let's write a simple in-memory chat agent so we can query our data. Every time it gets a question, it answers using a RAG chat engine powered by the vector index. ```python class ChatSchema(BaseModel): message: str chat_history = [] @app.post("/chat") def chat(chat: ChatSchema): query = {"content": chat.message, "isUser": False} chat_history.append(query) responseMessage = str(chat_engine.chat(chat.message)) response = {"content": responseMessage, "isUser": True} chat_history.append(response) return response @app.get("/history") def history_endpoint(): return chat_history ``` We'll serve the app's frontend from an HTML file using FastAPI. ```python @app.get("/") def frontend(): with open(os.path.join("html", "app.html")) as file: html = file.read() return HTMLResponse(html) ``` Finally, let's write a main function to launch DBOS and start our app: ```python if __name__ == "__main__": DBOS.launch() DBOS.register_queue("indexing_queue") uvicorn.run(app, host="0.0.0.0", port=8000) ``` ### Try it Yourself! Clone and enter the [dbos-demo-apps](https://github.com/dbos-inc/dbos-demo-apps) repository: ```shell git clone https://github.com/dbos-inc/dbos-demo-apps.git cd dbos-demo-apps/python/document-detective ``` Then follow the instructions in the [README](https://github.com/dbos-inc/dbos-demo-apps/tree/main/python/document-detective) to run the app. --- ## Hacker News Research Agent :::info This example is also available in [TypeScript](../../typescript/examples/hacker-news-agent). ::: In this example, we use DBOS to build an AI deep research agent that autonomously searches Hacker News for information on any topic. This example demonstrates how to build **reliable, durable AI agents** with durable workflows. The agent starts with a research topic, iteratively researches related topics, then synthesizes findings into a comprehensive report. Because the agent is implemented as a durable workflow, it can recover from any failure and continue research from where it left off, ensuring no work is lost. This example also demonstrates how easy it is to add DBOS to an existing agentic application. Adding DBOS to this agent required changing **<20 lines of code**. All you have to do is annotate workflows and steps. All source code is [available on GitHub](https://github.com/dbos-inc/dbos-demo-apps/tree/main/python/hacker-news-agent). ### Main Research Workflow The core of the agent is the main research workflow. It starts with a topic and autonomously explores related queries until it has enough information, then synthesizes a final report. ```python @DBOS.workflow() def agentic_research_workflow(topic: str, max_iterations: int = 3): """ This agent starts with a research topic then: 1. Searches Hacker News for information on that topic. 2. Iteratively searches related topics, collecting information. 3. Makes decisions about when to continue. 4. Synthesizes findings into a final report. """ console.print(f"[dim]🎯 Starting agentic research for: {topic}[/dim]") # Set and update an agent status the frontend can display agent_status = AgentStatus( created_at=datetime.now().isoformat(), topic=topic, iterations=0, report=None, status="PENDING", ) DBOS.set_event(AGENT_STATUS, agent_status) all_findings = [] current_iteration = 0 current_topic = topic # Main agentic research loop while current_iteration < max_iterations: current_iteration += 1 agent_status.iterations = current_iteration DBOS.set_event(AGENT_STATUS, agent_status) # Research the next topic in a child workflow evaluation = research_topic(topic, current_topic) all_findings.append(evaluation.model_dump()) # Evaluate whether to continue research should_continue = should_continue_step( topic, all_findings, current_iteration, max_iterations ) if not should_continue: break # Generate next research question based on findings if current_iteration < max_iterations: follow_up_topic = generate_follow_ups_step( topic, all_findings, current_iteration ) if follow_up_topic: current_topic = follow_up_topic else: break # Final step: Synthesize all findings into comprehensive report final_report = synthesize_findings_step(topic, all_findings) agent_status.report = final_report.report DBOS.set_event(AGENT_STATUS, agent_status) ``` ### Research Query Workflow Each iteration of the main research workflow calls a child workflow that searches Hacker News for information about a query, then evaluates and returns its findings. ```python @DBOS.workflow() def research_topic(topic: str, query: str) -> EvaluationResult: """Research a topic selected by the main agentic workflow.""" # Step 1: Search Hacker News for stories about the topic stories = search_hackernews_step(query, max_results=30) # Step 2: Gather comments from all stories found comments = [] if stories: for i, story in enumerate(stories): story_id = story.get("objectID") title = story.get("title", "Unknown")[:50] num_comments = story.get("num_comments", 0) if story_id and num_comments > 0: story_comments = get_comments_step(story_id, max_comments=10) comments.extend(story_comments) # Step 3: Evaluate gathered data and return findings return evaluate_results_step(topic, query, stories, comments) ``` ### Agent Decision-Making Steps The agent's intelligence comes from three key step functions that handle decision-making:
Agent Evaluation Step ```python @DBOS.step() def evaluate_results_step( topic: str, query: str, stories: List[Dict[str, Any]], comments: Optional[List[Dict[str, Any]]] = None, ) -> EvaluationResult: """Agent evaluates search results and extracts insights.""" # Create detailed content digest for LLM stories_text = "" top_stories = [] # Evaluate only the top 10 most relevant (per HN search) stories for i, story in enumerate(stories[:10]): title = story.get("title", "No title") url = story.get("url", "No URL") hn_url = f"https://news.ycombinator.com/item?id={story.get('objectID', '')}" points = story.get("points", 0) num_comments = story.get("num_comments", 0) author = story.get("author", "Unknown") stories_text += f"Story {i+1}:\n" stories_text += f" Title: {title}\n" stories_text += f" Points: {points}, Comments: {num_comments}\n" stories_text += f" URL: {url}\n" stories_text += f" HN Discussion: {hn_url}\n" stories_text += f" Author: {author}\n\n" # Store top stories for reference top_stories.append( StoryReference( title=title, url=url, hn_url=hn_url, points=points, num_comments=num_comments, author=author, objectID=story.get("objectID", ""), ) ) comments_text = "" if comments: for i, comment in enumerate(comments[:20]): # Limit to top 20 comments comment_text = comment.get("comment_text", "") if comment_text: author = comment.get("author", "Unknown") # Get longer excerpts for better analysis excerpt = ( comment_text[:400] + "..." if len(comment_text) > 400 else comment_text ) comments_text += f"Comment {i+1}:\n" comments_text += f" Author: {author}\n" comments_text += f" Text: {excerpt}\n\n" prompt = f""" You are a research agent evaluating search results for: {topic} Query used: {query} Stories found: {stories_text} Comments analyzed: {comments_text} Provide a DETAILED analysis with specific insights, not generalizations. Focus on: - Specific technical details, metrics, or benchmarks mentioned - Concrete tools, libraries, frameworks, or techniques discussed - Interesting problems, solutions, or approaches described - Performance data, comparison results, or quantitative insights - Notable opinions, debates, or community perspectives - Specific use cases, implementation details, or real-world examples Return JSON with: - "insights": String array of specific, technical insights with context - "relevance_score": Number 1-10 - "summary": Brief summary of findings - "key_points": Array of most important points discovered """ messages = [ { "role": "system", "content": "You are a research evaluation agent. Analyze search results and provide structured insights in JSON format.", }, {"role": "user", "content": prompt}, ] response = call_llm(messages, max_tokens=2000) cleaned_response = clean_json_response(response) evaluation_dict = json.loads(cleaned_response) evaluation_dict["query"] = query evaluation_dict["top_stories"] = top_stories return EvaluationResult(**evaluation_dict) ```
Follow-up Query Generation Step ```python @DBOS.step() def generate_follow_ups_step( topic: str, current_findings: List[Dict[str, Any]], iteration: int ) -> Optional[str]: """Agent generates follow-up research queries based on current findings.""" findings_summary = "" for finding in current_findings: findings_summary += f"Query: {finding.get('query', 'Unknown')}\n" findings_summary += f"Summary: {finding.get('summary', 'No summary')}\n" findings_summary += f"Key insights: {finding.get('insights', [])}\n" findings_summary += ( f"Unanswered questions: {finding.get('unanswered_questions', [])}\n\n" ) prompt = f""" You are a research agent investigating: {topic} This is iteration {iteration} of your research. Current findings: {findings_summary} Generate 2-4 SHORT KEYWORD-BASED search queries for Hacker News that explore DIVERSE aspects of {topic}. CRITICAL RULES: 1. Use SHORT keywords (2-4 words max) - NOT long sentences 2. Focus on DIFFERENT aspects of {topic}, not just one narrow area 3. Use terms that appear in actual Hacker News story titles 4. Avoid repeating previous focus areas 5. Think about what tech people actually discuss about {topic} For {topic}, consider diverse areas like: - Performance/optimization - Tools/extensions - Comparisons with other technologies - Use cases/applications - Configuration/deployment - Recent developments GOOD examples: ["postgres performance", "database tools", "sql optimization"] BAD examples: ["What are the best practices for PostgreSQL optimization?"] Return only a JSON array of SHORT keyword queries: ["query1", "query2", "query3"] """ messages = [ { "role": "system", "content": "You are a research agent. Generate focused follow-up queries based on current findings. Return only JSON array.", }, {"role": "user", "content": prompt}, ] response = call_llm(messages) cleaned_response = clean_json_response(response) queries = json.loads(cleaned_response) return queries[0] if isinstance(queries, list) and len(queries) > 0 else None ```
Continuation Decision Step ```python @DBOS.step() def should_continue_step( topic: str, all_findings: List[Dict[str, Any]], current_iteration: int, max_iterations: int, ) -> bool: """Agent decides whether to continue research or conclude.""" if current_iteration >= max_iterations: return { "should_continue": False, "reason": f"Reached maximum iterations ({max_iterations})", } # Analyze findings completeness findings_summary = "" total_relevance = 0 for finding in all_findings: findings_summary += f"Query: {finding.get('query', 'Unknown')}\n" findings_summary += f"Summary: {finding.get('summary', 'No summary')}\n" findings_summary += f"Relevance: {finding.get('relevance_score', 5)}/10\n" total_relevance += finding.get("relevance_score", 5) avg_relevance = total_relevance / len(all_findings) if all_findings else 0 prompt = f""" You are a research agent investigating: {topic} Current iteration: {current_iteration}/{max_iterations} Findings so far: {findings_summary} Average relevance score: {avg_relevance:.1f}/10 Decide whether to continue research or conclude. PRIORITIZE THOROUGH EXPLORATION - continue if: 1. Current iteration is less than 75% of max_iterations 2. Average relevance is above 6.0 and there are likely unexplored aspects 3. Recent queries found significant new information 4. The research seems to be discovering diverse perspectives on the topic Only stop early if: - Average relevance is below 5.0 for multiple iterations - No new meaningful information in the last 2 iterations - Research appears to be hitting diminishing returns Return JSON with: - "should_continue": boolean """ messages = [ { "role": "system", "content": "You are a research decision agent. Evaluate research completeness and decide whether to continue. Return JSON.", }, {"role": "user", "content": prompt}, ] raw_response = call_llm(messages) cleaned_response = clean_json_response(raw_response) json_response = json.loads(cleaned_response) response = ShouldContinueResult(**json_response) return response.should_continue ```
### Search API Steps After deciding what terms to search for, the agent calls these steps to retrieve stories and comments from Hacker News.
Hacker News API Steps ```python @DBOS.step() def search_hackernews_step(query: str, max_results: int = 20) -> List[Dict[str, Any]]: """Search Hacker News stories using Algolia API.""" params = {"query": query, "hitsPerPage": max_results, "tags": "story"} with httpx.Client(timeout=30.0) as client: response = client.get("https://hn.algolia.com/api/v1/search", params=params) response.raise_for_status() return response.json()["hits"] @DBOS.step() def get_comments_step(story_id: str, max_comments: int = 50) -> List[Dict[str, Any]]: """Get comments for a specific Hacker News story.""" params = {"tags": f"comment,story_{story_id}", "hitsPerPage": max_comments} with httpx.Client(timeout=30.0) as client: response = client.get("https://hn.algolia.com/api/v1/search", params=params) response.raise_for_status() return response.json()["hits"] ```
### Synthesize Findings Step Finally, after concluding its research, the agentic workflow calls this step to synthesize its findings into a report.
Synthesize Findings Step ```python @DBOS.step() def synthesize_findings_step( topic: str, all_findings: List[Dict[str, Any]] ) -> ResearchReport: """Synthesize all research findings into a comprehensive report.""" findings_text = "" story_links = [] for i, finding in enumerate(all_findings, 1): findings_text += f"\n=== Finding {i} ===\n" findings_text += f"Query: {finding.get('query', 'Unknown')}\n" findings_text += f"Summary: {finding.get('summary', 'No summary')}\n" findings_text += f"Key Points: {finding.get('key_points', [])}\n" findings_text += f"Insights: {finding.get('insights', [])}\n" # Extract story links and details for reference if finding.get("top_stories"): for story in finding["top_stories"]: story_links.append( { "title": story.get("title", "Unknown"), "url": story.get("url", ""), "hn_url": f"https://news.ycombinator.com/item?id={story.get('objectID', '')}", "points": story.get("points", 0), "comments": story.get("num_comments", 0), } ) # Build comprehensive story and citation data story_citations = {} citation_id = 1 for finding in all_findings: if finding.get("top_stories"): for story in finding["top_stories"]: story_id = story.get("objectID", "") if story_id and story_id not in story_citations: story_citations[story_id] = { "id": citation_id, "title": story.get("title", "Unknown"), "url": story.get("url", ""), "hn_url": story.get("hn_url", ""), "points": story.get("points", 0), "comments": story.get("num_comments", 0), } citation_id += 1 # Create citation references text citations_text = "\n".join( [ f"[{cite['id']}] {cite['title']} ({cite['points']} points, {cite['comments']} comments) - {cite['hn_url']}" + (f" - {cite['url']}" if cite["url"] else "") for cite in story_citations.values() ] ) prompt = f""" You are a research analyst. Synthesize the following research findings into a comprehensive, detailed report about: {topic} Research Findings: {findings_text} Available Citations: {citations_text} IMPORTANT: You must return ONLY a valid JSON object with no additional text, explanations, or formatting. Create a comprehensive research report that flows naturally as a single narrative. Include: - Specific technical details and concrete examples - Actionable insights practitioners can use - Interesting discoveries and surprising findings - Specific tools, libraries, or techniques mentioned - Performance metrics, benchmarks, or quantitative data when available - Notable opinions or debates in the community - INLINE LINKS: When making claims, include clickable links directly in the text using this format: [link text](HN_URL) - Use MANY inline links throughout the report. Aim for at least 4-5 links per paragraph. CRITICAL CITATION RULES - FOLLOW EXACTLY: 1. NEVER replace words with bare URLs like "(https://news.ycombinator.com/item?id=123)" 2. ALWAYS write complete sentences with all words present 3. Add citations using descriptive link text in brackets: [descriptive text](URL) 4. Every sentence must be grammatically complete and readable without the links 5. Links should ALWAYS be to the Hacker News discussion, NEVER directly to the article. CORRECT examples: "PostgreSQL's performance improvements have been significant in recent versions, as discussed in [community forums](https://news.ycombinator.com/item?id=123456), with developers highlighting [specific optimizations](https://news.ycombinator.com/item?id=789012) in query processing." "Redis performance issues can stem from common configuration mistakes, which are well-documented in [troubleshooting guides](https://news.ycombinator.com/item?id=345678) and [community discussions](https://news.ycombinator.com/item?id=901234)." "React's licensing changes have sparked significant community debate, as seen in [detailed discussions](https://news.ycombinator.com/item?id=15316175) about the implications for open-source projects." WRONG examples (NEVER DO THIS): "Community discussions reveal a strong interest in the (https://news.ycombinator.com/item?id=18717168) and the common pitfalls" "One significant topic is the (https://news.ycombinator.com/item?id=15316175), which raises important legal considerations" Always link to relevant discussions for: - Every specific tool, library, or technology mentioned - Performance claims and benchmarks - Community opinions and debates - Technical implementation details - Companies or projects referenced - Version releases or updates - Problem reports or solutions Return a JSON object with this exact structure: {{ "report": "A comprehensive research report written as flowing narrative text with inline clickable links [like this](https://news.ycombinator.com/item?id=123). Include specific technical details, tools, performance metrics, community opinions, and actionable insights. Make it detailed and informative, not just a summary." }} """ messages = [ { "role": "system", "content": "You are a research analyst. Provide comprehensive synthesis in JSON format.", }, {"role": "user", "content": prompt}, ] raw_response = call_llm(messages, max_tokens=3000) cleaned_response = clean_json_response(raw_response) json_response = json.loads(cleaned_response) return ResearchReport(**json_response) ```
### API Endpoints The agent has two API endpoints used by its frontend: one that starts a new agent researching a topic and one that retrieves agent statuses. This endpoint starts a durable agent in the background: ```python @app.post("/agents") def start_agent(request: AgentStartRequest): # Start a durable agent in the background DBOS.start_workflow(agentic_research_workflow, request.topic) return {"ok": True} ``` This endpoint returns the statuses of all agents. It lists all agents with [`DBOS.list_workflows`](../reference/contexts.md#list_workflows), then retrieves the status of each using [`DBOS.get_event`](../tutorials/workflow-communication.md#workflow-events). ```python @app.get("/agents", response_model=list[AgentStatus]) async def list_agents(): # List all active agents and retrieve their statuses agent_workflows = await DBOS.list_workflows_async( name=agentic_research_workflow.__qualname__, sort_desc=True, ) statuses: list[AgentStatus] = await asyncio.gather( *[DBOS.get_event_async(w.workflow_id, AGENT_STATUS) for w in agent_workflows] ) for workflow, status in zip(agent_workflows, statuses): status.status = workflow.status status.agent_id = workflow.workflow_id return statuses ``` ### Try it Yourself! Clone and enter the [dbos-demo-apps](https://github.com/dbos-inc/dbos-demo-apps) repository: ```shell git clone https://github.com/dbos-inc/dbos-demo-apps.git cd dbos-demo-apps/python/hacker-news-agent ``` Then follow the instructions in the [README](https://github.com/dbos-inc/dbos-demo-apps/tree/main/python/hacker-news-agent) to run the app. --- ## Transactional Outbox A **transactional outbox** is a common pattern that solves an important problem: how to reliably update a database record and send a message to another system. This is trickier than it sounds because the operations usually need to be **atomic**: they either both happen or neither do, even if there are failures (such as process crashes or network glitches) while performing them. Otherwise, the database might go out of sync with other systems, which could cause serious data integrity issues. A transactional outbox is typically implemented by adding a new "outbox" table to our database. When we need to perform an atomic update, we run a single database transaction that both: - Updates the database record - Writes the message we want to send to the "outbox" table. A separate background process then polls the outbox table and sends the messages there to the other system. Performing the database record update and writing the message to the "outbox" table in one transaction guarantees atomicity: either both records are updated or neither are, and once the message is written to the outbox, it will asynchronously be consumed and sent by the background process even if failures occur later. #### Performing Multiple Operations Atomically With DBOS In DBOS, we can use **durable workflows** instead of an explicit outbox table to atomically perform multiple operations, such as updating a database record and sending a message to another system. To do this, we simply perform each operation as a separate step in a durable workflow. For example: ```python @ds.transaction() def insert_order(customer: str, item: str, quantity: int) -> int: """Insert an order and return its ID. In the classic outbox pattern you would also INSERT an outbox row here. With DBOS the workflow itself provides that guarantee, so no outbox table is needed. """ result = ds.sql_session().execute( orders.insert().values(customer=customer, item=item, quantity=quantity) ) order_id: int = result.inserted_primary_key[0] DBOS.logger.info(f"Inserted order {order_id}: {quantity}x {item} for {customer}") return order_id @DBOS.step() def send_order_notification(order_id: int, customer: str, item: str) -> None: """Simulate sending an order confirmation (e.g. email, Kafka, webhook). In the classic pattern a background poller would read the outbox and call this. With DBOS the workflow calls it directly and guarantees it will be retried until it succeeds. """ DBOS.logger.info( f"Sending notification for order {order_id}: {item} for {customer}" ) time.sleep(3) # simulate network latency DBOS.logger.info(f"Notification sent for order {order_id}: {item} for {customer}") @DBOS.workflow() def place_order_workflow(customer: str, item: str, quantity: int) -> int: """Place an order and send a notification, atomically. If this process crashes after insert_order but before send_order_notification, DBOS will automatically recover and complete the notification on restart. """ order_id = insert_order(customer, item, quantity) send_order_notification(order_id, customer, item) ``` This works because **durable workflows are atomic**. If a failure occurs after writing to the database but before sending the message to the external system, the workflow will recover from its last completed step (writing to the database) and retry the next step (sending the message) until the message is successfully sent. This is the same guarantee a conventional transactional outbox provides: assuming the message is eventually delivered after enough retries, either both operations occur or neither do. One noteworthy detail is that we perform the initial database write in a [transactional step](../tutorials/transaction-tutorial.md), which performs the workflow checkpoint in the same database transaction as the step logic. This way, the database write is guaranteed to execute exactly-once no matter what failures occur during workflow execution. Other operations may execute at-least-once, and so should be idempotent (the same is true in a conventional transactional outbox pattern, where messages are sent from the outbox with at-least-once semantics). #### Transactionally Enqueuing a Workflow The durable workflow above replaces the outbox entirely. If you instead want a pattern that more closely mirrors a conventional outbox, where you write a database record and durably schedule some follow-up work in the same transaction, you can **transactionally enqueue a workflow**. Inside the transaction that inserts the order, we call the [`dbos.enqueue_workflow` Postgres function](../../explanations/system-tables.md) to enqueue a notification workflow. Because the enqueue happens in the same transaction as the order insert, the order row and the enqueued workflow commit (or roll back) together: the notification workflow is durably enqueued if and only if the order is created. This is exactly the guarantee a conventional outbox provides, except the "outbox" is DBOS's own queue table instead of one you build and poll yourself. ```python @ds.transaction() def insert_order(customer: str, item: str, quantity: int) -> int: """Insert an order and transactionally enqueue its notification workflow.""" session = ds.sql_session() result = session.execute( orders.insert().values(customer=customer, item=item, quantity=quantity) ) order_id: int = result.inserted_primary_key[0] # Enqueue the notification workflow as part of this transaction. session.execute( sa.text(""" SELECT dbos.enqueue_workflow( workflow_name => :workflow_name, queue_name => :queue_name, positional_args => ARRAY[ CAST(:arg_order_id AS json), CAST(:arg_customer AS json), CAST(:arg_item AS json) ] ) """), { "workflow_name": "send_notification_workflow", "queue_name": NOTIFICATION_QUEUE, "arg_order_id": json.dumps(order_id), "arg_customer": json.dumps(customer), "arg_item": json.dumps(item), }, ) DBOS.logger.info(f"Inserted order {order_id}: {quantity}x {item} for {customer}") return order_id ``` The enqueued workflow acts as the "consumer" of the outbox. DBOS guarantees it runs exactly once for every committed order, recovering automatically if the process crashes partway through: ```python @DBOS.workflow() def send_notification_workflow(order_id: int, customer: str, item: str) -> None: """Send a notification for an order, then mark it sent.""" send_order_notification(order_id, customer, item) update_notification_status(order_id, "SENT") ``` #### Try it Yourself Full source code for both patterns, demoing how they can recover from any failure, is [available on GitHub](https://github.com/dbos-inc/dbos-demo-apps/tree/main/python/transactional-outbox). The single-workflow pattern is in `atomic_workflow.py` and the transactional enqueue pattern is in `transactional_enqueue.py`. To run it, clone and enter the [dbos-demo-apps](https://github.com/dbos-inc/dbos-demo-apps) repository: ```shell git clone https://github.com/dbos-inc/dbos-demo-apps.git cd dbos-demo-apps/python/transactional-outbox ``` Then follow the instructions in the [README](https://github.com/dbos-inc/dbos-demo-apps/tree/main/python/transactional-outbox) to run the app. --- ## Advanced Queue Patterns This example demonstrates how to build several advanced queue patterns with DBOS. For the full queues documentation, check out the [queues tutorial](../tutorials/queue-tutorial.md). All source code is [available on GitHub](https://github.com/dbos-inc/dbos-demo-apps/tree/main/python/queue-patterns). ### Fair Queueing Often, you have a queue with limited capacity and need to fairly divide that capacity among multiple tenants. Suppose your application can only process 5 workflows per server, on a limited number of servers. You don't want one tenant to monopolize the capacity. With fair queueing, you can limit each tenant to 1 workflow at a time while still allowing up to 5 workflows per server. You can implement fair queueing in DBOS by combining a [**partitioned queue**](../tutorials/queue-tutorial.md#partitioning-queues) with a **regular (non-partitioned) queue**. You enforce per-tenant limits on the partitioned queue and per-server limits on the non-partitioned queue. To do that, first let's register the two queues and define a workflow: ```python DBOS.register_queue("concurrency-queue", worker_concurrency=5) DBOS.register_queue("partitioned-queue", partition_concurrency=1) # This workflow is fairly queued: at most five workflows can run concurrently, # but no more than one per tenant. @DBOS.workflow() def fair_queue_workflow(): time.sleep(5) ``` Because `DBOS.register_queue` writes the queue's configuration to the system database, call it after `DBOS.launch()`. Next, let's create an endpoint to enqueue the workflow. It does not enqueue the workflow directly, but instead enqueues a "concurrency manager" workflow to the partitioned queue to enforce per-tenant limits: ```python @api.post("/workflows/fair_queue") def submit_fair_queue(tenant_id: str): # Enqueue a "concurrency manager" workflow to the partitioned # queue to enforce per-partition limits. with SetEnqueueOptions(queue_partition_key=tenant_id): DBOS.enqueue_workflow("partitioned-queue", fair_queue_concurrency_manager) ``` The "concurrency manager" bridges the two queues, enqueueing the workflow on the non-partitioned queue and waiting for it to complete: ```python @DBOS.workflow() def fair_queue_concurrency_manager(): # The "concurrency manager" workflow enqueues the # workflow on the non-partitioned queue and # awaits its results to enforce global flow control limits. return DBOS.enqueue_workflow("concurrency-queue", fair_queue_workflow).get_result() ``` Because the "concurrency manager" has the same lifetime as the actual workflow, this pattern ensures both the partitioned queue's per-tenant limits and the non-partitioned queue's worker concurrency limits are respected. You can adapt this pattern to combine any per-tenant limits with any global limits. ### Rate Limiting Sometimes, you need to **rate limit** a workflow, limiting the number of workflows that can start in a given period. This is especially useful when using a rate-limited API, like many LLM APIs. You can do this by applying a rate limit to a queue. For example, here's a rate-limited queue and workflow: ```python DBOS.register_queue("rate-limited-queue", limiter={"limit": 2, "period": 10}) # This workflow is rate-limited: No more than two workflows can start per 10 seconds @DBOS.workflow() def rate_limited_queue_workflow(): time.sleep(5) ``` If a rate-limit is defined with limit X and period Y, no more than X workflows can start per Y seconds. Rate limits are global across all DBOS processes using a queue. You can enqueue a workflow on a rate-limited queue like with any other queue: ```python @api.post("/workflows/rate_limited_queue") def submit_rate_limited_queue(): DBOS.enqueue_workflow("rate-limited-queue", rate_limited_queue_workflow) ``` ### Debouncing Sometimes, you may receive many requests to start a workflow in quick succession, but you only want to start it once. For example, if a user is editing an input field, you may want to start a processing workflow only after some time has passed since the last edit. **Debouncing** delays a workflow's execution until some time has passed since it was last called. To debounce a workflow, we define the workflow and queue and create a debouncer for the workflow: ```python DBOS.register_queue("debouncer-queue") @DBOS.workflow() def debouncer_workflow(tenant_id: str, input: str): print(f"Executing debounced workflow for tenant {tenant_id} with input {input}") time.sleep(5) debouncer = Debouncer.create(debouncer_workflow, queue="debouncer-queue") ``` Then, we submit the workflow with the debouncer. This delays the workflow until a set time has passed since the last input is submitted for a tenant. When the workflow starts, it uses the last input received by the debouncer. ```python # Each time a new input is submitted for a tenant, debounce debouncer_workflow. # The debouncer waits until 5 seconds after input stops being submitted for the tenant, # then enqueues the workflow with the last input submitted. @api.post("/workflows/debouncer") def submit_debounced_workflow(tenant_id: str, input: str): debounce_key = tenant_id debounce_period_sec = 5 debouncer.debounce(debounce_key, debounce_period_sec, tenant_id, input) ``` Learn more about debouncing in the [reference](../reference/contexts.md#debouncing). ### Try it Yourself! Clone and enter the [dbos-demo-apps](https://github.com/dbos-inc/dbos-demo-apps) repository: ```shell git clone https://github.com/dbos-inc/dbos-demo-apps.git cd dbos-demo-apps/python/queue-patterns ``` Then follow the instructions in the [README](https://github.com/dbos-inc/dbos-demo-apps/tree/main/python/queue-patterns) to run the example. --- ## Queue Worker :::info This example is also available in [TypeScript](../../typescript/examples/queue-worker.md). ::: This example demonstrates how to run DBOS workflows in their own "queue worker" service while enqueueing and managing them from other services. This design pattern lets you separate concerns and separately scale the workers that execute your durable workflows from your other services. Architecturally, this example contains two services: a web server and a worker service. The web server uses the [DBOS Client](../reference/client.md) to enqueue workflows and monitor their status. The worker service dequeues and executes workflows. All source code is [available on GitHub](https://github.com/dbos-inc/dbos-demo-apps/tree/main/python/queue-worker). ### Worker Service The worker service implements your durable workflows and their steps. Notably, this workflow periodically reports its progress using [`DBOS.set_event`](../tutorials/workflow-communication.md). This lets the web server query the event to monitor workflow progress. ```python @DBOS.workflow() def workflow(num_steps: int): progress = { "steps_completed": 0, "num_steps": num_steps, } # The server can query this event to obtain # the current progress of the workflow DBOS.set_event(WF_PROGRESS_KEY, progress) for i in range(num_steps): step(i) # Update workflow progress each time a step completes progress["steps_completed"] = i + 1 DBOS.set_event(WF_PROGRESS_KEY, progress) @DBOS.step() def step(i: int): print(f"Step {i} completed!") time.sleep(1) ``` In its main function, the worker service configures and launches DBOS, registers the queue on which the web server can submit workflows for execution, then waits indefinitely, dequeuing and executing workflows: ```python if __name__ == "__main__": system_database_url = os.environ.get( "DBOS_SYSTEM_DATABASE_URL", "sqlite:///dbos_queue_worker.sqlite" ) config: DBOSConfig = { "name": "dbos-queue-worker", "application_version": "0.1.0", "system_database_url": system_database_url, } DBOS(config=config) DBOS.launch() # Define a queue on which the web server # can submit workflows for execution. DBOS.register_queue("workflow-queue") # After launching DBOS, the worker waits indefinitely, # dequeuing and executing workflows. threading.Event().wait() ``` ### Web Server The web server first creates a DBOS Client: ```python system_database_url = os.environ.get( "DBOS_SYSTEM_DATABASE_URL", "sqlite:///dbos_queue_worker.sqlite" ) client = DBOSClient(system_database_url=system_database_url) ``` It then enqueues workflows using the client: ```python @api.post("/workflows") def enqueue_workflow(): options: EnqueueOptions = { "queue_name": "workflow-queue", "workflow_name": "workflow", } num_steps = 10 client.enqueue(options, num_steps) return {"status": "enqueued"} ``` The web server can also report workflow status. This function first lists all workflows, then uses [`get_event`](../tutorials/workflow-communication.md) to query the progress of each workflow. This is a useful pattern for showing workflow progress or status to end users of your application. ```python @api.get("/workflows") def list_workflows() -> List[WorkflowStatus]: # Use the DBOS client to list all workflows workflows = client.list_workflows(name="workflow", sort_desc=True) statuses: List[WorkflowStatus] = [] for workflow in workflows: # Query each workflow's progress event. This may not be available # if the workflow has not yet started executing. progress = client.get_event( workflow.workflow_id, WF_PROGRESS_KEY, timeout_seconds=0 ) status = WorkflowStatus( workflow_id=workflow.workflow_id, workflow_status=workflow.status, steps_completed=progress.get("steps_completed") if progress else None, num_steps=progress.get("num_steps") if progress else None, ) statuses.append(status) return statuses ``` ### Try it Yourself! Clone and enter the [dbos-demo-apps](https://github.com/dbos-inc/dbos-demo-apps) repository: ```shell git clone https://github.com/dbos-inc/dbos-demo-apps.git cd dbos-demo-apps/python/queue-worker ``` Then follow the instructions in the [README](https://github.com/dbos-inc/dbos-demo-apps/tree/main/python/queue-worker) to run the app. --- ## S3 Mirror S3Mirror is a DBOS-powered utility for performant, durable and observable data transfers between S3 buckets. This app was created in collaboration with Bristol Myers Squibb for reliable transfers of genomic datasets. Read our joint manuscript, including a performance benchmark, on bioRxiv [here](https://www.biorxiv.org/content/10.1101/2025.06.13.657723v1). Structurally, the app uses a "fanout" pattern. A transfer starts with a list of files. We use a DBOS queue to process the files as independent tasks. We configure the queue to tune the number of simultaneous API calls we submit to S3 - to achieve peak performance without exceeding the limit. On DBOS Cloud Pro, the app auto-scales based on queue length, automatically launching new VMs to speed up large transfers. Each file is processed as a separate Step, which the queue automatically wraps in its own Workflow. This means that, if the app crashes and restarts, the transfer resumes only the files that have not yet completed. Meanwhile, the [DBOS workflow management API](../tutorials/workflow-management.md) makes it effortless to determine the state of each file at any point. So the app offers instant file-wise transfer observability. ### API The app implements the following endpoints: 1. POST `/start_transfer`: given a list of files, starts a new transfer and returns its `transfer_id` 2. POST `/cancel/{transfer_id}`: cancel a perviously-started transfer 3. GET `/transfer_status/{transfer_id}`: returns the file-wise status of a transfer 4. POST `/crash_application`: immediately terminates the app process for durability demonstration The following sections describe how the transfers are implemented. ### Defining a Step to Transfer a File AWS recommends using the `boto3` package to split the file into chunks and transfer using many concurrent requests per file. Where possible, boto3 will actually tranfer the data in the S3 backplane, without having to download and re-upload the chunks. We define a simple [DBOS Step](../tutorials/step-tutorial.md) wrapper around the `s3.copy` routine. Just in case there are transient errors, we decorate this step with `retries_allowed=True, max_attempts=3`. We add some simple logging as well: ```python s3 = boto3.client('s3', config=Config(max_pool_connections=MAX_FILES_PER_WORKER * MAX_CONCURRENT_REQUESTS_PER_FILE)) @DBOS.step(retries_allowed=True, max_attempts=3) def s3_transfer_file(buckets: BucketPaths, task: FileTransferTask): DBOS.span.set_attribute("s3mirror_key", task.key) DBOS.logger.info(f"{DBOS.workflow_id} starting transfer {task.idx}: {task.key}") s3.copy( CopySource= { 'Bucket': buckets.src_bucket, 'Key': buckets.src_prefix + task.key }, Bucket = buckets.dst_bucket, Key = buckets.dst_prefix + task.key, Config = TransferConfig( use_threads=True, max_concurrency=MAX_CONCURRENT_REQUESTS_PER_FILE, multipart_chunksize=FILE_CHUNK_SIZE_BYTES ) ) DBOS.logger.info(f"{DBOS.workflow_id} finished transfer {task.idx}: {task.key}") ``` ### Starting a Transfer We start transferring a batch of files using a [DBOS workflow](../tutorials/workflow-tutorial.md). The workflow enqueues one `s3_transfer_file` step for each file. The queue automatically wraps each step in its own workflow and we capture the list of file-wise Workflow IDs. We then use [DBOS.set_event](../tutorials/workflow-communication.md#set_event) to record those Workflow IDs, along with metadata about the files, for later retrieval. As of this writing, S3 supports up to 3500 concurrent requests per prefix. So we set `concurrency` and `worker_concurrency` on our queue to allow for some parallelism. ```python transfer_queue = Queue("transfer_queue", concurrency = MAX_FILES_AT_A_TIME, worker_concurrency = MAX_FILES_PER_WORKER) @DBOS.workflow() def transfer_job(buckets: BucketPaths, tasks: List[FileTransferTask]): DBOS.logger.info(f"{DBOS.workflow_id} starting {len(tasks)} transfers from {buckets.src_bucket}/{buckets.src_prefix} to {buckets.dst_bucket}/{buckets.dst_prefix}") # For each task, start a workflow on the queue for task in tasks: handle = transfer_queue.enqueue(s3_transfer_file, task = task, buckets = buckets) task.workflow_id = handle.workflow_id # Store the description and ID of each transfer in the workflow context DBOS.set_event('tasks', tasks) ``` This workflow terminates as soon as all of its child workflows are enqueued. Once enqueued, DBOS ensures that they will continue to completion. ### Cancelling a Transfer To stop a previously-started transfer, we use the workflow ID of the `transfer_job` that started it. We call [DBOS.get_event](../tutorials/workflow-communication.md#get_event) to retrieve the transfer data stored previously. This gives us the list of transferred files, some metadata about them, and the workflow ID of each respective `s3_transfer_file` step. We then simply call [DBOS.cancel_workflow](../reference/contexts.md#cancel_workflow) for each of those IDs. ```python @app.post("/cancel/{transfer_id}") def cancel(transfer_id: str): tasks = DBOS.get_event(transfer_id, 'tasks', timeout_seconds=0) if tasks is None: raise HTTPException(status_code=404, detail="Transfer not found") for task in tasks: DBOS.cancel_workflow(task.workflow_id) ``` ### Monitoring an Existing Transfer Just like the cancel method, we start by calling [DBOS.get_event](../tutorials/workflow-communication.md#get_event) to retrieve file-wise metadata including the enqueued workflow IDs. Instead of calling [DBOS.cancel_workflow](../reference/contexts.md#cancel_workflow), we call [DBOS.list_workflows](../reference/contexts.md#list_workflows) for each ID to retrieve its summary and current status. We can count how many files have a status of "SUCCESS" versus "ERROR". DBOS also tracks the start and update time of each workflow. So, with a handful of lines of code, we can calculate the gigabyte-per-second transfer rate. ```python @app.get("/transfer_status/{transfer_id}") def transfer_status(transfer_id: str): tasks = DBOS.get_event(transfer_id, 'tasks', timeout_seconds=0) if tasks is None: raise HTTPException(status_code=404, detail="Transfer not found") filewise_status = [] n_transferred = n_error = transferred_size = 0 t_start = t_end = None for task in tasks: workflow_summary = DBOS.list_workflows(workflow_ids=[task.workflow_id])[0] filewise_status.append({ 'file': task.key, 'size': task.size, 'status': workflow_summary.status, 'tstart': workflow_summary.created_at, 'tend': (workflow_summary.updated_at if workflow_summary.status == "SUCCESS" else None), 'error': str(workflow_summary.error) }) if workflow_summary.status == "SUCCESS": t_start = workflow_summary.created_at if t_start is None else min(t_start, workflow_summary.created_at) t_end = workflow_summary.updated_at if t_end is None else max(t_end, workflow_summary.updated_at) n_transferred += 1 transferred_size += task.size elif workflow_summary.status == "ERROR": n_error += 1 transfer_rate = transferred_size * 1000.0 / (t_end - t_start) if transferred_size > 0 and (t_start != t_end) else 0 return { 'files': len(tasks), 'transferred': n_transferred, 'errors': n_error, 'rate': transfer_rate, 'filewise': filewise_status } ``` ### Try it Yourself! Clone the [dbos-demo-apps](https://github.com/dbos-inc/dbos-demo-apps) repository: ```shell git clone https://github.com/dbos-inc/dbos-demo-apps.git cd python/s3mirror ``` Then follow the instructions in the [README](https://github.com/dbos-inc/dbos-demo-apps/tree/main/python/s3mirror) to run the app. --- ## Fault-Tolerant Checkout(3) :::info This example is also available in [TypeScript](../../typescript/examples/checkout-tutorial), [Java](../../java/examples/widget-store.md), and [Go](../../golang/examples/widget-store.md). ::: In this example, we use DBOS and FastAPI to build an online storefront that's resilient to any failure. You can [run the application yourself](#try-it-yourself) and press its crash button as often as you want. Within a few seconds, the app will recover and resume as if nothing happened. All source code is [available on GitHub](https://github.com/dbos-inc/dbos-demo-apps/tree/main/python/widget-store). ![Widget store UI](./assets/widget_store_ui.png) ### Import and Set Up the App Let's begin with imports and creating a FastAPI app. We also declare a [SQLAlchemy datasource](../tutorials/transaction-tutorial.md), which the app uses to run its database operations. We configure DBOS and the datasource at startup, so here we only declare `ds`. Finally, we define some constants. ```python import os import uvicorn from dbos import DBOS, DBOSConfig, SetWorkflowID, SQLAlchemyDatasource from fastapi import FastAPI, HTTPException, Response from fastapi.responses import HTMLResponse from .schema import OrderStatus, orders, products app = FastAPI() ds: SQLAlchemyDatasource WIDGET_ID = 1 PAYMENT_STATUS = "payment_status" PAYMENT_ID = "payment_id" ORDER_ID = "order_id" ``` ### Building the Checkout Workflow The heart of this application is the checkout workflow, which orchestrates the entire purchase process. This workflow is triggered whenever a customer buys a widget and handles the complete order lifecycle: 1. Creates a new order in the system 2. Reserves inventory to ensure the item is available 3. Processes payment 4. Marks the order as paid and initiates fulfillment 5. Handles failures gracefully by releasing reserved inventory and canceling orders when necessary DBOS **durably executes** this workflow. It checkpoints each step in the database so that if the app fails or is interrupted during checkout, it will automatically recover from the last completed step. This means that customers never lose their order progress, no matter what breaks. You can try this yourself! [Run the application](#try-it-yourself), start an order, and press the crash button at any time. Within seconds, your app will recover to exactly the state it was in before the crash and continue as if nothing happened. ```python @DBOS.workflow() def checkout_workflow(): # Create a new order order_id = ds.run_tx_step({"name": "create_order"}, create_order) # Attempt to reserve inventory, cancelling the order if no inventory remains. inventory_reserved = ds.run_tx_step( {"name": "reserve_inventory"}, reserve_inventory ) if not inventory_reserved: DBOS.logger.error(f"Failed to reserve inventory for order {order_id}") ds.run_tx_step( {"name": "update_order_status"}, update_order_status, order_id=order_id, status=OrderStatus.CANCELLED.value, ) DBOS.set_event(PAYMENT_ID, None) return # Send a unique payment ID to the checkout endpoint so it # can redirect the customer to the payments page. DBOS.set_event(PAYMENT_ID, DBOS.workflow_id) # Wait for a message that the customer has completed payment. payment_status = DBOS.recv(PAYMENT_STATUS) # If payment succeeded, mark the order as paid and start the order dispatch workflow. # Otherwise, return reserved inventory and cancel the order. if payment_status == "paid": DBOS.logger.info(f"Payment successful for order {order_id}") ds.run_tx_step( {"name": "update_order_status"}, update_order_status, order_id=order_id, status=OrderStatus.PAID.value, ) DBOS.start_workflow(dispatch_order_workflow, order_id) else: DBOS.logger.warning(f"Payment failed for order {order_id}") ds.run_tx_step({"name": "undo_reserve_inventory"}, undo_reserve_inventory) ds.run_tx_step( {"name": "update_order_status"}, update_order_status, order_id=order_id, status=OrderStatus.CANCELLED.value, ) # Finally, send the order ID to the payment endpoint so it # can redirect the customer to the order status page. DBOS.set_event(ORDER_ID, str(order_id)) ``` ### The Checkout and Payment Endpoints Now let's implement the HTTP endpoints that handle customer interactions with the checkout system. The checkout endpoint is triggered when a customer clicks the "Buy Now" button. It starts the checkout workflow in the background, then waits for the workflow to generate and send it a unique payment ID. It then returns the payment ID so the browser can redirect the user to the payments page. The endpoint accepts an [idempotency key](../tutorials/workflow-tutorial.md#workflow-ids-and-idempotency) so that even if the customer presses "buy now" multiple times, only one checkout workflow is started. ```python @app.post("/checkout/{idempotency_key}") def checkout_endpoint(idempotency_key: str) -> Response: # Idempotently start the checkout workflow in the background. with SetWorkflowID(idempotency_key): handle = DBOS.start_workflow(checkout_workflow) # Wait for the checkout workflow to send a payment ID, then return it. payment_id = DBOS.get_event(handle.workflow_id, PAYMENT_ID) if payment_id is None: raise HTTPException(status_code=404, detail="Checkout failed to start") return Response(payment_id) ``` The payment endpoint handles the communication between the payment system and the checkout workflow. It uses the payment ID to signal the checkout workflow whether the payment succeeded or failed. It then retrieves the order ID from the checkout workflow so the browser can redirect the customer to the order status page. ```python @app.post("/payment_webhook/{payment_id}/{payment_status}") def payment_endpoint(payment_id: str, payment_status: str) -> Response: # Send the payment status to the checkout workflow. DBOS.send(payment_id, payment_status, PAYMENT_STATUS) # Wait for the checkout workflow to send an order ID, then return it. order_url = DBOS.get_event(payment_id, ORDER_ID) if order_url is None: raise HTTPException(status_code=404, detail="Payment failed to process") return Response(order_url) ``` ### Database Operations Now, let's implement the checkout workflow's steps. Each step performs a simple CRUD operation, like updating inventory or order status. Each is an ordinary Python function that the workflow runs through [`ds.run_tx_step`](../tutorials/transaction-tutorial.md#inline-with-run_tx_step--run_tx_step_async), so it executes as a durable, exactly-once [database transaction](../tutorials/transaction-tutorial.md). We also expose some of them as HTTP endpoints with FastAPI so the frontend can access them.
Database Operations ```python def reserve_inventory() -> bool: rows_affected = ds.sql_session().execute( products.update() .where(products.c.product_id == WIDGET_ID) .where(products.c.inventory > 0) .values(inventory=products.c.inventory - 1) ).rowcount return rows_affected > 0 def undo_reserve_inventory() -> None: ds.sql_session().execute( products.update() .where(products.c.product_id == WIDGET_ID) .values(inventory=products.c.inventory + 1) ) def create_order() -> int: result = ds.sql_session().execute( orders.insert().values(order_status=OrderStatus.PENDING.value) ) return result.inserted_primary_key[0] def get_order(order_id: int): return ( ds.sql_session().execute(orders.select().where(orders.c.order_id == order_id)) .mappings() .first() ) @app.get("/order/{order_id}") def order_endpoint(order_id: int): return ds.run_tx_step({"name": "get_order"}, get_order, order_id) def update_order_status(order_id: int, status: int) -> None: ds.sql_session().execute( orders.update().where(orders.c.order_id == order_id).values(order_status=status) ) def get_product(): return ds.sql_session().execute(products.select()).mappings().first() @app.get("/product") def product_endpoint(): return ds.run_tx_step({"name": "get_product"}, get_product) def get_orders(): rows = ds.sql_session().execute(orders.select()) return [dict(row) for row in rows.mappings()] @app.get("/orders") def orders_endpoint(): return ds.run_tx_step({"name": "get_orders"}, get_orders) def restock(): ds.sql_session().execute(products.update().values(inventory=100)) @app.post("/restock") def restock_endpoint(): return ds.run_tx_step({"name": "restock"}, restock) @DBOS.workflow() def dispatch_order_workflow(order_id): for _ in range(10): DBOS.sleep(1) ds.run_tx_step( {"name": "update_order_progress"}, update_order_progress, order_id ) def update_order_progress(order_id): # Update the progress of paid orders. progress_remaining = ds.sql_session().execute( orders.update() .where(orders.c.order_id == order_id) .values(progress_remaining=orders.c.progress_remaining - 1) .returning(orders.c.progress_remaining) ).scalar_one() # Dispatch if the order is fully-progressed. if progress_remaining == 0: ds.sql_session().execute( orders.update() .where(orders.c.order_id == order_id) .values(order_status=OrderStatus.DISPATCHED.value) ) ```
### Launching and Serving the App Let's add the final touches to the app. This FastAPI endpoint serves its frontend: ```python @app.get("/") def frontend(): with open(os.path.join("html", "app.html")) as file: html = file.read() return HTMLResponse(html) ``` This FastAPI endpoint crashes the app. Trigger it as many times as you want—DBOS always comes back, resuming from exactly where it left off! ```python @app.post("/crash_application") def crash_application(): os._exit(1) ``` Finally, configure and launch DBOS, then launch the FastAPI server. This is where we create the datasource and initialize DBOS with the app's database connection. ```python if __name__ == "__main__": database_url = os.environ.get("DBOS_DATABASE_URL") if database_url is None: raise Exception("DBOS_DATABASE_URL not set") ds = SQLAlchemyDatasource.create(database_url) config: DBOSConfig = { "name": "widget-store", "application_version": "0.1.0", "system_database_url": database_url, } DBOS(config=config) DBOS.launch() uvicorn.run(app, host="0.0.0.0", port=8000) ``` ### Try it Yourself! Clone and enter the [dbos-demo-apps](https://github.com/dbos-inc/dbos-demo-apps) repository: ```shell git clone https://github.com/dbos-inc/dbos-demo-apps.git cd dbos-demo-apps/python/widget-store ``` Then follow the instructions in the [README](https://github.com/dbos-inc/dbos-demo-apps/tree/main/python/widget-store) to run the app. --- ## Add DBOS To Your App(Python) This guide shows you how to add the open-source [DBOS Transact](https://github.com/dbos-inc/dbos-transact-py) library to your existing application to **durably execute** it and make it resilient to any failure. #### 1. Install DBOS `pip install` DBOS into your application. ```shell pip install dbos ``` #### 2. Add the DBOS Initializer Add these lines of code to your program's main function. They initialize DBOS when your program starts. ```python import os from dbos import DBOS, DBOSConfig config: DBOSConfig = { "name": "my-app", "application_version": "0.1.0", "system_database_url": os.environ.get("DBOS_SYSTEM_DATABASE_URL"), } DBOS(config=config) DBOS.launch() ``` :::info DBOS uses a database to durably store workflow and step state. By default, it uses SQLite, which requires no configuration. For production use, we recommend connecting your DBOS application to a Postgres database. When you're ready for production, you can connect this initialization code to Postgres by setting the `DBOS_SYSTEM_DATABASE_URL` environment variable to a connection string to your Postgres database. ::: #### 3. Start Your Application Try starting your application. If everything is set up correctly, your app should run normally, but log `Initializing DBOS` and `DBOS launched!` on startup. Congratulations! You've integrated DBOS into your application. #### 4. Start Building With DBOS At this point, you can add any DBOS decorator or method to your application. For example, you can annotate one of your functions as a [workflow](./tutorials/workflow-tutorial.md) and the functions it calls as [steps](./tutorials/step-tutorial.md). DBOS durably executes the workflow so if it is ever interrupted, upon restart it automatically resumes from the last completed step. You can add DBOS to your application incrementally—it won't interfere with code that's already there. It's totally okay for your application to have one DBOS workflow alongside thousands of lines of non-DBOS code. To learn more about programming with DBOS, check out [the guide](./programming-guide.md). ```python @DBOS.step() def step_one(): ... @DBOS.step() def step_two(): ... @DBOS.workflow() def workflow(): step_one() step_two() ``` --- ## Learn DBOS Python This guide shows you how to use DBOS to build Python apps that are **resilient to any failure**. :::tip To teach your AI coding assistant to build with DBOS, try out [skills](./prompting.md) and [MCP](../integrations/mcp.md). ::: ### 1. Setting Up Your Environment Create a folder for your app with a virtual environment, then enter the folder and activate the virtual environment. **macOS or Linux** ```shell python3 -m venv dbos-starter/.venv cd dbos-starter source .venv/bin/activate ``` **Windows (PowerShell)** ```shell python3 -m venv dbos-starter/.venv cd dbos-starter .venv\Scripts\activate.ps1 ``` **Windows (cmd)** ```shell python3 -m venv dbos-starter/.venv cd dbos-starter .venv\Scripts\activate.bat ``` Then, install DBOS: ```shell pip install dbos ``` ### 2. Workflows and Steps DBOS helps you add reliability to your Python programs. The key feature of DBOS is **workflow functions** comprised of **steps**. DBOS checkpoints the state of your workflows and steps to its system database. If your program crashes or is interrupted, DBOS uses this checkpointed state to recover each of your workflows from its last completed step. Thus, DBOS makes your application **resilient to any failure**. :::info DBOS uses a database to durably store workflow and step state. By default, it uses SQLite, which requires no configuration. For production use, we recommend connecting your DBOS application to a Postgres database. You can optionally run these examples with Postgres by setting the `DBOS_SYSTEM_DATABASE_URL` environment variable to a connection string to your Postgres database. ::: Let's create a simple DBOS program that runs a workflow of two steps. Create `main.py` and add this code to it: ```python showLineNumbers title="main.py" import os from dbos import DBOS, DBOSConfig @DBOS.step() def step_one(): print("Step one completed!") @DBOS.step() def step_two(): print("Step two completed!") @DBOS.workflow() def dbos_workflow(): step_one() step_two() if __name__ == "__main__": config: DBOSConfig = { "name": "dbos-starter", "application_version": "0.1.0", "system_database_url": os.environ.get("DBOS_SYSTEM_DATABASE_URL"), } DBOS(config=config) DBOS.launch() dbos_workflow() ``` Now, run this code with `python3 main.py`. Your program should print output like: ``` 15:41:06 [ INFO] (dbos:_dbos.py:735) DBOS launched! Step one completed! Step two completed! ``` To see durable execution in action, let's modify the app to serve a DBOS workflow from an HTTP endpoint using FastAPI. Copy this code into `main.py`: ```python showLineNumbers title="main.py" import os import uvicorn from dbos import DBOS, DBOSConfig from fastapi import FastAPI app = FastAPI() @DBOS.step() def step_one(): print("Step one completed!") @DBOS.step() def step_two(): print("Step two completed!") @app.get("/") @DBOS.workflow() def dbos_workflow(): step_one() for _ in range(5): print("Press Control + C to stop the app...") DBOS.sleep(1) step_two() if __name__ == "__main__": config: DBOSConfig = { "name": "dbos-starter", "application_version": "0.1.0", "system_database_url": os.environ.get("DBOS_SYSTEM_DATABASE_URL"), } DBOS(config=config) DBOS.launch() uvicorn.run(app, host="0.0.0.0", port=8000) ``` Now, install FastAPI with: ```shell pip install 'fastapi[standard]' ``` Then, start your app with `python3 main.py`. Then, visit this URL: http://localhost:8000. In your terminal, you should see an output like: ```shell INFO: Uvicorn running on http://0.0.0.0:8000 (Press CTRL+C to quit) Step one completed! Press Control + C to stop the app... Press Control + C to stop the app... ``` Now, press CTRL+C stop your app (press CTRL+C multiple times to force quit it). Then, run `python3 main.py` to restart it. You should see an output like: ```shell INFO: Uvicorn running on http://0.0.0.0:8000 (Press CTRL+C to quit) Press Control + C to stop the app... Press Control + C to stop the app... Press Control + C to stop the app... Press Control + C to stop the app... Press Control + C to stop the app... Step two completed! ``` You can see how DBOS **recovers your workflow from the last completed step**, executing step two without re-executing step one. Learn more about workflows, steps, and their guarantees [here](./tutorials/workflow-tutorial.md). ### 3. Queues and Parallelism If you need to run many functions concurrently, use DBOS _queues_. To try them out, copy this code into `main.py`: ```python showLineNumbers title="main.py" import os import time import uvicorn from dbos import DBOS, DBOSConfig from fastapi import FastAPI app = FastAPI() @DBOS.step() def dbos_step(n: int): time.sleep(5) print(f"Step {n} completed!") @app.get("/") @DBOS.workflow() def dbos_workflow(): print("Enqueueing steps") handles = [] for i in range(10): handle = DBOS.enqueue_workflow("example-queue", dbos_step, i) handles.append(handle) results = [handle.get_result() for handle in handles] print(f"Successfully completed {len(results)} steps") if __name__ == "__main__": config: DBOSConfig = { "name": "dbos-starter", "application_version": "0.1.0", "system_database_url": os.environ.get("DBOS_SYSTEM_DATABASE_URL"), } DBOS(config=config) DBOS.launch() DBOS.register_queue("example-queue") uvicorn.run(app, host="0.0.0.0", port=8000) ``` When you enqueue a function with `DBOS.enqueue_workflow`, DBOS executes it _asynchronously_, running it in the background without waiting for it to finish. `enqueue_workflow` returns a handle representing the state of the enqueued function. This example enqueues ten functions, then waits for them all to finish using `handle.get_result()` to wait for each of their handles. Start your app with `python3 main.py`. Then, visit this URL: http://localhost:8000. Wait five seconds and you should see an output like: ```shell INFO: Uvicorn running on http://0.0.0.0:8000 (Press CTRL+C to quit) Enqueueing steps Step 0 completed! Step 1 completed! Step 2 completed! Step 3 completed! Step 4 completed! Step 5 completed! Step 6 completed! Step 7 completed! Step 8 completed! Step 9 completed! Successfully completed 10 steps ``` You can see how all ten steps run concurrently—even though each takes five seconds, they all finish at the same time. Learn more about DBOS queues [here](./tutorials/queue-tutorial.md). ### 4. Connecting to DBOS Conductor [Conductor](../conductor/overview.md) is the control plane for your durable workflows, providing distributed workflow recovery, observability, and management. Once you connect your app to Conductor, you can view and manage all its workflows and queued tasks from the [DBOS Console](https://console.dbos.dev). To connect your app to Conductor, first sign up for an account on the [DBOS Console](https://console.dbos.dev/login-redirect). Then, install [`dbosctl`](../conductor/reference/dbosctl.md), the Conductor command-line client. On Windows, [download a release binary](https://github.com/dbos-inc/dbos-ctl/releases) instead. ```shell curl -sSfL https://raw.githubusercontent.com/dbos-inc/dbos-ctl/main/install.sh | sh ``` Next, configure a `dbosctl` profile and log in. `dbosctl login` prints a URL and a code for you to approve in your browser. ```shell dbosctl config set dbos --managed dbosctl login ``` Then, register your application with Conductor and create an API key. The name you register must match the `name` in your DBOS configuration. The key's secret is printed once and cannot be retrieved afterwards, so copy it now. ```shell dbosctl app register dbos-starter dbosctl api-key create dbos-starter-key ``` Next, supply your API key to your app through the `conductor_key` configuration option. Update the configuration in `main.py` to read the key from an environment variable: ```python showLineNumbers title="main.py" config: DBOSConfig = { "name": "dbos-starter", "application_version": "0.1.0", "system_database_url": os.environ.get("DBOS_SYSTEM_DATABASE_URL"), "conductor_key": os.environ.get("DBOS_CONDUCTOR_KEY"), } ``` Finally, set the `DBOS_CONDUCTOR_KEY` environment variable to the key you created and restart your app: ```shell export DBOS_CONDUCTOR_KEY= python3 main.py ``` Your app is now connected to Conductor! Launch a workflow by visiting http://localhost:8000, then watch it execute in real time from the [DBOS Console](https://console.dbos.dev). Learn more about Conductor [here](../conductor/overview.md). Congratulations! You've finished the DBOS Python guide. Next, you should: - Learn how to [**add DBOS to your own application**](./integrating-dbos.md). - Check out some [**example applications**](../examples/index.md). --- ## AI-Assisted Development If you're using an AI coding agent to build a DBOS application, make sure it has the latest information on DBOS by either: 1. [Installing DBOS skills.](#dbos-agent-skills) 2. [Providing your agent with a DBOS prompt.](#dbos-prompt) You may also want to use the [DBOS MCP server](../integrations/mcp.md) so your model can directly access your application's workflows and steps. ### DBOS Agent Skills [Agent Skills](https://agentskills.io/home) help developers use AI agents to add DBOS durable workflows to their applications. DBOS provides open-source skills you can check out [here](https://github.com/dbos-inc/agent-skills). To install them into your coding agent, run: ``` npx skills add dbos-inc/agent-skills ``` The [Skills CLI](https://skills.sh/) is compatible with most coding agents, including Claude Code, Codex, Antigravity, and Cursor. ### DBOS Prompt You can use this prompt to add rich information about DBOS to your AI coding agent's context. You can copy and paste it directly into your context, or follow these directions to add it to your AI-powered IDE or coding agent of choice: - Claude Code: Add the prompt, or a link to it, to your CLAUDE.md file. - Cursor: Add the prompt to [your project rules](https://docs.cursor.com/context/rules-for-ai). - GitHub Copilot: Create a [`.github/copilot-instructions.md`](https://docs.github.com/en/copilot/customizing-copilot/adding-repository-custom-instructions-for-github-copilot) file in your repository and add the prompt to it.
DBOS Python Prompt ````markdown # Build Reliable Applications With DBOS ## Guidelines - Respond in a friendly and concise manner - Ask clarifying questions when requirements are ambiguous - Generate code in Python using the DBOS library - You MUST import all methods and classes used in the code you generate - You SHALL keep all code in a single file unless otherwise specified. - DBOS does NOT stand for anything. ## Workflow Guidelines Workflows provide durable execution so you can write programs that are resilient to any failure. Workflows are comprised of steps, which are ordinary Python functions annotated with @DBOS.step(). When using DBOS workflows, you should annotate any function that performs complex operations or accesses external APIs or services as a step. You can turn any Python function into a step by annotating it with the @DBOS.step decorator. The only requirement is that its inputs and outputs should be serializable. If a workflow is interrupted for any reason (e.g., an executor restarts or crashes), when your program restarts the workflow automatically resumes execution from the last completed step. - If asked to add DBOS to existing code, you MUST ask which function to make a workflow. Do NOT recommend any changes until they have told you what function to make a workflow. Do NOT make a function a workflow unless SPECIFICALLY requested. - When making a function a workflow, you should make all functions it calls steps. Do NOT change the functions in any way except by adding the @DBOS.step() annotation. - Do NOT make functions steps unless they are DIRECTLY called by a workflow. - If the workflow function performs a non-deterministic action, you MUST move that action to its own function and make that function a step. Examples of non-deterministic actions include accessing an external API or service, accessing files on disk, generating a random number, or getting the current time. - Do NOT use threads to start workflows or to start steps in workflows. You should instead use DBOS.start_workflow and DBOS queues. - DBOS workflows and steps should NOT have side effects in memory outside of their own scope. They can access global variables, but they should NOT create or update global variables or variables outside their scope. - Do NOT call DBOS.start_workflow or DBOS.recv from a step - Do NOT start workflows from inside a step. - Do NOT call DBOS.set_event and DBOS.recv from outside a workflow. ## DBOS Lifecycle Guidelines A DBOS application MUST always be configured like so, unless otherwise specified, configuring and launching DBOS in its main function: ```python if __name__ == "__main__": config: DBOSConfig = { "name": "my-app", "application_version": "0.1.0", "system_database_url": os.environ.get("DBOS_SYSTEM_DATABASE_URL"), } DBOS(config=config) DBOS.launch() ``` In a FastAPI application, the server should ALWAYS be started explicitly after a DBOS.launch in the main function: ```python if __name__ == "__main__": config: DBOSConfig = { "name": "my-app", "application_version": "0.1.0", "system_database_url": os.environ.get("DBOS_SYSTEM_DATABASE_URL"), } DBOS(config=config) DBOS.launch() uvicorn.run(app, host="0.0.0.0", port=8000) ``` If an app contains scheduled workflows and NOTHING ELSE (no HTTP server), then the main thread should block forever while the scheduled workflows run like this: ```python if __name__ == "__main__": config: DBOSConfig = { "name": "my-app", "application_version": "0.1.0", "system_database_url": os.environ.get("DBOS_SYSTEM_DATABASE_URL"), } DBOS(config=config) DBOS.launch() threading.Event().wait() ``` Or if using asyncio: ```python import asyncio from dbos import DBOS, DBOSConfig async def main(): config: DBOSConfig = { "name": "dbos-app", "application_version": "0.1.0", } DBOS(config=config) DBOS.launch() await asyncio.Event().wait() if __name__ == "__main__": asyncio.run(main()) ``` ### Workflow and Steps Examples Simple example: ```python import os from dbos import DBOS, DBOSConfig @DBOS.step() def step_one(): print("Step one completed!") @DBOS.step() def step_two(): print("Step two completed!") @DBOS.workflow() def dbos_workflow(): step_one() step_two() if __name__ == "__main__": config: DBOSConfig = { "name": "dbos-starter", "application_version": "0.1.0", "system_database_url": os.environ.get("DBOS_SYSTEM_DATABASE_URL"), } DBOS(config=config) DBOS.launch() dbos_workflow() ``` Example with FastAPI: ```python import os import uvicorn from dbos import DBOS, DBOSConfig from fastapi import FastAPI app = FastAPI() @DBOS.step() def step_one(): print("Step one completed!") @DBOS.step() def step_two(): print("Step two completed!") @app.get("/") @DBOS.workflow() def dbos_workflow(): step_one() step_two() if __name__ == "__main__": config: DBOSConfig = { "name": "dbos-starter", "application_version": "0.1.0", "system_database_url": os.environ.get("DBOS_SYSTEM_DATABASE_URL"), } DBOS(config=config) DBOS.launch() uvicorn.run(app, host="0.0.0.0", port=8000) ``` Example with queues: ```python import os import time import uvicorn from dbos import DBOS, DBOSConfig from fastapi import FastAPI app = FastAPI() @DBOS.step() def dbos_step(n: int): time.sleep(5) print(f"Step {n} completed!") @app.get("/") @DBOS.workflow() def dbos_workflow(): print("Enqueueing steps") handles = [] for i in range(10): handle = DBOS.enqueue_workflow("example-queue", dbos_step, i) handles.append(handle) results = [handle.get_result() for handle in handles] print(f"Successfully completed {len(results)} steps") if __name__ == "__main__": config: DBOSConfig = { "name": "dbos-starter", "application_version": "0.1.0", "system_database_url": os.environ.get("DBOS_SYSTEM_DATABASE_URL"), } DBOS(config=config) DBOS.launch() DBOS.register_queue("example-queue") uvicorn.run(app, host="0.0.0.0", port=8000) ``` #### Scheduled Workflows You can schedule DBOS workflows to run on a cron schedule. Schedules are stored in the database and can be created, paused, resumed, and deleted at runtime. A scheduled workflow MUST take two arguments: a `datetime` (the scheduled execution time) and a context object: ```python from datetime import datetime from typing import Any from dbos import DBOS @DBOS.workflow() def my_periodic_task(scheduled_time: datetime, context: Any): DBOS.logger.info(f"Running task scheduled for {scheduled_time}") DBOS.create_schedule( schedule_name="my-task-schedule", workflow_fn=my_periodic_task, schedule="*/5 * * * *", # Every 5 minutes ) ``` - You MUST create schedules after `DBOS.launch()`; schedule methods raise if DBOS has not been launched. - Use `DBOS.create_schedule` to create a schedule with a crontab expression. It raises if a schedule with that name already exists. - Use `DBOS.pause_schedule` and `DBOS.resume_schedule` to pause and resume schedules. - Use `DBOS.delete_schedule` to delete a schedule. - Use `DBOS.apply_schedules` to atomically create or update multiple schedules at once. To define static schedules on program start, use `DBOS.apply_schedules`, which updates schedules that already exist. - Use `DBOS.list_schedules` and `DBOS.get_schedule` to inspect schedules. - Use `DBOS.backfill_schedule` to enqueue missed executions for a time range. - Use `DBOS.trigger_schedule` to immediately trigger a schedule. - Each workflow enqueued by a schedule is tagged with its schedule's name (recorded in the workflow's status). Retrieve all runs of a schedule with `DBOS.list_workflows(schedule_name="my-task-schedule")`. ### Workflow Documentation: --- sidebar_position: 10 title: Workflows toc_max_heading_level: 3 --- Workflows provide **durable execution** so you can write programs that are **resilient to any failure**. Workflows help you write fault-tolerant background tasks, data processing pipelines, AI agents, and more. You can make a function a workflow by annotating it with `@DBOS.workflow()`. Workflows call steps, which are Python functions annotated with `@DBOS.step()`. If a workflow is interrupted for any reason, DBOS automatically recovers its execution from the last completed step. Here's an example of a workflow: ```python @DBOS.step() def step_one(): print("Step one completed!") @DBOS.step() def step_two(): print("Step two completed!") @DBOS.workflow() def workflow(): step_one() step_two() ``` ### Starting Workflows In The Background One common use-case for workflows is building reliable background tasks that keep running even when the program is interrupted, restarted, or crashes. You can use `DBOS.start_workflow` to start a workflow in the background. If you start a workflow this way, it returns a workflow handle, from which you can access information about the workflow or wait for it to complete and retrieve its result. Here's an example: ```python @DBOS.workflow() def background_task(input): # ... return output # Start the background task handle: WorkflowHandle = DBOS.start_workflow(background_task, input) # Wait for the background task to complete and retrieve its result. output = handle.get_result() ``` After starting a workflow in the background, you can use `DBOS.retrieve_workflow` to retrieve a workflow's handle from its ID. You can also retrieve a workflow's handle from outside of your DBOS application with `DBOSClient.retrieve_workflow`. If you need to run many workflows in the background and manage their concurrency or flow control, you can also use DBOS queues. ### Workflow IDs and Idempotency Every time you execute a workflow, that execution is assigned a unique ID, by default a UUID. You can access this ID through the `DBOS.workflow_id` context variable. Workflow IDs are useful for communicating with workflows and developing interactive workflows. You can set the workflow ID of a workflow with `SetWorkflowID`. Workflow IDs must be **globally unique** for your application. An assigned workflow ID acts as an idempotency key: if a workflow is called multiple times with the same ID, it executes only once. This is useful if your operations have side effects like making a payment or sending an email. For example: ```python @DBOS.workflow() def example_workflow(): DBOS.logger.info(f"I am a workflow with ID {DBOS.workflow_id}") with SetWorkflowID("very-unique-id"): example_workflow() ``` By default, starting a workflow with an ID that is already in use returns the existing workflow instead of starting a new one. To instead raise `DBOSWorkflowIDInUseError`, use `SetWorkflowID("very-unique-id", workflow_id_reuse_policy="reject")` (catch it with `from dbos import error as dboserror` and `except dboserror.DBOSWorkflowIDInUseError`). The same option is available as `workflow_id_reuse_policy` in `DBOSClient.enqueue` options. ### Determinism Workflows are in most respects normal Python functions. They can have loops, branches, conditionals, and so on. However, a workflow function must be **deterministic**: if called multiple times with the same inputs, it should invoke the same steps with the same inputs in the same order (given the same return values from those steps). If you need to perform a non-deterministic operation like accessing the database, calling a third-party API, generating a random number, or getting the local time, you shouldn't do it directly in a workflow function. Instead, you should do all database operations in datasource transactions and all other non-deterministic operations in steps. For example, **don't do this**: ```python @DBOS.workflow() def example_workflow(): choice = random.randint(0, 1) if choice == 0: step_one() else: step_two() ``` Do this instead: ```python @DBOS.step() def generate_choice(): return random.randint(0, 1) @DBOS.workflow() def example_workflow(friend: str): choice = generate_choice() if choice == 0: step_one() else: step_two() ``` If DBOS detects that a single execution of a workflow recorded different results for the same step, it raises a `DBOSStepNondeterminismError`, indicating the workflow is not deterministic. ### Workflow Timeouts You can set a timeout for a workflow with `SetWorkflowTimeout`. When the timeout expires, the workflow **and all its children** are cancelled. Cancelling a workflow sets its status to `CANCELLED` and preempts its execution at the beginning of its next step. Timeouts are **start-to-completion**: if a workflow is enqueued, the timeout does not begin until the workflow is dequeued and starts execution. Also, timeouts are **durable**: they are stored in the database and persist across restarts, so workflows can have very long timeouts. Example syntax: ```python @DBOS.workflow() def example_workflow(): ... # If the workflow does not complete within 10 seconds, it times out and is cancelled with SetWorkflowTimeout(10): example_workflow() ``` ### Workflow Attributes You can attach a dictionary of custom, JSON-serializable key-value attributes to your workflows with `SetWorkflowAttributes`. Every workflow started or enqueued inside the block is recorded with those attributes. This is useful for tagging workflows with application-specific metadata (such as a customer ID, tenant, or region) so you can find them later. Attributes are recorded at creation time and are **not** inherited by child workflows. ```python from dbos import DBOS, SetWorkflowAttributes with SetWorkflowAttributes({"customer": "acme", "region": "us-east-1"}): example_workflow() ``` Attributes are stored in Postgres as GIN-indexed JSONB, so you can efficiently search for workflows by attribute by passing the `attributes` filter to `DBOS.list_workflows` or `DBOS.list_queued_workflows`. A workflow matches if its attributes contain all the key-value pairs you provide (filtering by attribute requires a Postgres system database): ```python # Retrieve all workflows tagged with this customer workflows = DBOS.list_workflows(attributes={"customer": "acme"}) ``` To change a workflow's attributes after it is created, use `DBOS.update_workflow_attributes`, which replaces the workflow's entire attributes dictionary (pass `None` to clear them). This is safe to call from within a workflow, including on the workflow itself, because the update is recorded as a step. ```python # Replace the attributes of a workflow by ID DBOS.update_workflow_attributes(workflow_id, {"customer": "acme", "phase": "processing"}) ``` ### Durable Sleep You can use `DBOS.sleep()` to put your workflow to sleep for any period of time. This sleep is **durable**—DBOS saves the wakeup time in the database so that even if the workflow is interrupted and restarted multiple times while sleeping, it still wakes up on schedule. Sleeping is useful for scheduling a workflow to run in the future (even days, weeks, or months from now). For example: ```python @DBOS.workflow() def schedule_task(time_to_sleep, task): # Durably sleep for some time before running the task DBOS.sleep(time_to_sleep) run_task(task) ``` ### Debouncing Workflows You can create a `Debouncer` to debounce your workflows. Debouncing delays workflow execution until some time has passed since the workflow has last been called. This is useful for preventing wasted work when a workflow may be triggered multiple times in quick succession. For example, if a user is editing an input field, you can debounce their changes to execute a processing workflow only after they haven't edited the field for some time: #### Debouncer.create ```python Debouncer.create( workflow: Callable[P, R], *, debounce_timeout_sec: Optional[float] = None, queue: Optional[Union[Queue, str]] = None, ) -> Debouncer[P, R] ``` **Parameters:** - `workflow`: The workflow to debounce. - `debounce_timeout_sec`: After this time elapses since the first time a workflow is submitted from this debouncer, the workflow is started regardless of the debounce period. - `queue`: When starting a workflow after debouncing, enqueue it on this queue (a queue name or `Queue`) instead of an internal queue. #### debounce ```python debouncer.debounce( debounce_key: str, debounce_period_sec: float, *args: P.args, **kwargs: P.kwargs, ) -> WorkflowHandle[R] ``` Submit a workflow for execution but delay it by `debounce_period_sec`. Returns a handle to the workflow. The workflow may be debounced again, which further delays its execution (up to `debounce_timeout_sec`). When the workflow eventually executes, it uses the **last** set of inputs passed into `debounce`. Once the debounce period expires and the workflow is released for execution, the next call to `debounce` starts the debouncing process again for a new workflow execution. **Parameters:** - `debounce_key`: A key used to group workflow executions that will be debounced together. For example, if the debounce key is set to customer ID, each customer's workflows would be debounced separately. - `debounce_period_sec`: Delay this workflow's execution by this period. - `*args`: Variadic workflow arguments. - `**kwargs`: Variadic workflow keyword arguments. **Example Syntax**: ```python @DBOS.workflow() def process_input(user_input): ... # Each time a user submits a new input, debounce the process_input workflow. # The workflow will wait until 60 seconds after the user stops submitting new inputs, # then process the last input submitted. debouncer = Debouncer.create(process_input) def on_user_input_submit(user_id, user_input): debounce_key = user_id debounce_period_sec = 60 debouncer.debounce(debounce_key, debounce_period_sec, user_input) ``` #### Debouncer.create_async ```python Debouncer.create_async( workflow: Callable[P, Coroutine[Any, Any, R]], *, debounce_timeout_sec: Optional[float] = None, queue: Optional[Union[Queue, str]] = None, ) -> Debouncer[P, R] ``` Async version of `Debouncer.create`. #### debounce_async ```python debouncer.debounce_async( debounce_key: str, debounce_period_sec: float, *args: P.args, **kwargs: P.kwargs, ) -> WorkflowHandleAsync[R]: ``` Async version of `debouncer.debounce`. ### Coroutine (Async) Workflows Coroutines (functions defined with `async def`, also known as async functions) can also be DBOS workflows. Coroutine workflows may invoke coroutine steps via await expressions. You should start coroutine workflows using `DBOS.start_workflow_async` and enqueue them using `DBOS.enqueue_workflow_async`. Calling a coroutine workflow or starting it with `DBOS.start_workflow_async` always runs it in the same event loop as its caller, but a workflow enqueued with `DBOS.enqueue_workflow_async` is started by DBOS in the event loop in which `DBOS.launch()` was called (if that loop is still running) or otherwise in a separate background event loop. Additionally, coroutine workflows should use the asynchronous versions of the workflow communication context methods. Many synchronous DBOS methods (such as `DBOS.sleep`, `DBOS.recv`, `DBOS.send`, `DBOS.set_event`, `DBOS.get_event`, and `DBOS.register_queue`) raise a `RuntimeError` when called while an event loop is running; you MUST use their `_async` variants (such as `await DBOS.sleep_async(...)`) in async code. :::tip For async database operations, use an `AsyncSQLAlchemyDatasource`, which supports `async def` transaction functions (see Datasources below). To call a blocking function from an async workflow without blocking the event loop, use `asyncio.to_thread`. ::: ```python @DBOS.step() async def example_step(): async with aiohttp.ClientSession() as session: async with session.get("https://example.com") as response: return await response.text() @DBOS.workflow() async def example_workflow(friend: str): await DBOS.sleep_async(10) body = await example_step() # asyncio.to_thread runs a synchronous, blocking function without stalling the event loop result = await asyncio.to_thread(blocking_function, body) return result ``` #### Running Async Steps In Parallel Initiating several concurrent steps in an `async` workflow, followed by awaiting them with `asyncio.gather(..., return_exceptions=True)`, is valid as long as the steps are started in a **deterministic order**. For example, the following is allowed: ```python # Start steps in a deterministic order (step1, step2, step3, step4), # then await them all together. tasks = [ asyncio.create_task(step1("arg1")), asyncio.create_task(step2("arg2")), asyncio.create_task(step3("arg3")), asyncio.create_task(step4("arg4")), ] # Collects exceptions instead of raising immediately results = await asyncio.gather(*tasks, return_exceptions=True) return results ``` This is allowed because each step is started in a well-defined sequence before awaiting. By contrast, the following is not allowed: ```python async def seq_a(): await step1("arg1") await step2("arg3") async def seq_b(): await step3("arg2") await step4("arg4") tasks = [ asyncio.create_task(seq_a()), asyncio.create_task(seq_b()), ] results = await asyncio.gather(*tasks, return_exceptions=True) return results ``` Here, `step2` and `step4` may be started in either order since their execution depends on the relative time taken by `step1` and `step3`. If you need to run sequences of operations concurrently, start child workflows and await their results, rather than interleaving step execution inside a single workflow. For proper error handling, when using `asyncio.gather()`, specify `return_exceptions=True`. Without `return_exceptions=True`, `gather` will raise any exception immediately and stop awaiting the rest of the tasks. If one of the remaining tasks later fails, its exception may go unobserved. Instead, prefer `asyncio.gather(..., return_exceptions=True)`, which safely waits for all tasks to complete and reports their outcomes. ### Communicating with Workflows DBOS provides a few different ways to communicate with your workflows. You can: - Send messages to workflows - Publish events from workflows for clients to read - Stream values from workflows to clients ### Workflow Messaging and Notifications You can send messages to a specific workflow. This is useful for signaling a workflow or sending notifications to it while it's running. #### Send ```python DBOS.send( destination_id: str, message: Any, topic: Optional[str] = None, *, idempotency_key: Optional[str] = None, ) -> None ``` You can call `DBOS.send()` to send a message to a workflow. Messages can optionally be associated with a topic and are queued on the receiver per topic. You can also call `send` from outside of your DBOS application with the DBOS Client. #### Recv ```python DBOS.recv( topic: Optional[str] = None, timeout_seconds: float = 60, ) -> Any ``` Workflows can call `DBOS.recv()` to receive messages sent to them, optionally for a particular topic. Each call to `recv()` waits for and consumes the next message to arrive in the queue for the specified topic, returning `None` if the wait times out. If the topic is not specified, this method only receives messages sent without a topic. #### Messages Example Messages are especially useful for sending notifications to a workflow. For example, in the widget store demo, the checkout workflow, after redirecting customers to a payments page, must wait for a notification that the user has paid. To wait for this notification, the payments workflow uses `recv()`, executing failure-handling code if the notification doesn't arrive in time: ```python @DBOS.workflow() def checkout_workflow(): ... # Validate the order, then redirect customers to a payments service. payment_status = DBOS.recv(PAYMENT_STATUS) if payment_status is not None and payment_status == "paid": ... # Handle a successful payment. else: ... # Handle a failed payment or timeout. ``` An endpoint waits for the payment processor to send the notification, then uses `send()` to forward it to the workflow: ```python @app.post("/payment_webhook/{payment_id}/{payment_status}") def payment_endpoint(payment_id: str, payment_status: str) -> Response: # Send the payment status to the checkout workflow. DBOS.send(payment_id, payment_status, PAYMENT_STATUS) ``` #### Reliability Guarantees All messages are persisted to the database, so if `send` completes successfully, the destination workflow is guaranteed to be able to `recv` it. If you're sending a message from a workflow, DBOS guarantees exactly-once delivery. If you're sending a message from normal Python code, you can pass an `idempotency_key` to `DBOS.send` to guarantee exactly-once delivery: the message is sent only once to each destination no matter how many times `DBOS.send` is called with that key. ### Workflow Events Workflows can publish _events_, which are key-value pairs associated with the workflow. They are useful for publishing information about the status of a workflow or to send a result to clients while the workflow is running. #### set_event ```python DBOS.set_event( key: str, value: Any, ) -> None ``` Any workflow or step can call `DBOS.set_event` to publish a key-value pair, or update its value if it has already been published. #### get_event ```python DBOS.get_event( workflow_id: str, key: str, timeout_seconds: float = 60, ) -> Any ``` You can call `DBOS.get_event` to retrieve the value published by a particular workflow identity for a particular key. If the event does not yet exist, this call waits for it to be published, returning `None` if the wait times out. You can also call `get_event` from outside of your DBOS application with DBOS Client. #### get_all_events ```python DBOS.get_all_events( workflow_id: str ) -> Dict[str, Any] ``` You can use `DBOS.get_all_events` to retrieve the latest values of all events published by a workflow. #### Events Example Events are especially useful for writing interactive workflows that communicate information to their caller. For example, in the widget store demo, the checkout workflow, after validating an order, needs to send the customer a unique payment ID. To communicate the payment ID to the customer, it uses events. The payments workflow emits the payment ID using `set_event()`: ```python @DBOS.workflow() def checkout_workflow(): ... payment_id = ... DBOS.set_event(PAYMENT_ID, payment_id) ... ``` The FastAPI handler that originally started the workflow uses `get_event()` to await this payment ID, then returns it: ```python @app.post("/checkout/{idempotency_key}") def checkout_endpoint(idempotency_key: str) -> Response: # Idempotently start the checkout workflow in the background. with SetWorkflowID(idempotency_key): handle = DBOS.start_workflow(checkout_workflow) # Wait for the checkout workflow to send a payment ID, then return it. payment_id = DBOS.get_event(handle.workflow_id, PAYMENT_ID) if payment_id is None: raise HTTPException(status_code=404, detail="Checkout failed to start") return Response(payment_id) ``` #### Reliability Guarantees All events are persisted to the database, so the latest version of an event is always retrievable. Additionally, if `get_event` is called in a workflow, the retrieved value is persisted in the database so workflow recovery can use that value, even if the event is later updated. ### Workflow Streaming Workflows can stream data in real time to clients. This is useful for streaming results from a long-running workflow or LLM call or for monitoring or progress reporting. #### Writing to Streams ```python DBOS.write_stream( key: str, value: Any ) -> None: ``` You can write values to a stream from a workflow or its steps using `DBOS.write_stream`. A workflow may have any number of streams, each identified by a unique key. When you are done writing to a stream, you should close it with `DBOS.close_stream`, which you can call from a workflow or its steps. Otherwise, streams are automatically closed when the workflow terminates. ```python DBOS.close_stream( key: str ) -> None ``` DBOS streams are immutable and append-only. Writes to a stream from a workflow happen exactly-once. Writes to a stream from a step happen at-least-once; if a step fails and is retried, it may write to the stream multiple times. Readers will see all values written to the stream from all tries of the step in the order in which they were written. **Example syntax:** ```python @DBOS.workflow() def producer_workflow(): DBOS.write_stream(example_key, {"step": 1, "data": "value1"}) DBOS.write_stream(example_key, {"step": 2, "data": "value2"}) DBOS.close_stream(example_key) # Signal completion ``` #### Reading from Streams ```python DBOS.read_stream( workflow_id: str, key: str, *, offset: int = 0, polling_interval_sec: Optional[float] = None, timeout_seconds: Optional[float] = None, ) -> Generator[Any, Any, None] ``` You can read values from a stream from anywhere using `DBOS.read_stream`. This function reads values from a stream identified by a workflow ID and key, yielding each value in order until the stream is closed or the workflow terminates. You can also read from a stream from outside a DBOS application with a DBOS Client. **Parameters:** - `offset`: The offset to start reading from. Defaults to `0`, the start of the stream. - `polling_interval_sec`: Polling interval in seconds when waiting for new values when not using LISTEN/NOTIFY. Must be at least `0.001`. Defaults to the configured `notification_listener_polling_interval_sec` (`1.0` if not configured). - `timeout_seconds`: How long to wait for **each** value before raising `DBOSStreamTimeoutError`. The clock restarts every time a value is delivered, so this bounds the gap between values, not the total duration of the read. Defaults to `None`, waiting indefinitely. **Example syntax:** ```python for value in DBOS.read_stream(workflow_id, example_key): print(f"Received: {value}") ``` ```python from dbos import error as dboserror try: for value in DBOS.read_stream(workflow_id, example_key, timeout_seconds=30): print(f"Received: {value}") except dboserror.DBOSStreamTimeoutError: print("The producer stopped sending values") ``` #### Configurable Retries You can optionally configure a step to automatically retry any exception a set number of times with exponential backoff. This is useful for automatically handling transient failures, like making requests to unreliable APIs. Retries are configurable through arguments to the step decorator: ```python DBOS.step( *, retries_allowed: bool = False, interval_seconds: float = 1.0, max_attempts: int = 3, backoff_rate: float = 2.0, timeout_seconds: Optional[float] = None, ) ``` For example, we configure this step to retry exceptions (such as if `example.com` is temporarily down) up to 10 times: ```python @DBOS.step(retries_allowed=True, max_attempts=10) def example_step(): return requests.get("https://example.com").text ``` #### Step Timeouts You can set a timeout for an async step with `timeout_seconds`. If the step runs longer than that, it is cancelled and `DBOSStepTimeoutError` is raised to the calling workflow. ```python @DBOS.step(timeout_seconds=30) async def example_step(): async with aiohttp.ClientSession() as session: async with session.get("https://example.com") as response: return await response.text() ``` Step timeouts are only supported for async steps, because Python provides no way to preempt a running synchronous function. Setting `timeout_seconds` on a sync step raises an exception. The timeout must be positive and finite. If the step also has retries enabled, each attempt gets its own timeout, and time spent waiting between retries is not counted against it. If every attempt times out, the workflow sees `DBOSMaxStepRetriesExceeded` rather than `DBOSStepTimeoutError`. Like any other step failure, the timeout is checkpointed as the step's outcome, so a workflow that recovers after one re-raises the same error instead of re-running the step. The timeout is enforced only when the step runs as part of a workflow; calling the function directly outside a workflow runs it as an ordinary function call, with no timeout. ### DBOS Queues You can use queues to run many workflows at once with managed concurrency. Queues provide _flow control_, letting you manage how many workflows run at once or how often workflows are started. Register a queue with `DBOS.register_queue`, specifying its name and optional flow control parameters: ```python DBOS.register_queue( name: str, *, # Applied to the queue as a whole global_concurrency: Optional[int] = None, worker_concurrency: Optional[int] = None, limiter: Optional[QueueRateLimit] = None, # Applied to each partition separately partition_concurrency: Optional[int] = None, partition_worker_concurrency: Optional[int] = None, partition_limiter: Optional[QueueRateLimit] = None, polling_interval_sec: float = 1.0, on_conflict: QueueConflictResolution = "update_if_latest_version", ) -> Queue class QueueRateLimit(TypedDict): limit: int period: float # In seconds ``` **Parameters:** - `name`: The name of the queue. Must be unique among all queues in the system database, including those of other applications sharing it. Names starting with `_dbos_` are reserved. - `global_concurrency`: The maximum number of functions from this queue that may run concurrently across all DBOS processes. If not provided, any number of functions may run concurrently. - `worker_concurrency`: The maximum number of functions from this queue that may run concurrently on a given DBOS process. Must be less than or equal to `global_concurrency`. - `limiter`: A limit on the maximum number of functions which may be started in a given period. - `partition_concurrency`: The maximum number of functions from any one partition of this queue that may run concurrently across all DBOS processes. Must be at least 1 and less than or equal to `global_concurrency`. - `partition_worker_concurrency`: The maximum number of functions from any one partition of this queue that may run concurrently on a given DBOS process. Must be at least 1 and less than or equal to `partition_concurrency`, `worker_concurrency`, and `global_concurrency`. - `partition_limiter`: A limit on the maximum number of functions which may be started from any one partition in a given period. - `polling_interval_sec`: The interval at which DBOS polls the database for new workflows on this queue. - `on_conflict`: How to behave when a queue with this name already exists in the system database. Defaults to `"update_if_latest_version"` (only overwrite if the running app is the latest version) to keep older workers in a rolling deploy from clobbering newer config. Other options: `"always_update"`, `"never_update"`. Setting any `partition_*` limit makes the queue partitioned: every enqueue must supply a `queue_partition_key`, and deduplication is not supported. The queue-wide limits still apply across all partitions. Queues are persisted to the system database, so they are visible to every DBOS process and client connected to that database. You MUST register your queues after `DBOS.launch()` (for example, in your main function); `DBOS.register_queue` raises if DBOS has not been launched. For brevity, some snippets below omit the surrounding launch code. In async code, use `await DBOS.register_queue_async(...)` instead. **Example syntax:** ```python DBOS.register_queue("example_queue") ``` You can then enqueue any DBOS workflow or step with `DBOS.enqueue_workflow`. Enqueuing a function submits it for execution and returns a handle to it. Queued tasks are started in first-in, first-out (FIFO) order. ```python DBOS.register_queue("example_queue") @DBOS.workflow() def process_task(task): ... task = ... handle = DBOS.enqueue_workflow("example_queue", process_task, task) ``` #### Queue Example Here's an example of a workflow using a queue to process tasks concurrently: ```python from dbos import DBOS DBOS.register_queue("example_queue") @DBOS.workflow() def process_task(task): ... @DBOS.workflow() def process_tasks(tasks): task_handles = [] # Enqueue each task so all tasks are processed concurrently. for task in tasks: handle = DBOS.enqueue_workflow("example_queue", process_task, task) task_handles.append(handle) # Wait for each task to complete and retrieve its result. # Return the results of all tasks. return [handle.get_result() for handle in task_handles] ``` #### Enqueueing from Another Application Often, you want to enqueue a workflow from outside your DBOS application. For example, let's say you have an API server and a data processing service. You're using DBOS to build a durable data pipeline in the data processing service. When the API server receives a request, it should enqueue the data pipeline for execution on the data processing service. You can use the DBOS Client to register queues and enqueue workflows from outside your DBOS application by connecting directly to your DBOS application's system database. Since the DBOS Client is designed to be used from outside your DBOS application, workflow and queue metadata must be specified explicitly. For example, this code registers `pipeline_queue` and enqueues the `data_pipeline` workflow on it with `task` as an argument. ```python import os from dbos import DBOSClient, EnqueueOptions client = DBOSClient(system_database_url=os.environ["DBOS_SYSTEM_DATABASE_URL"]) # Register the queue from the client. client.register_queue("pipeline_queue") options: EnqueueOptions = { "queue_name": "pipeline_queue", "workflow_name": "data_pipeline", } handle = client.enqueue(options, task) result = handle.get_result() ``` #### Managing Concurrency You can control how many workflows from a queue run simultaneously by configuring concurrency limits. This helps prevent resource exhaustion when workflows consume significant memory or processing power. ##### Worker Concurrency Worker concurrency sets the maximum number of workflows from a queue that can run concurrently on a single DBOS process. This is particularly useful for resource-intensive workflows to avoid exhausting the resources of any process. For example, this queue has a worker concurrency of 5, so each process will run at most 5 workflows from this queue simultaneously: ```python DBOS.register_queue("example_queue", worker_concurrency=5) ``` ##### Global Concurrency Global concurrency limits the total number of workflows from a queue that can run concurrently across all DBOS processes in your application. For example, this queue will have a maximum of 10 workflows running simultaneously across your entire application. :::warning Worker concurrency limits are recommended for most use cases. Take care when using a global concurrency limit as any `PENDING` workflow on the queue counts toward the limit, including workflows from previous application versions ::: ```python DBOS.register_queue("example_queue", global_concurrency=10) ``` ##### In-Order Processing You can use a queue with `global_concurrency=1` to guarantee sequential, in-order processing of events. Only a single event will be processed at a time. For example, this app processes events sequentially in the order of their arrival: ```python from dbos import DBOS DBOS.register_queue("in_order_queue", global_concurrency=1) @DBOS.step() def process_event(event: str): ... def event_endpoint(event: str): DBOS.enqueue_workflow("in_order_queue", process_event, event) ``` #### Rate Limiting You can set _rate limits_ for a queue, limiting the number of functions that it can start in a given period. Rate limits are global across all DBOS processes using this queue. For example, this queue has a limit of 50 with a period of 30 seconds, so it may not start more than 50 functions in 30 seconds: ```python DBOS.register_queue("example_queue", limiter={"limit": 50, "period": 30}) ``` Rate limits are especially useful when working with a rate-limited API, such as many LLM APIs. #### Reconfiguring Queues at Runtime Because queue configuration lives in the system database, you can change a queue's configuration at runtime without redeploying or restarting your workers. Use `DBOS.retrieve_queue` to fetch a queue, then call its `set_*` methods. Workers pick up the new configuration on their next polling iteration. Available mutators: `set_global_concurrency`, `set_worker_concurrency`, `set_limiter`, `set_partition_concurrency`, `set_partition_worker_concurrency`, `set_partition_limiter`, `set_polling_interval_sec`. Pass `None` to a setter to remove the limit (where applicable). ```python queue = DBOS.retrieve_queue("example_queue") # Double the queue's global concurrency. queue.set_global_concurrency(20) # Tighten its rate limit. queue.set_limiter({"limit": 25, "period": 30}) ``` You can also do this from a `DBOSClient`. ### Setting Timeouts You can set a timeout for an enqueued workflow with `SetWorkflowTimeout`. When the timeout expires, the workflow **and all its children** are cancelled. Cancelling a workflow sets its status to `CANCELLED` and preempts its execution at the beginning of its next step. Timeouts are **start-to-completion**: a workflow's timeout does not begin until the workflow is dequeued and starts execution. Also, timeouts are **durable**: they are stored in the database and persist across restarts, so workflows can have very long timeouts. Example syntax: ```python @DBOS.workflow() def example_workflow(): ... DBOS.register_queue("example-queue") # If the workflow does not complete within 10 seconds after being dequeued, it times out and is cancelled with SetWorkflowTimeout(10): DBOS.enqueue_workflow("example-queue", example_workflow) ``` ### Partitioning Queues You can **partition** queues to distribute work across dynamically created queue partitions. A queue is partitioned if you register it with any per-partition flow control limit: - `partition_concurrency`: Maximum workflows from any one partition running at once across all processes. - `partition_worker_concurrency`: Maximum workflows from any one partition running at once on a single process. - `partition_limiter`: Maximum workflows that may be started from any one partition in a given period. When you enqueue a workflow on a partitioned queue, you must supply a queue partition key. Essentially, you can think of each partition as a "subqueue" you dynamically create by enqueueing a workflow with a partition key. For example, suppose you want your users to each be able to run at most one task at a time. You can do this with a queue whose `partition_concurrency` is 1, where the partition key is user ID. **Example Syntax** ```python DBOS.register_queue("partitioned_queue", partition_concurrency=1) @DBOS.workflow() def process_task(task: Task): ... def on_user_task_submission(user_id: str, task: Task): # Partition the task queue by user ID. As the queue has a # per-partition concurrency of 1, this means that at most one # task can run at once per user (but tasks from different # users can run concurrently). with SetEnqueueOptions(queue_partition_key=user_id): DBOS.enqueue_workflow("partitioned_queue", process_task, task) ``` A partitioned queue enforces its per-partition limits **and** its queue-wide limits (`global_concurrency`, `worker_concurrency`, and `limiter`) at the same time. This lets you protect your workers from overload while still fairly distributing work between partitions. For example, this "fair queue" runs at most one task per user, but no more than 10 tasks on any single process: ```python DBOS.register_queue("fair_queue", partition_concurrency=1, worker_concurrency=10) ``` Each queue-wide limit has a per-partition counterpart, so you can mix and match them freely: ```python # At most 100 tasks running globally and 25 running per tenant, # at most 10 tasks running per process and 2 per tenant per process, # and at most 1000 tasks started per minute globally and 50 per tenant. DBOS.register_queue( "tenant_queue", global_concurrency=100, worker_concurrency=10, limiter={"limit": 1000, "period": 60}, partition_concurrency=25, partition_worker_concurrency=2, partition_limiter={"limit": 50, "period": 60}, ) ``` Each per-partition concurrency limit must be less than or equal to its queue-wide counterpart, and `partition_worker_concurrency` must be less than or equal to `partition_concurrency`. Deduplication is not supported on partitioned queues. ### Deduplication You can set a deduplication ID for an enqueued workflow with `SetEnqueueOptions`. At any given time, only one workflow with a specific deduplication ID can be enqueued in the specified queue. If a workflow with a deduplication ID is currently enqueued, delayed, or actively executing (status `ENQUEUED`, `DELAYED`, or `PENDING`), subsequent workflow enqueue attempt with the same deduplication ID in the same queue will raise a `DBOSQueueDeduplicatedError` exception. For example, this is useful if you only want to have one workflow active at a time per user—set the deduplication ID to the user's ID. Example syntax: ```python from dbos import DBOS, SetEnqueueOptions from dbos import error as dboserror DBOS.register_queue("example_queue") with SetEnqueueOptions(deduplication_id="my_dedup_id"): try: handle = DBOS.enqueue_workflow("example_queue", example_workflow, ...) except dboserror.DBOSQueueDeduplicatedError as e: # Handle deduplication error ... ``` ### Priority You can set a priority for an enqueued workflow with `SetEnqueueOptions`. Workflows with the same priority are dequeued in **FIFO (first in, first out)** order. Priority values can range from `1` to `2,147,483,647`, where **a low number indicates a higher priority**. Priority is enabled on every queue; no extra configuration is needed. :::tip Workflows without assigned priorities have the highest priority and are dequeued before workflows with assigned priorities. ::: Example syntax: ```python DBOS.register_queue("priority_queue") with SetEnqueueOptions(priority=10): # All workflows are enqueued with priority set to 10 # They will be dequeued in FIFO order for task in tasks: DBOS.enqueue_workflow("priority_queue", task_workflow, task) # first_workflow (priority=1) will be dequeued before all task_workflows (priority=10) with SetEnqueueOptions(priority=1): DBOS.enqueue_workflow("priority_queue", first_workflow) ``` ### Explicit Queue Listening By default, a process running DBOS listens to (dequeues workflows from) all queues owned by its application in its system database. However, sometimes you only want a process to listen to a specific list of queues. You can use `DBOS.listen_queues` to explicitly tell a process running DBOS to only listen to a specific set of queues. You must call `DBOS.listen_queues` before DBOS is launched. This is particularly useful when managing heterogeneous workers, where specific tasks should execute on specific physical servers. For example, say you have a mix of CPU workers and GPU workers and you want CPU tasks to only execute on CPU workers and GPU tasks to only execute on GPU workers. You can configure each type of worker to only listen to the appropriate queue: ```python if __name__ == "__main__": worker_type = ... # "cpu' or 'gpu' config: DBOSConfig = ... DBOS(config=config) if worker_type == "gpu": # GPU workers will only dequeue and execute workflows from the GPU queue DBOS.listen_queues(["gpu_queue"]) elif worker_type == "cpu": # CPU workers will only dequeue and execute workflows from the CPU queue DBOS.listen_queues(["cpu_queue"]) DBOS.launch() DBOS.register_queue("cpu_queue") DBOS.register_queue("gpu_queue") ``` Note that `DBOS.listen_queues` only controls what workflows are dequeued, not what workflows can be enqueued, so you can freely enqueue tasks onto the GPU queue from a CPU worker for execution on a GPU worker, and vice versa. ### Python Classes You can add DBOS decorators to your Python class instance methods. You can add step decorators to any class methods, but to add a workflow decorator to a class method, its class must inherit from `DBOSConfiguredInstance` and must be decorated with `@DBOS.dbos_class`. For example: ```python @DBOS.dbos_class() class URLFetcher(DBOSConfiguredInstance): def __init__(self, url: str): self.url = url super().__init__(config_name=url) @DBOS.workflow() def fetch_workflow(self): return self.fetch_url() @DBOS.step() def fetch_url(self): return requests.get(self.url).text example_fetcher = URLFetcher("https://example.com") print(example_fetcher.fetch_workflow()) ``` When you create a new instance of a DBOS class, `DBOSConfiguredInstance` must be instantiated with a `config_name`. This `config_name` should be a unique identifier of the instance. Additionally, all DBOS-decorated classes must be instantiated before `DBOS.launch()` is called. The reason for these requirements is to enable workflow recovery. When you create a new instance of a DBOS class, DBOS stores it in a global registry indexed by its class name and `config_name`. When DBOS needs to recover a workflow belonging to that class, it looks up the class instance using `config_name` so it can run the workflow using the right instance of its class. If DBOS classes are dynamically instantiated after `DBOS.launch()`, then DBOS may not find the class instance it needs to recover a workflow. #### Testing DBOS Functions Unless otherwise specified, recommend testing DBOS applications with pytest. You MUST use this fixture to reset DBOS between tests: ```python @pytest.fixture() def reset_dbos(): DBOS.destroy() config: DBOSConfig = { "name": "my-app", "application_version": "0.1.0", "system_database_url": os.environ.get("TESTING_DATABASE_URL"), } DBOS(config=config) DBOS.reset_system_database(truncate=True) DBOS.launch() ``` ### Upgrading Workflow Code One challenge you may encounter when operating long-running durable workflows in production is **how to deploy breaking changes without disrupting in-progress workflows.** A breaking change to a workflow is any change in what steps run or the order in which steps run. The issue is that if a breaking change was made to a workflow, the checkpoints created by a workflow that started on the previous version of the code may not match the steps called by the workflow in the new version of the code, which makes the workflow difficult to recover. DBOS supports two strategies for safely upgrading workflow code: **patching** and **versioning**. #### Patching When using patching, you use `DBOS.patch()` to make a breaking change in a conditional. `DBOS.patch()` returns `True` for new workflows (those started after the breaking change) and `False` for old workflows (those started before the breaking change). Therefore, if `DBOS.patch()` is `True`, call the new code, else, call the old code. To use patching, you must enable it in configuration: ```python config: DBOSConfig = { "name": "dbos-app", "application_version": "0.1.0", "system_database_url": os.environ.get("DBOS_SYSTEM_DATABASE_URL"), "enable_patching": True, } DBOS(config=config) ``` For example, let's say our workflow is: ```python @DBOS.workflow() def workflow(): foo() bar() ``` We want to replace the call to `foo()` with a call to `baz()`. This is a breaking change because it changes what steps run. We can make this breaking change safely using a patch: ```python @DBOS.workflow() def workflow(): if DBOS.patch("use-baz"): baz() else: foo() bar() ``` Now, new workflows will run `baz()`, while old workflows will safely continue through `foo()`. In coroutine workflows, you MUST use `await DBOS.patch_async()` and `await DBOS.deprecate_patch_async()` instead, as `DBOS.patch()` and `DBOS.deprecate_patch()` raise an error when called from a running event loop. ##### Deprecating and Removing Patches Patches don't need to stay in your code forever. Once all workflows that started before you deployed the patch are complete, you can safely remove patches from your code. You can use the list workflows APIs to see what workflows are still active. First, you must deprecate the patch with `DBOS.deprecate_patch()`. This safely runs all workflows that contain the patch marker, but does not insert the patch marker into new workflows. For example, here's how to deprecate the patch above: ```python @DBOS.workflow() def workflow(): DBOS.deprecate_patch("use-baz") baz() bar() ``` Then, when all workflows that started before you deprecated the patch are complete, you can remove the patch entirely: ```python @DBOS.workflow() def workflow(): baz() bar() ``` If any mistakes happen during the process (a breaking change is not patched, or a patch is deprecated or removed prematurely), the workflow will throw a `DBOSUnexpectedStepError` error clearly pointing to the step where the problem occurred. #### Versioning When using versioning, DBOS **versions** applications and workflows. All workflows are tagged with the application version on which they started. By default, application version is automatically computed from a hash of workflow source code. However, you can set your own version through configuration. ```python config: DBOSConfig = { "name": "dbos-app", "system_database_url": os.environ.get("DBOS_SYSTEM_DATABASE_URL"), "application_version": "1.0.0", } DBOS(config=config) ``` When DBOS tries to recover workflows, it only recovers workflows whose version matches the current application version. This prevents unsafe recovery of workflows that depend on different code. When using versioning, we recommend **blue-green** code upgrades. When deploying a new version of your code, launch new processes running your new code version, but retain some processes running your old code version. Direct new traffic to your new processes while your old processes "drain" and complete all workflows of the old code version. Then, once all workflows of the old version are complete (you can use `DBOS.list_workflows` to check), you can retire the old code version. ### Workflow Handle DBOS.start_workflow, DBOS.retrieve_workflow, and enqueue return workflow handles. ##### get_workflow_id ```python handle.get_workflow_id() -> str ``` Retrieve the ID of the workflow. ##### get_result ```python handle.get_result( *, polling_interval_sec: float = 1.0, ) -> R ``` Wait for the workflow to complete, then return its result. **Parameters:** - **polling_interval_sec**: The interval at which DBOS polls the database for the workflow's result. Only used for enqueued workflows or retrieved handles. ##### get_status ```python handle.get_status() -> WorkflowStatus ``` ### Workflow Management Methods #### list_workflows ```python def list_workflows( *, workflow_ids: Optional[List[str]] = None, status: Optional[Union[str, List[str]]] = None, start_time: Optional[str] = None, end_time: Optional[str] = None, name: Optional[Union[str, List[str]]] = None, app_version: Optional[Union[str, List[str]]] = None, forked_from: Optional[Union[str, List[str]]] = None, user: Optional[Union[str, List[str]]] = None, queue_name: Optional[Union[str, List[str]]] = None, limit: Optional[int] = None, offset: Optional[int] = None, sort_desc: bool = False, workflow_id_prefix: Optional[Union[str, List[str]]] = None, load_input: bool = True, load_output: bool = True, executor_id: Optional[Union[str, List[str]]] = None, queues_only: bool = False, has_parent: Optional[bool] = None, attributes: Optional[Dict[str, Any]] = None, schedule_name: Optional[Union[str, List[str]]] = None, ) -> List[WorkflowStatus]: ``` Retrieve a list of `WorkflowStatus` of all workflows matching specified criteria. **Parameters:** - **workflow_ids**: Retrieve workflows with these IDs. - **status**: Retrieve workflows with this status (or one of these statuses) (Must be `ENQUEUED`, `DELAYED`, `PENDING`, `SUCCESS`, `ERROR`, `CANCELLED`, or `MAX_RECOVERY_ATTEMPTS_EXCEEDED`) - **start_time**: Retrieve workflows started after this (RFC 3339-compliant) timestamp. - **end_time**: Retrieve workflows started before this (RFC 3339-compliant) timestamp. - **name**: Retrieve workflows with this name. - **app_version**: Retrieve workflows tagged with this application version. - **forked_from**: Retrieve workflows forked from this workflow ID. - **user**: Retrieve workflows run by this authenticated user. - **queue_name**: Retrieve workflows that were enqueued on this queue. - **limit**: Retrieve up to this many workflows. - **offset**: Skip this many workflows from the results returned (for pagination). - **sort_desc**: Whether to sort the results in descending (`True`) or ascending (`False`) order by workflow start time. - **workflow_id_prefix**: Retrieve workflows whose IDs start with the specified string. - **load_input**: Whether to load and deserialize workflow inputs. Set to `False` to improve performance when inputs are not needed. - **load_output**: Whether to load and deserialize workflow outputs. Set to `False` to improve performance when outputs are not needed. - **executor_id**: Retrieve workflows with this executor ID. - **queues_only**: If `True`, only retrieve workflows that are currently queued (status `DELAYED`, `ENQUEUED`, or `PENDING` and `queue_name` not null). - **has_parent**: If `True`, only retrieve workflows that have a parent workflow. If `False`, only retrieve workflows without a parent. - **attributes**: Retrieve workflows whose custom attributes contain all the given key-value pairs. Only supported when using a Postgres system database. - **schedule_name**: Retrieve workflows that were enqueued by this schedule (or one of these schedules). #### list_queued_workflows ```python def list_queued_workflows( *, workflow_ids: Optional[List[str]] = None, status: Optional[Union[str, List[str]]] = None, start_time: Optional[str] = None, end_time: Optional[str] = None, name: Optional[Union[str, List[str]]] = None, app_version: Optional[Union[str, List[str]]] = None, forked_from: Optional[Union[str, List[str]]] = None, user: Optional[Union[str, List[str]]] = None, queue_name: Optional[Union[str, List[str]]] = None, limit: Optional[int] = None, offset: Optional[int] = None, sort_desc: bool = False, workflow_id_prefix: Optional[Union[str, List[str]]] = None, load_input: bool = True, load_output: bool = True, executor_id: Optional[Union[str, List[str]]] = None, has_parent: Optional[bool] = None, attributes: Optional[Dict[str, Any]] = None, ) -> List[WorkflowStatus]: ``` Retrieve a list of `WorkflowStatus` of all **queued** workflows (status `DELAYED`, `ENQUEUED`, or `PENDING` and `queue_name` not null) matching specified criteria. **Parameters:** - **workflow_ids**: Retrieve workflows with these IDs. - **status**: Retrieve workflows with this status (or one of these statuses) (Must be `DELAYED`, `ENQUEUED`, or `PENDING`) - **start_time**: Retrieve workflows enqueued after this (RFC 3339-compliant) timestamp. - **end_time**: Retrieve workflows enqueued before this (RFC 3339-compliant) timestamp. - **name**: Retrieve workflows with this name. - **app_version**: Retrieve workflows tagged with this application version. - **forked_from**: Retrieve workflows forked from this workflow ID. - **user**: Retrieve workflows run by this authenticated user. - **queue_name**: Retrieve workflows running on this queue. - **limit**: Retrieve up to this many workflows. - **offset**: Skip this many workflows from the results returned (for pagination). - **sort_desc**: Whether to sort the results in descending (`True`) or ascending (`False`) order by workflow start time. - **workflow_id_prefix**: Retrieve workflows whose IDs start with the specified string. - **load_input**: Whether to load and deserialize workflow inputs. Set to `False` to improve performance when inputs are not needed. - **load_output**: Whether to load and deserialize workflow outputs. Set to `False` to improve performance when outputs are not needed. - **executor_id**: Retrieve workflows with this executor ID. - **has_parent**: If `True`, only retrieve workflows that have a parent workflow. If `False`, only retrieve workflows without a parent. - **attributes**: Retrieve workflows whose custom attributes contain all the given key-value pairs. Only supported when using a Postgres system database. #### list_workflow_steps ```python def list_workflow_steps( workflow_id: str, *, load_output: bool = True, limit: Optional[int] = None, offset: Optional[int] = None, ) -> List[StepInfo] ``` Retrieve the steps of a workflow. Steps are ordered by `function_id`. Use `limit` and `offset` to paginate results. Set `load_output` to `False` to improve performance when step outputs and errors are not needed; the `output` and `error` fields are then always `None`. This is a list of `StepInfo` objects, with the following structure: ```python class StepInfo(TypedDict): # The unique ID of the step in the workflow. One-indexed. function_id: int # The name of the step function_name: str # The step's output, if any output: Optional[Any] # The error the step threw, if any error: Optional[Exception] # If the step starts or retrieves the result of a workflow, its ID child_workflow_id: Optional[str] # The Unix epoch timestamp at which this step started started_at_epoch_ms: Optional[int] # The Unix epoch timestamp at which this step completed completed_at_epoch_ms: Optional[int] ``` #### set_workflow_delay ```python DBOS.set_workflow_delay( workflow_id: str, *, delay_seconds: Optional[float] = None, delay_until_epoch_ms: Optional[int] = None, ) -> None ``` Set or update the delay on a workflow. Only affects workflows with `DELAYED` status. Provide exactly one of `delay_seconds` (relative) or `delay_until_epoch_ms` (absolute). #### cancel_workflow ```python DBOS.cancel_workflow( workflow_id: str, *, cancel_children: bool = False, ) -> None ``` Cancel a workflow. This sets its status to `CANCELLED`, removes it from its queue (if it is enqueued) and preempts its execution (interrupting it at the beginning of its next step) If `cancel_children` is `True`, also recursively cancels all child workflows started by this workflow. #### resume_workflow ```python DBOS.resume_workflow( workflow_id: str, *, queue_name: Optional[str] = None, ) -> WorkflowHandle[Any] ``` Resume a workflow. This immediately starts it from its last completed step. You can use this to resume workflows that are cancelled or have exceeded their maximum recovery attempts. You can also use this to start an enqueued workflow immediately, bypassing its queue. If `queue_name` is provided, the resumed workflow is enqueued on the specified queue instead of starting immediately. #### fork_workflow ```python DBOS.fork_workflow( workflow_id: str, start_step: int, *, application_version: Optional[str] = None, queue_name: Optional[str] = None, queue_partition_key: Optional[str] = None, ) -> WorkflowHandle[Any] ``` Start a new execution of a workflow from a specific step. The input step ID must match the `function_id` of the step returned by `list_workflow_steps`. The specified `start_step` is the step from which the new workflow will start, so any steps whose ID is less than `start_step` will not be re-executed. The forked workflow will have a new workflow ID, which can be set with `SetWorkflowID`. It is possible to specify the application version on which the forked workflow will run by setting `application_version`, this is useful for "patching" workflows that failed due to a bug in a previous application version. If `queue_name` is provided, the forked workflow is enqueued on the specified queue instead of starting immediately. If the queue is partitioned, you can also specify `queue_partition_key`. #### rewind_workflow ```python DBOS.rewind_workflow( workflow_id: str, *, start_step: Optional[int] = None, application_version: Optional[str] = None, queue_name: Optional[str] = None, queue_partition_key: Optional[str] = None, ) -> WorkflowHandle[Any] ``` Rewind a workflow to a specific step and re-execute it from that step, keeping its workflow ID (unlike `fork_workflow`, which creates a new workflow). Steps with IDs greater than or equal to `start_step` are discarded and re-executed; if `start_step` is not provided, the whole workflow re-executes. Only a workflow in a terminal state (`SUCCESS`, `ERROR`, `CANCELLED`, or `MAX_RECOVERY_ATTEMPTS_EXCEEDED`) can be rewound; cancel a running workflow first. Rewinding clears the workflow's output, rolls back events it set at or after `start_step`, deletes messages it received at or after `start_step`, and deletes the checkpoints of datasource transactions at or after `start_step`. `application_version`, `queue_name`, and `queue_partition_key` behave as in `fork_workflow`. #### Workflow Status Some workflow introspection and management methods return a `WorkflowStatus`. This object has the following definition: ```python class WorkflowStatus: # The workflow ID workflow_id: str # The workflow status. Must be one of ENQUEUED, DELAYED, PENDING, SUCCESS, ERROR, CANCELLED, or MAX_RECOVERY_ATTEMPTS_EXCEEDED status: str # The name of the workflow function name: str # The name of the workflow's class, if any class_name: Optional[str] # The name with which the workflow's class instance was configured, if any config_name: Optional[str] # The user who ran the workflow, if specified authenticated_user: Optional[str] # The role with which the workflow ran, if specified assumed_role: Optional[str] # All roles which the authenticated user could assume authenticated_roles: Optional[list[str]] # The deserialized workflow input object input: Optional[WorkflowInputs] # The workflow's output, if any output: Optional[Any] # The error the workflow threw, if any error: Optional[Exception] # Workflow start time, as a Unix epoch timestamp in ms created_at: Optional[int] # Last time the workflow status was updated, as a Unix epoch timestamp in ms updated_at: Optional[int] # If this workflow was enqueued, on which queue queue_name: Optional[str] # The executor to most recently execute this workflow executor_id: Optional[str] # The application version on which this workflow was started app_version: Optional[str] # The start-to-close timeout of the workflow in ms workflow_timeout_ms: Optional[int] # The deadline of a workflow, computed by adding its timeout to its start time. workflow_deadline_epoch_ms: Optional[int] # Unique ID for deduplication on a queue deduplication_id: Optional[str] # Priority of the workflow on the queue, starting from 1 ~ 2,147,483,647. Default 0 (highest priority). priority: Optional[int] # If this workflow is enqueued on a partitioned queue, its partition key queue_partition_key: Optional[str] # If this workflow was forked from another, that workflow's ID. forked_from: Optional[str] # Whether this workflow has ever been forked from by another workflow. was_forked_from: bool # If this workflow was started as a child of another workflow, that workflow's ID. parent_workflow_id: Optional[str] # The UNIX epoch timestamp at which the workflow was last dequeued, if it had been enqueued dequeued_at: Optional[int] # The UNIX epoch timestamp before which the workflow should not be dequeued delay_until_epoch_ms: Optional[int] # The UNIX epoch timestamp at which the workflow completed (SUCCESS, ERROR, or CANCELLED), if it has completed_at: Optional[int] # Custom key-value attributes attached to the workflow attributes: Optional[Dict[str, Any]] # If this workflow was enqueued by a schedule, that schedule's name schedule_name: Optional[str] # The application that owns this workflow application_name: Optional[str] ``` #### Configuring DBOS To configure DBOS, pass a `DBOSConfig` object to its constructor. For example: ```python config: DBOSConfig = { "name": "dbos-example", "application_version": "0.1.0", "system_database_url": os.environ["DBOS_SYSTEM_DATABASE_URL"], } DBOS(config=config) ``` The `DBOSConfig` object has the following fields. All fields except `name` are optional. ```python class DBOSConfig(TypedDict): name: str enable_patching: Optional[bool] application_version: Optional[str] executor_id: Optional[str] system_database_url: Optional[str] sys_db_pool_size: Optional[int] sys_db_polling_concurrency: Optional[int] dbos_system_schema: Optional[str] system_database_engine: Optional[sqlalchemy.Engine] use_listen_notify: Optional[bool] observability_query_timeout_sec: Optional[float] sys_db_idle_transaction_timeout_sec: Optional[float] conductor_key: Optional[str] conductor_url: Optional[str] enable_otlp: Optional[bool] otlp_traces_endpoints: Optional[List[str]] otlp_logs_endpoints: Optional[List[str]] otlp_attributes: Optional[dict[str, str]] log_level: Optional[str] serializer: Optional[Serializer] ``` - **name**: Your application's name. It must be between 3 and 256 characters long and contain only lowercase letters, numbers, dashes, and underscores. - **enable_patching** Enable the patching strategy for safely upgrading workflow code. - **application_version**: If using the versioning strategy for safely upgrading workflow code, the code version for this application and its workflows. - **executor_id**: A unique process ID used to identify the application instance in distributed environments. If using DBOS Conductor or Cloud, this is set automatically. - **system_database_url**: A connection string to your system database. This is the database in which DBOS stores workflow and step state. This may be either Postgres or SQLite, though Postgres is recommended for production. DBOS uses this connection string to create a SQLAlchemy Engine (for Postgres, DBOS always uses the psycopg driver). A valid connection string looks like: ``` postgresql://[username]:[password]@[hostname]:[port]/[database name] ``` Or with SQLite: ``` sqlite:///[path to database file] ``` :::info Passwords in connection strings must be escaped (for example with urllib) if they contain special characters. ::: If no connection string is provided, DBOS uses a SQLite database (with any dashes in the application name replaced by underscores): ```shell sqlite:///[application_name].sqlite ``` - **sys_db_pool_size**: The size of the connection pool used for the DBOS system database. Defaults to 20. - **sys_db_polling_concurrency**: The maximum number of database-backed polling reads from wait operations (such as `get_result`, `recv`, `get_event`, and `read_stream`) that may run concurrently against the system database pool. This prevents high-fan-out polling from checking out every connection in the pool and starving control-plane operations (such as enqueue/dequeue, status writes, recovery, and cancellation). Defaults to half the `sys_db_pool_size` (minimum 1). Set to a non-positive value to disable the limit. - **dbos_system_schema**: Postgres schema name for DBOS system tables. Defaults to "dbos". - **system_database_engine**: A custom SQLAlchemy engine to use to connect to your system database. If provided, DBOS will not create an engine but use this instead. - **use_listen_notify**: Whether to use PostgreSQL LISTEN/NOTIFY (`True`) or polling (`False`) to await notifications and events. Defaults to `True`. Ignored in SQLite, which always uses polling. - **observability_query_timeout_sec**: The statement timeout, in seconds, applied to observability queries (such as listing workflows, queued workflows, and workflow steps) on a Postgres system database, so a slow query on a large database does not hold resources indefinitely. A query that exceeds the timeout raises `DBOSQueryTimeoutError`. Defaults to 30 seconds. Set to zero or a negative value to disable the timeout. - **sys_db_idle_transaction_timeout_sec**: The Postgres `idle_in_transaction_session_timeout`, in seconds, set on the system database connections DBOS creates. Defaults to 60 seconds. - **conductor_key**: An API key for DBOS Conductor. If provided, application is connected to Conductor. API keys can be created from the DBOS Console. - **conductor_url**: The URL of the Conductor service to connect to. Only set if you are self-hosting Conductor. - **enable_otlp**: Enable DBOS OpenTelemetry tracing and export. Defaults to False. - **otlp_traces_endpoints**: DBOS operations automatically generate OpenTelemetry Traces. Use this field to declare a list of OTLP-compatible trace receivers. Requires `enable_otlp` to be True. - **otlp_logs_endpoints**: the DBOS logger can export OTLP-formatted log signals. Use this field to declare a list of OTLP-compatible log receivers. Requires `enable_otlp` to be True. - **otlp_attributes**: A set of attributes (key-value pairs) to apply to all OTLP-exported logs and traces. - **log_level**: Configure the DBOS logger severity. Defaults to `INFO`. - **serializer**: A custom serializer for the system database. ##### Custom Serialization DBOS must serialize data such as workflow inputs and outputs and step outputs to store it in the system database. By default, data is serialized with `pickle` then Base64-encoded, but you can optionally supply a custom serializer through DBOS configuration. A custom serializer must match this interface: ```python class Serializer(ABC): @abstractmethod def serialize(self, data: Any) -> str: pass @abstractmethod def deserialize(self, serialized_data: str) -> Any: pass def name(self) -> str: return "custom_serializer" ``` For example, here is how to configure DBOS to use a JSON serializer: ```python import json import os from typing import Any from dbos import DBOS, DBOSConfig, Serializer class JsonSerializer(Serializer): def serialize(self, data: Any) -> str: return json.dumps(data) def deserialize(self, serialized_data: str) -> Any: return json.loads(serialized_data) def name(self) -> str: return "basic_json" serializer = JsonSerializer() config: DBOSConfig = { "name": "dbos-starter", "application_version": "0.1.0", "system_database_url": os.environ.get("DBOS_SYSTEM_DATABASE_URL"), "serializer": serializer } DBOS(config=config) DBOS.launch() ``` The serializer's `name` is stored with serialized values and used to ensure that the correct deserializer is used. #### Datasources Datasources are the recommended way to durably perform database operations inside workflows. A datasource wraps a SQLAlchemy engine so that each database transaction inside a workflow runs exactly once, even if the workflow is interrupted and retried. ONLY use datasource transactions if you are SPECIFICALLY requested to perform database operations, DO NOT USE THEM OTHERWISE. If asked to add DBOS to code that already contains database operations, ALWAYS make it a step, do NOT attempt to make it a datasource transaction unless requested. ONLY use datasource transactions with a PostgreSQL or SQLite database. To access any other database, ALWAYS use steps. Create a datasource with the `create` factory method, passing your database URL. Use `SQLAlchemyDatasource` for synchronous code and `AsyncSQLAlchemyDatasource` for async code. Annotate a function with `@ds.transaction()` to run it as a tracked database transaction, and access the current SQLAlchemy session with `ds.sql_session()`. ##### Synchronous datasource ```python import os from typing import Optional from dbos import DBOS, SQLAlchemyDatasource from sqlalchemy import text ds = SQLAlchemyDatasource.create(os.environ["APP_DATABASE_URL"]) @ds.transaction() def example_insert(name: str, note: str) -> None: # Insert a new greeting into the database session = ds.sql_session() # sqlalchemy.orm.Session session.execute( text("INSERT INTO greetings (name, note) VALUES (:name, :note)"), {"name": name, "note": note}, ) @ds.transaction() def example_select(name: str) -> Optional[str]: # Select the first greeting to a particular name session = ds.sql_session() row = session.execute( text("SELECT note FROM greetings WHERE name = :name LIMIT 1"), {"name": name}, ).first() return row[0] if row else None @DBOS.workflow() def greeting_workflow(name: str, note: str) -> None: example_insert(name, note) ``` ##### Asynchronous datasource `AsyncSQLAlchemyDatasource.create` is a coroutine and it ONLY accepts `async def` transaction functions. Because Python does NOT allow `await` at module scope, create the datasource with `asyncio.run` so it is a module-level global that the `@ads.transaction()` decorators can use. Only use `await ...create(...)` if you are already inside a coroutine. To use `AsyncSQLAlchemyDatasource` with SQLite, you MUST use an async driver URL such as `sqlite+aiosqlite:///app.sqlite` (install the driver with `pip install "dbos[aiosqlite]"`); a plain `sqlite:///` URL raises an error. ```python import asyncio import os from dbos import DBOS, AsyncSQLAlchemyDatasource from sqlalchemy import text ads = asyncio.run(AsyncSQLAlchemyDatasource.create(os.environ["APP_DATABASE_URL"])) @ads.transaction() async def example_insert(name: str, note: str) -> None: # Insert a new greeting into the database session = ads.sql_session() # sqlalchemy.ext.asyncio.AsyncSession await session.execute( text("INSERT INTO greetings (name, note) VALUES (:name, :note)"), {"name": name, "note": note}, ) @DBOS.workflow() async def greeting_workflow(name: str, note: str) -> None: await example_insert(name, note) ``` `SQLAlchemyDatasource` ONLY supports synchronous (non-`async def`) functions and `AsyncSQLAlchemyDatasource` ONLY supports `async def` functions. Decorating the wrong function type raises a `DBOSException`. ````
--- ## DBOS CLI(Reference) ### Workflow Management Commands These commands all require the URL of your DBOS system database. You can supply this URL through the `--sys-db-url` argument or through a [`dbos-config.yaml` configuration file](./configuration.md#dbos-configuration-file). #### dbos workflow list **Description:** List workflows run by your application in JSON format ordered by recency (most recently started workflows last). **Arguments:** - `-s, --sys-db-url URL`: Your DBOS system database URL. * `-l, --limit INTEGER`: Limit the results returned [default: 10] * `-u, --user TEXT`: Retrieve workflows run by this user * `-t, --start-time TEXT`: Retrieve workflows starting after this timestamp (ISO 8601 format) * `-e, --end-time TEXT`: Retrieve workflows starting before this timestamp (ISO 8601 format) * `-S, --status TEXT`: Retrieve workflows with this status (PENDING, SUCCESS, ERROR, MAX_RECOVERY_ATTEMPTS_EXCEEDED, ENQUEUED, DELAYED, or CANCELLED) * `-v, --application-version TEXT`: Retrieve workflows with this application version * `-n, --name TEXT`: Retrieve workflows with this name * `-a, --application-name TEXT`: Retrieve workflows owned by this application (workflows owned by no application are always included) * `-d, --sort-desc`: Sort the results in descending order (newest first) * `-o, --offset INTEGER`: Offset for pagination * `--schema TEXT`: The schema name for the DBOS system tables. Defaults to `dbos`. **Output:** A JSON-formatted list of [workflow statuses](./contexts#workflow-status). #### dbos workflow get **Description:** Retrieve information on a workflow run by your application. **Arguments:** - ``: The ID of the workflow to retrieve - `-s, --sys-db-url URL`: Your DBOS system database URL. - `--schema TEXT`: The schema name for the DBOS system tables. Defaults to `dbos`. **Output:** A JSON-formatted [workflow status](./contexts#workflow-status). #### dbos workflow steps **Arguments:** - ``: The ID of the workflow to retrieve - `-s, --sys-db-url URL`: Your DBOS system database URL. - `--schema TEXT`: The schema name for the DBOS system tables. Defaults to `dbos`. **Output:** A JSON-formatted list of [workflow steps](./contexts#list_workflow_steps). #### dbos workflow cancel **Description:** Cancel a workflow so it is no longer automatically retried or restarted. If the workflow is executing, it is interrupted at the beginning of its next step. **Arguments:** - ``: The ID of the workflow to cancel - `-s, --sys-db-url URL`: Your DBOS system database URL. - `--schema TEXT`: The schema name for the DBOS system tables. Defaults to `dbos`. #### dbos workflow resume **Description:** Resume a workflow from its last completed step. You can use this to resume workflows that are cancelled or that have exceeded their maximum recovery attempts. You can also use this to start an `ENQUEUED` workflow, bypassing its queue. **Arguments:** - ``: The ID of the workflow to resume. - `-s, --sys-db-url URL`: Your DBOS system database URL. - `--schema TEXT`: The schema name for the DBOS system tables. Defaults to `dbos`. #### dbos workflow fork **Description:** Fork a new execution of a workflow, starting at a given step. This new workflow has a new workflow ID. Unless you pin it with `--application-version`, it is not tagged with any application version, so it is dequeued by an executor running the latest application version. Forking from step N copies the results of all previous steps to the new workflow, which then starts running from step N. **Arguments:** * ``: The ID of the workflow to fork. - `-s, --sys-db-url URL`: Your DBOS system database URL. * `-f, --forked-workflow-id`: Custom ID for the forked workflow * `-v, --application-version`: Custom application version for the forked workflow * `-S, --step INTEGER`: Restart from this step [default: 1] - `--schema TEXT`: The schema name for the DBOS system tables. Defaults to `dbos`. **Output:** A JSON-formatted [workflow status](./contexts#workflow-status). #### dbos workflow queue list **Description:** Lists all currently enqueued tasks in JSON format ordered by recency (most recently enqueued functions last). **Arguments:** - `-s, --sys-db-url URL`: Your DBOS system database URL. * `-l, --limit INTEGER`: Limit the results returned * `-t, --start-time TEXT`: Retrieve functions starting after this timestamp (ISO 8601 format) * `-e, --end-time TEXT`: Retrieve functions starting before this timestamp (ISO 8601 format) * `-S, --status TEXT`: Retrieve functions with this status (PENDING, SUCCESS, ERROR, MAX_RECOVERY_ATTEMPTS_EXCEEDED, ENQUEUED, DELAYED, or CANCELLED) * `-q, --queue-name TEXT`: Retrieve functions on this queue * `-n, --name TEXT`: Retrieve functions with this name * `-a, --application-name TEXT`: Retrieve functions owned by this application (functions owned by no application are always included) * `-d, --sort-desc`: Sort the results in descending order (newest first) * `-o, --offset INTEGER`: Offset for pagination * `--schema TEXT`: The schema name for the DBOS system tables. Defaults to `dbos`. **Output:** A JSON-formatted list of [workflow statuses](./contexts#workflow-status). ### Application Management Commands #### dbos migrate Create the DBOS system database and internal tables. By default, a DBOS application automatically creates these on startup. However, in production environments, a DBOS application may not run with sufficient privilege to create databases or tables. In that case, this command can be run with a privileged user to create all DBOS database tables. After creating the DBOS database tables with this command, a DBOS application can run with minimum permissions, requiring only access to the DBOS schema in the system database. Use the `-r` flag to grant a role access to that schema. Such an application should also be configured with [`run_migrations=False`](./configuration.md#database-connection-settings), so it never attempts to alter the schema and instead verifies at launch that this command has brought the system database up to date. You can also run these migrations from code with [`DBOS.migrate`](./dbos-class.md#migrate). This command does not migrate the tables used by [datasources](./datasources.md); use [`SQLAlchemyDatasource.migrate`](./datasources.md#sqlalchemydatasourcemigrate) for those. **Arguments:** - `-s, --sys-db-url URL`: A connection string for your DBOS [system database](../../explanations/system-tables.md), in which DBOS stores its internal state. This command will create that database if it does not exist and create or update the DBOS system tables within it. - `-r, --app-role`: The role with which you will run your DBOS app. This role is granted the minimum permissions needed to access the DBOS schema in your system database. - `--schema TEXT`: The schema name for the DBOS system tables. Defaults to `dbos`. - `--print-migrations [all|NUMBER]`: Instead of running the migrations, print their SQL to standard output, either all of them (for a fresh database) or starting from a migration number (to upgrade an existing database). Postgres only. - `--print-user-role`: Instead of executing them, print the SQL statements granting `--app-role` access to the DBOS system tables. Use these last two flags to emit SQL you can apply yourself, for example if your database is managed by a DBA. The output is only SQL and comments, but it contains `CREATE INDEX CONCURRENTLY` and `DROP INDEX CONCURRENTLY`, so it must run outside a transaction block. ```shell dbos migrate --print-migrations all -s ${DBOS_SYSTEM_DATABASE_URL} > migrations.sql dbos migrate --print-user-role -r my_app_role -s ${DBOS_SYSTEM_DATABASE_URL} > grants.sql ``` #### dbos start Start your DBOS application by executing the `start` command defined in [`dbos-config.yaml`](./configuration.md#dbos-configuration-file). For example: ```yaml runtimeConfig: start: - "fastapi run" ``` DBOS Cloud executes this command to start your app. #### dbos init Initialize the local directory with a DBOS template application. **Arguments:** - ``: The name of your application. If not specified, will be prompted for (the `dbos-toolbox` and `dbos-app-starter` templates instead default to the template name). - `-t, --template TEXT`: Specify a template to use. ("dbos-toolbox", "dbos-app-starter", "dbos-db-starter") - `--config, -c`: If this flag is set, only the `dbos-config.yaml` file is added from the template. Useful to add DBOS to an existing project. #### dbos reset Reset your DBOS [system database](../../explanations/system-tables.md), deleting metadata about past workflows and steps. This drops the entire system database (for SQLite, it deletes the database file), including any application data stored in that database; data in other databases is not affected. **Arguments:** * `--yes, -y`: Skip confirmation prompt. - `-s, --sys-db-url URL`: Your DBOS system database URL. #### dbos rename-application After renaming an application, transfer ownership of everything in the system database (workflows, steps, queues, schedules, and application versions) from its old name to its new name. Equivalent to [`DBOSClient.rename_application`](./client.md#rename_application); see there for details. Prints the number of rows transferred, by table. **Stop the application being renamed before running this.** **Arguments:** - `-s, --sys-db-url URL`: Your DBOS system database URL. - `-f, --from TEXT`: The application's previous name. Omit to only adopt rows owned by no application (requires `--adopt-unclaimed-rows`). - `-t, --to TEXT`: The application that ends up owning the rows. Required. - `--adopt-unclaimed-rows`: Also transfer rows owned by no application. - `--batch-size INTEGER`: The number of workflows per batch when transferring completed workflows and steps [default: 10000] - `--schema TEXT`: The schema name for the DBOS system tables. Defaults to `dbos`. - `-y, --yes`: Skip confirmation prompt. --- ## DBOS Client(Reference) `DBOSClient` provides a programmatic way to interact with your DBOS application from external code or from another DBOS application. `DBOSClient` includes methods similar to [`DBOS`](./contexts.md) that can be used outside of a DBOS application, such as [`enqueue`](./queues.md#enqueue) or [`get_event`](./contexts.md#get_event). :::note `DBOSClient` is included in the `dbos` package, the same package used by DBOS applications. Where DBOS applications use the [`DBOS` methods](./contexts.md), external applications use `DBOSClient` instead. ::: #### Constructor ```python DBOSClient( *, system_database_url: Optional[str] = None, system_database_engine: Optional[sa.Engine] = None, dbos_system_schema: Optional[str] = "dbos", serializer: Serializer = DefaultSerializer(), system_database_pool_size: Optional[int] = None, system_database_polling_concurrency: Optional[int] = None, use_listen_notify: bool = False, application_name: Optional[str] = None, lazy: bool = False, retry_connection_errors: bool = True, observability_query_timeout_sec: Optional[float] = None, sys_db_idle_transaction_timeout_sec: Optional[float] = None, ) ``` **Parameters:** - `system_database_url`: A connection string to your DBOS system database, with the same format as in [DBOSConfig](./configuration.md). Required unless `system_database_engine` is provided (raises `DBOSException` if neither is set). - `system_database_engine`: A custom SQLAlchemy engine to use to connect to your system database. If provided, the client will not create an engine but use this instead. - `dbos_system_schema`: Postgres schema name for DBOS system tables. Defaults to "dbos". - `serializer`: A custom [serializer](./contexts.md#custom-serialization) for workflow inputs and outputs. Must match the serializer used by the DBOS application. - `system_database_pool_size`: The maximum size of the client's system database connection pool. Defaults to 5. - `system_database_polling_concurrency`: The maximum number of concurrent database-backed polling reads from wait operations. See [`sys_db_polling_concurrency`](./configuration.md#database-connection-settings) in the configuration reference. Defaults to half the pool size (minimum 1). - `use_listen_notify`: Whether the client runs a background listener thread that uses PostgreSQL LISTEN, so wait operations such as [`get_event`](#get_event) and [`read_stream`](#read_stream) are woken by notifications instead of polling the database. Defaults to `False`. Only enable this if your DBOS application's system database is Postgres and was created with [`use_listen_notify`](./configuration.md#database-connection-settings) enabled (the Postgres default). - `lazy`: Whether to defer connecting to the system database until the client is first used. Defaults to `False`, meaning the connection is checked on construction and the constructor raises if the system database is unreachable. If `True`, the constructor does not connect; use [`check_connection`](#check_connection) to check the connection explicitly. Cannot be combined with `use_listen_notify`, whose listener thread connects immediately (raises `DBOSException` if both are set). - `retry_connection_errors`: Whether a client operation that loses its database connection blocks and retries until the connection recovers. Defaults to `True`. Set to `False` to raise connection errors instead, so an unreachable database surfaces as an error rather than a wait. - `application_name`: The application on whose behalf this client acts. Workflows the client enqueues and queues and schedules it registers are owned by that application, and the client's listing operations default to that application's rows. Always set this if multiple applications share a system database. - `observability_query_timeout_sec`: The statement timeout, in seconds, applied to the client's observability queries (such as listing workflows, queued workflows, and workflow steps) on a Postgres system database. A query that exceeds the timeout raises `DBOSQueryTimeoutError`. Defaults to 30 seconds. Set to zero or a negative value to disable the timeout. See [`observability_query_timeout_sec`](./configuration.md#database-connection-settings) in the configuration reference. - `sys_db_idle_transaction_timeout_sec`: The Postgres `idle_in_transaction_session_timeout`, in seconds, set on the system database connections the client creates. Defaults to 60 seconds. **Example syntax:** This DBOS client connects to the system database specified in the `DBOS_SYSTEM_DATABASE_URL` environment variable. ```python import os from dbos import DBOSClient client = DBOSClient(system_database_url=os.environ["DBOS_SYSTEM_DATABASE_URL"]) ``` #### check_connection ```python client.check_connection() -> None ``` Verify that the client can reach the system database, raising an exception if it cannot. Useful for checking the connection of a client constructed with [`lazy=True`](#constructor), or for verifying at any time that the connection is still healthy. #### check_connection_async ```python await client.check_connection_async() -> None ``` Asynchronous version of [`check_connection`](#check_connection). #### destroy ```python client.destroy() -> None ``` Clean up database connections and release resources. Call this method when you are done using the client. ### Workflow Interaction Methods #### enqueue ```python class EnqueueOptions(TypedDict): workflow_name: str queue_name: str workflow_id: NotRequired[str] workflow_id_reuse_policy: NotRequired[WorkflowIDReusePolicy] app_version: NotRequired[str] workflow_timeout: NotRequired[float] deduplication_id: NotRequired[str] duplication_policy: NotRequired[DuplicationPolicy] priority: NotRequired[int] delay_seconds: NotRequired[float] queue_partition_key: NotRequired[str] authenticated_user: NotRequired[str] authenticated_roles: NotRequired[list[str]] serialization_type: NotRequired[WorkflowSerializationFormat] class_name: NotRequired[str] instance_name: NotRequired[str] attributes: NotRequired[Dict[str, Any]] otel_context: NotRequired[opentelemetry.context.Context] application_name: NotRequired[Optional[str]] client.enqueue( options: EnqueueOptions, *args: Any, **kwargs: Any ) -> WorkflowHandle[R] ``` Enqueue a workflow for processing and return a handle to it, similar to [Queue.enqueue](queues.md#enqueue). Returns a [WorkflowHandle](./workflow_handles.md#workflowhandle). When enqueuing a workflow from within a DBOS application, the workflow and queue metadata can be retrieved automatically. However, since `DBOSClient` runs outside the DBOS application, the metadata must be specified explicitly. Required metadata includes: * `workflow_name`: The registered name of the workflow being enqueued: the `name` passed to [`@DBOS.workflow`](./decorators.md#workflow), or by default the function's qualified name (for example, `URLFetcher.fetch_workflow` for a method on a class). * `queue_name`: The name of the [Queue](./queues.md) to enqueue the workflow on. Additional but optional metadata includes: * `workflow_id`: The unique ID for the enqueued workflow. If left undefined, DBOS Client will generate a [UUID](https://en.wikipedia.org/wiki/Universally_unique_identifier). Please see [Workflow IDs and Idempotency](../tutorials/workflow-tutorial#workflow-ids-and-idempotency) for more information. * `workflow_id_reuse_policy`: What to do if a workflow with ID `workflow_id` already exists, whatever its status. Defaults to `"return-existing"`. - `"return-existing"`: return a handle to the existing workflow without enqueueing a new one. - `"reject"`: raise `DBOSWorkflowIDInUseError` without enqueueing a new workflow or modifying the existing one. * `app_version`: The version of your application that should process this workflow. If left undefined, the workflow is only dequeued by an executor running the latest application version, and its version is set to that executor's version when it is first dequeued. - `workflow_timeout`: Set a timeout for the enqueued workflow. When the timeout expires, the workflow **and all its children** are cancelled. The timeout does not begin until the workflow is dequeued and starts execution. - `deduplication_id`: At any given time, only one workflow with a specific deduplication ID can be enqueued in the specified queue. If a workflow with a deduplication ID is currently delayed, enqueued, or actively executing (status `DELAYED`, `ENQUEUED`, or `PENDING`), subsequent workflow enqueue attempt with the same deduplication ID in the same queue will raise a `DBOSQueueDeduplicatedError` exception. - `duplication_policy`: How to handle a collision with another workflow that has the same `deduplication_id` on the same queue. Defaults to `"reject"`. - `"reject"`: raise `DBOSQueueDeduplicatedError`. - `"return-existing"`: return a handle to the existing workflow instead of raising. Requires `deduplication_id`. Arguments passed by the colliding caller are discarded and the returned handle resolves with the original workflow's result. See [Singleton Workflows](../tutorials/queue-tutorial.md#singleton-workflows). - `priority`: The priority of the enqueued workflow in the specified queue. Workflows with the same priority are dequeued in **FIFO (first in, first out)** order. Priority values can range from `1` to `2,147,483,647`, where **a low number indicates a higher priority**. Workflows without assigned priorities have the highest priority and are dequeued before workflows with assigned priorities. - `delay_seconds`: Delay the workflow by this many seconds before it becomes eligible for execution. The workflow is initially placed in `DELAYED` status and transitions to `ENQUEUED` after the delay expires. - `queue_partition_key`: The queue partition in which to enqueue this workflow. Use if and only if the queue is [partitioned](../tutorials/queue-tutorial.md#partitioning-queues) (registered with at least one `partition_*` limit). A partitioned queue applies its `partition_*` limits to each partition separately, while its `global_concurrency`, `worker_concurrency`, and `limiter` still apply across all partitions. - `authenticated_user`: An authenticated user to associate with the workflow. - `authenticated_roles`: Authenticated roles to associate with the workflow. - `serialization_type`: The [serialization strategy](./contexts.md#serialization-strategy) for the workflow arguments. - `class_name`: If the workflow is a class method (`@classmethod`) or a method on a [configured instance](../tutorials/classes.md), the registered name of its class: the `class_name` passed to [`@DBOS.dbos_class`](./decorators.md#dbos_class), or by default the class's qualified name. Not needed for static methods. - `instance_name`: If the workflow is a method on a [configured instance](../tutorials/classes.md), the `config_name` of the instance that runs it. Requires `class_name`. - `attributes`: A dictionary of custom, JSON-serializable key-value [attributes](./contexts.md#setworkflowattributes) to attach to the workflow. Recorded in the workflow's [status](./contexts.md#workflow-status) and searchable via the `attributes` filter on [`list_workflows`](#list_workflows). - `otel_context`: An OpenTelemetry context to propagate to the enqueued workflow, so that when the workflow runs, its span joins that context's trace. The client-side equivalent of [`PropagateOtelContext`](./contexts.md#propagateotelcontext). Only the W3C trace context (`traceparent`/`tracestate`) is propagated, not baggage. See [the tracing tutorial](../tutorials/logging-and-tracing.md#keeping-enqueued-workflows-on-the-callers-trace) for details. - `application_name`: The application that owns and runs the enqueued workflow. Defaults to the client's own [`application_name`](#constructor). Always set `application_name` either here or in the client constructor if multiple applications share a system database. To enqueue a workflow that is a method on a [Python class](../tutorials/classes.md), also set `class_name` (for a class method) or both `class_name` and `instance_name` (for a method on a configured instance). The class and instance must be registered in the application that dequeues the workflow; otherwise, the workflow cannot run. **Example syntax:** ```python from dbos import EnqueueOptions options: EnqueueOptions = { "queue_name": "example_queue", "workflow_name": "process_task", } handle = client.enqueue(options, task) result = handle.get_result() ``` To enqueue a method on a configured instance, such as `fetch_workflow` on the `URLFetcher("https://example.com")` instance from the [classes tutorial](../tutorials/classes.md): ```python options: EnqueueOptions = { "queue_name": "example_queue", "workflow_name": "URLFetcher.fetch_workflow", "class_name": "URLFetcher", "instance_name": "https://example.com", } handle = client.enqueue(options) ``` #### enqueue_async ```python client.enqueue_async( options: EnqueueOptions, *args: Any, **kwargs: Any ) -> WorkflowHandleAsync[R] ``` Similar to [enqueue](#enqueue), but enqueues asynchronously and returns a [WorkflowHandleAsync](workflow_handles.md#workflowhandleasync). **Example syntax:** ```python options: EnqueueOptions = { "queue_name": "example_queue", "workflow_name": "process_task", } handle = await client.enqueue_async(options, task) result = await handle.get_result() ``` #### enqueue_in_transaction ```python client.enqueue_in_transaction( conn_or_session: Union[sqlalchemy.Connection, sqlalchemy.orm.Session], options: EnqueueOptions, *args: Any, **kwargs: Any ) -> WorkflowHandle[R] ``` Similar to [enqueue](#enqueue), but performs the enqueue write inside a caller-owned SQLAlchemy transaction instead of in its own transaction. This lets you enqueue a workflow **atomically** with your own database writes: either both are committed or both are rolled back. Pass either a SQLAlchemy [`Connection`](https://docs.sqlalchemy.org/en/20/core/connections.html) or an ORM [`Session`](https://docs.sqlalchemy.org/en/20/orm/session_basics.html) as `conn_or_session`. The remaining parameters are the same as [enqueue](#enqueue), except that `duplication_policy="return-existing"` is not supported (raises `DBOSException`). You own the transaction: `enqueue_in_transaction` does not begin, commit, or roll back the transaction, and does not retry on database errors. You must commit (or roll back) the transaction yourself. The returned [WorkflowHandle](./workflow_handles.md#workflowhandle) is created immediately, but the workflow is not enqueued until you commit, so do not call `get_result` on the handle until after the transaction commits. :::warning `conn_or_session` must target the DBOS system database. The enqueue cannot atomically span a separate application database. ::: **Example syntax:** ```python import os import sqlalchemy as sa from dbos import EnqueueOptions # For Postgres, use a postgresql+psycopg:// URL: DBOS installs the psycopg (v3) driver, # while SQLAlchemy uses psycopg2 for a plain postgresql:// URL. engine = sa.create_engine( sa.make_url(os.environ["DBOS_SYSTEM_DATABASE_URL"]).set(drivername="postgresql+psycopg") ) options: EnqueueOptions = { "queue_name": "example_queue", "workflow_name": "process_task", } with engine.connect() as conn: with conn.begin(): # Perform your own writes on conn here, in the same transaction... handle = client.enqueue_in_transaction(conn, options, task) # Once the transaction commits, the workflow is enqueued. result = handle.get_result() ``` There is no asynchronous variant of this method. From an async context, bridge to it using [`AsyncConnection.run_sync`](https://docs.sqlalchemy.org/en/20/orm/extensions/asyncio.html), which hands your callable the underlying synchronous `Connection` bound to the same transaction: ```python async with async_engine.connect() as conn: async with conn.begin(): handle = await conn.run_sync( lambda sync_conn: client.enqueue_in_transaction(sync_conn, options, task) ) ``` #### retrieve_workflow ```python client.retrieve_workflow( workflow_id: str, ) -> WorkflowHandle[R] ``` Retrieve the [handle](./workflow_handles.md#workflowhandle) of a workflow with identity `workflow_id`. Similar to [`DBOS.retrieve_workflow`](contexts.md#retrieve_workflow). **Parameters:** - `workflow_id`: The identifier of the workflow whose handle to retrieve. **Returns:** - The [WorkflowHandle](./workflow_handles.md#workflowhandle) of the workflow whose ID is `workflow_id`. **Raises:** - `DBOSNonExistentWorkflowError`: If no workflow with ID `workflow_id` exists. #### retrieve_workflow_async ```python client.retrieve_workflow_async( workflow_id: str, ) -> WorkflowHandleAsync[R] ``` Asynchronously retrieve the [handle](./workflow_handles.md#workflowhandleasync) of a workflow with identity `workflow_id`. Similar to [`DBOS.retrieve_workflow`](contexts.md#retrieve_workflow). **Parameters:** - `workflow_id`: The identifier of the workflow whose handle to retrieve. **Returns:** - The [WorkflowHandleAsync](./workflow_handles.md#workflowhandleasync) of the workflow whose ID is `workflow_id`. **Raises:** - `DBOSNonExistentWorkflowError`: If no workflow with ID `workflow_id` exists. #### wait_first ```python client.wait_first( handles: List[WorkflowHandle[Any]], *, polling_interval_sec: float = 1.0, ) -> WorkflowHandle[Any] ``` Wait for any one of the given workflow handles to complete and return the first completed handle. Similar to [`DBOS.wait_first`](contexts.md#wait_first). **Parameters:** - **handles**: A non-empty list of workflow handles to wait on. Raises `ValueError` if the list is empty or contains duplicate workflow IDs. - **polling_interval_sec**: The interval (in seconds) at which DBOS polls the database. Defaults to `1.0`. #### wait_first_async ```python client.wait_first_async( handles: List[WorkflowHandleAsync[Any]], *, polling_interval_sec: float = 1.0, ) -> WorkflowHandleAsync[Any] ``` Asynchronous version of [`wait_first`](#wait_first). #### send ```python client.send( destination_id: str, message: Any, topic: Optional[str] = None, idempotency_key: Optional[str] = None, *, serialization_type: Optional[WorkflowSerializationFormat] = WorkflowSerializationFormat.DEFAULT, send_to_forks: bool = False, ) -> None ``` Sends a message to a specified workflow. Similar to [`DBOS.send`](contexts.md#send). **Parameters:** - `destination_id`: The workflow to which to send the message. - `message`: The message to send. Must be serializable. - `topic`: An optional topic with which to associate the message. Messages are enqueued per-topic on the receiver. - `idempotency_key`: An optional string used to ensure exactly-once delivery, even from outside of the DBOS application. The key is scoped per destination workflow. - `serialization_type`: The [serialization strategy](./contexts.md#serialization-strategy) for the message. - `send_to_forks`: If `True`, also deliver the message to every workflow recursively forked from `destination_id`. Defaults to `False`. :::warning Since DBOS Client is running outside of a DBOS application, it is highly recommended that you use the `idempotency_key` parameter with both `send` and `send_async` in order to get exactly-once behavior. ::: #### send_async ```python client.send_async( destination_id: str, message: Any, topic: Optional[str] = None, idempotency_key: Optional[str] = None, *, serialization_type: Optional[WorkflowSerializationFormat] = WorkflowSerializationFormat.DEFAULT, send_to_forks: bool = False, ) -> None ``` Asynchronously sends a message to a specified workflow. Similar to [`DBOS.send_async`](contexts.md#send_async). **Parameters:** - `destination_id`: The workflow to which to send the message. - `message`: The message to send. Must be serializable. - `topic`: An optional topic with which to associate the message. Messages are enqueued per-topic on the receiver. - `idempotency_key`: An optional string used to ensure exactly-once delivery, even from outside of the DBOS application. The key is scoped per destination workflow. - `serialization_type`: The [serialization strategy](./contexts.md#serialization-strategy) for the message. - `send_to_forks`: If `True`, also deliver the message to every workflow recursively forked from `destination_id`. Defaults to `False`. #### send_in_transaction ```python client.send_in_transaction( conn_or_session: Union[sqlalchemy.Connection, sqlalchemy.orm.Session], destination_id: str, message: Any, topic: Optional[str] = None, idempotency_key: Optional[str] = None, *, serialization_type: Optional[WorkflowSerializationFormat] = WorkflowSerializationFormat.DEFAULT, send_to_forks: bool = False, ) -> None ``` Similar to [send](#send), but performs the send inside a caller-owned SQLAlchemy transaction instead of in its own transaction. This lets you send a message **atomically** with your own database writes: either both are committed or both are rolled back. Pass either a SQLAlchemy [`Connection`](https://docs.sqlalchemy.org/en/20/core/connections.html) or an ORM [`Session`](https://docs.sqlalchemy.org/en/20/orm/session_basics.html) as `conn_or_session`. The remaining parameters are the same as [send](#send). You own the transaction: `send_in_transaction` does not begin, commit, or roll back the transaction, and does not retry on database errors. You must commit (or roll back) the transaction yourself. The message is not visible to the destination workflow until the transaction commits. :::warning `conn_or_session` must target the DBOS system database. The send cannot atomically span a separate application database. ::: **Example syntax:** ```python import os import sqlalchemy as sa # For Postgres, use a postgresql+psycopg:// URL (see the enqueue_in_transaction example) engine = sa.create_engine( sa.make_url(os.environ["DBOS_SYSTEM_DATABASE_URL"]).set(drivername="postgresql+psycopg") ) with engine.connect() as conn: with conn.begin(): # Perform your own writes on conn here, in the same transaction... client.send_in_transaction(conn, destination_id, message, idempotency_key="my-key") # Once the transaction commits, the message is sent. ``` There is no asynchronous variant of this method. From an async context, bridge to it using [`AsyncConnection.run_sync`](https://docs.sqlalchemy.org/en/20/orm/extensions/asyncio.html), which hands your callable the underlying synchronous `Connection` bound to the same transaction: ```python async with async_engine.connect() as conn: async with conn.begin(): await conn.run_sync( lambda sync_conn: client.send_in_transaction(sync_conn, destination_id, message) ) ``` #### send_bulk ```python client.send_bulk( messages: List[SendMessage], *, serialization_type: Optional[WorkflowSerializationFormat] = WorkflowSerializationFormat.DEFAULT, send_to_forks: bool = False, ) -> None ``` Sends many messages to workflow executions in a single transaction. Similar to [`DBOS.send_bulk`](contexts.md#send_bulk). Each message is described by a `SendMessage` object specifying its destination, payload, and optional topic and idempotency key: ```python @dataclass class SendMessage: # The workflow to which to send the message destination_id: str # The message to send. Must be serializable. message: Any # A topic with which to associate the message. Messages are enqueued per-topic on the receiver. topic: Optional[str] = None # If set, the message is sent only once per destination no matter how many times it is submitted with this key. idempotency_key: Optional[str] = None ``` The send is atomic: if any message cannot be delivered, the entire batch is rolled back and no messages are sent. **Parameters:** - `messages`: The list of `SendMessage` objects to send. Two messages in the same call may not share an idempotency key. - `serialization_type`: The [serialization strategy](./contexts.md#serialization-strategy) for the messages. - `send_to_forks`: If `True`, every message is also delivered to all workflows recursively forked from its destination. Defaults to `False`. :::warning Since DBOS Client is running outside of a DBOS application, it is highly recommended that you set an idempotency key on each message in order to get exactly-once behavior. ::: #### send_bulk_async ```python client.send_bulk_async( messages: List[SendMessage], *, serialization_type: Optional[WorkflowSerializationFormat] = WorkflowSerializationFormat.DEFAULT, send_to_forks: bool = False, ) -> None ``` Asynchronously sends many messages to workflow executions in a single transaction. Similar to [`DBOS.send_bulk_async`](contexts.md#send_bulk_async). See [`send_bulk`](#send_bulk) for the `SendMessage` definition. **Parameters:** - `messages`: The list of `SendMessage` objects to send. Two messages in the same call may not share an idempotency key. - `serialization_type`: The [serialization strategy](./contexts.md#serialization-strategy) for the messages. - `send_to_forks`: If `True`, every message is also delivered to all workflows recursively forked from its destination. Defaults to `False`. #### send_bulk_in_transaction ```python client.send_bulk_in_transaction( conn_or_session: Union[sqlalchemy.Connection, sqlalchemy.orm.Session], messages: List[SendMessage], *, serialization_type: Optional[WorkflowSerializationFormat] = WorkflowSerializationFormat.DEFAULT, send_to_forks: bool = False, ) -> None ``` Similar to [send_bulk](#send_bulk), but performs the sends inside a caller-owned SQLAlchemy transaction instead of in its own transaction. This lets you send messages **atomically** with your own database writes: either both are committed or both are rolled back. Pass either a SQLAlchemy [`Connection`](https://docs.sqlalchemy.org/en/20/core/connections.html) or an ORM [`Session`](https://docs.sqlalchemy.org/en/20/orm/session_basics.html) as `conn_or_session`. See [`send_bulk`](#send_bulk) for the `SendMessage` definition. You own the transaction: `send_bulk_in_transaction` does not begin, commit, or roll back the transaction, and does not retry on database errors. You must commit (or roll back) the transaction yourself. The messages are not visible to their destination workflows until the transaction commits. :::warning `conn_or_session` must target the DBOS system database. The send cannot atomically span a separate application database. ::: There is no asynchronous variant of this method. From an async context, bridge to it using [`AsyncConnection.run_sync`](https://docs.sqlalchemy.org/en/20/orm/extensions/asyncio.html) as shown for [`send_in_transaction`](#send_in_transaction). **Parameters:** - `messages`: The list of `SendMessage` objects to send. Two messages in the same call may not share an idempotency key. - `serialization_type`: The [serialization strategy](./contexts.md#serialization-strategy) for the messages. - `send_to_forks`: If `True`, every message is also delivered to all workflows recursively forked from its destination. Defaults to `False`. #### get_event ```python client.get_event( workflow_id: str, key: str, timeout_seconds: float = 60 ) -> Any ``` Retrieve the latest value of an event published by the workflow identified by `workflow_id` to the key `key`. If the event does not yet exist, wait for it to be published, returning `None` if the wait times out. Similar to [`DBOS.get_event`](contexts.md#get_event). **Parameters:** - `workflow_id`: The identifier of the workflow whose events to retrieve. - `key`: The key of the event to retrieve. - `timeout_seconds`: A timeout in seconds. If the wait times out, return `None`. **Returns:** - The value of the event published by `workflow_id` with name `key`, or `None` if the wait times out. #### get_event_async ```python client.get_event_async( workflow_id: str, key: str, timeout_seconds: float = 60 ) -> Any ``` Asynchronously retrieve the latest value of an event published by the workflow identified by `workflow_id` to the key `key`. If the event does not yet exist, wait for it to be published, returning `None` if the wait times out. Similar to [`DBOS.get_event_async`](contexts.md#get_event_async). **Parameters:** - `workflow_id`: The identifier of the workflow whose events to retrieve. - `key`: The key of the event to retrieve. - `timeout_seconds`: A timeout in seconds. If the wait times out, return `None`. **Returns:** - The value of the event published by `workflow_id` with name `key`, or `None` if the wait times out. #### read_stream ```python client.read_stream( workflow_id: str, key: str, *, offset: int = 0, polling_interval_sec: Optional[float] = None, timeout_seconds: Optional[float] = None, ) -> Generator[Any, Any, None] ``` Read values from a stream as a generator. This function reads values from a stream identified by the workflow_id and key, yielding each value in order until the stream is closed or the workflow terminates. Similar to [`DBOS.read_stream`](contexts.md#read_stream), except that client reads are never checkpointed. **Parameters:** - `workflow_id`: The workflow instance ID that owns the stream - `key`: The stream key / name within the workflow - `offset`: The offset to start reading from. Defaults to `0`, the start of the stream. A higher offset skips that many values from the beginning of the stream. - `polling_interval_sec`: Polling interval in seconds when waiting for new values when not using LISTEN/NOTIFY. Must be at least `0.001`. Defaults to `1.0`. - `timeout_seconds`: How long to wait for **each** value before raising `DBOSStreamTimeoutError`. The clock restarts every time a value is delivered, so this bounds the gap between values, not the total duration of the read. Defaults to `None`, waiting indefinitely. **Yields:** - Each value in the stream until the stream is closed **Raises:** - `DBOSStreamTimeoutError`: If `timeout_seconds` passes without a value arriving. - `DBOSNonExistentWorkflowError`: If no workflow with ID `workflow_id` exists. **Example syntax:** ```python for value in client.read_stream(workflow_id, "results"): print(f"Received: {value}") ``` #### read_stream_async ```python client.read_stream_async( workflow_id: str, key: str, *, offset: int = 0, polling_interval_sec: Optional[float] = None, timeout_seconds: Optional[float] = None, ) -> AsyncGenerator[Any, None] ``` Coroutine version of [`read_stream`](#read_stream), returning an async generator. **Example syntax:** ```python async for value in client.read_stream_async(workflow_id, "results"): print(f"Received: {value}") ``` #### read_stream_offset ```python client.read_stream_offset( workflow_id: str, key: str, offset: int, *, polling_interval_sec: Optional[float] = None, timeout_seconds: Optional[float] = None, ) -> Any ``` Read the single value at one offset of a stream, waiting for it to be written. Similar to [`DBOS.read_stream_offset`](contexts.md#read_stream_offset). **Parameters:** - `workflow_id`: The workflow instance ID that owns the stream - `key`: The stream key / name within the workflow - `offset`: The offset to read - `polling_interval_sec`: Polling interval in seconds when waiting for the value when not using LISTEN/NOTIFY. Must be at least `0.001`. Defaults to `1.0`. - `timeout_seconds`: How long to wait for the value before raising `DBOSStreamTimeoutError`. Defaults to `None`, waiting indefinitely. **Returns:** - The value at the offset **Raises:** - `DBOSStreamTimeoutError`: If `timeout_seconds` passes, or if the stream ends before reaching `offset` (no value will ever arrive at that offset). - `DBOSNonExistentWorkflowError`: If no workflow with ID `workflow_id` exists. **Example syntax:** ```python value = client.read_stream_offset(workflow_id, "results", 5, timeout_seconds=30) ``` #### read_stream_offset_async ```python client.read_stream_offset_async( workflow_id: str, key: str, offset: int, *, polling_interval_sec: Optional[float] = None, timeout_seconds: Optional[float] = None, ) -> Coroutine[Any, Any, Any] ``` Coroutine version of [`read_stream_offset`](#read_stream_offset). #### set_workflow_delay ```python client.set_workflow_delay( workflow_id: str, *, delay_seconds: Optional[float] = None, delay_until_epoch_ms: Optional[int] = None, ) -> None ``` Set or update the delay on a workflow. Only affects workflows with `DELAYED` status. Provide exactly one of `delay_seconds` (relative) or `delay_until_epoch_ms` (absolute). Similar to [`DBOS.set_workflow_delay`](./contexts.md#set_workflow_delay). **Parameters:** - `workflow_id`: The ID of the workflow whose delay to set. - `delay_seconds`: Delay the workflow by this many seconds from now. Must be non-negative. - `delay_until_epoch_ms`: Delay the workflow until this absolute time, specified as a Unix epoch timestamp in milliseconds. Must be non-negative. #### set_workflow_delay_async ```python client.set_workflow_delay_async( workflow_id: str, *, delay_seconds: Optional[float] = None, delay_until_epoch_ms: Optional[int] = None, ) -> None ``` Asynchronous version of [`set_workflow_delay`](#set_workflow_delay). ### Queue Management Methods #### register_queue ```python client.register_queue( name: str, *, # Applied to the queue as a whole global_concurrency: Optional[int] = None, worker_concurrency: Optional[int] = None, limiter: Optional[QueueRateLimit] = None, # Applied to each partition separately partition_concurrency: Optional[int] = None, partition_worker_concurrency: Optional[int] = None, partition_limiter: Optional[QueueRateLimit] = None, polling_interval_sec: float = 1.0, on_conflict: QueueConflictResolution = "always_update", application_name: Optional[str] = None, ) -> Queue ``` Register a [queue](./queues.md) and persist its configuration to the system database, returning the [`Queue`](./queues.md#class-dbosqueue). Similar to [`DBOS.register_queue`](./contexts.md#register_queue). Parameters have the same meaning as on `DBOS.register_queue` except for `on_conflict` and `application_name`: - `on_conflict`: - `"always_update"` (default): always overwrite the existing configuration. - `"never_update"`: leave any existing configuration unchanged. - `"update_if_latest_version"` is **not** supported on the client because clients are not associated with an application version. Passing it raises `DBOSException`. - `application_name`: The application that owns this queue and dequeues workflows from it. Defaults to the client's own [`application_name`](#constructor). Registering a queue already owned by a different application raises an error. **Example syntax:** ```python import os from dbos import DBOSClient client = DBOSClient(system_database_url=os.environ["DBOS_SYSTEM_DATABASE_URL"]) client.register_queue("email", global_concurrency=10, limiter={"limit": 100, "period": 60}) client.enqueue({"queue_name": "email", "workflow_name": "send_email"}, "alice@example.com") ``` #### register_queue_async ```python client.register_queue_async( name: str, *, global_concurrency: Optional[int] = None, worker_concurrency: Optional[int] = None, limiter: Optional[QueueRateLimit] = None, partition_concurrency: Optional[int] = None, partition_worker_concurrency: Optional[int] = None, partition_limiter: Optional[QueueRateLimit] = None, polling_interval_sec: float = 1.0, on_conflict: QueueConflictResolution = "always_update", application_name: Optional[str] = None, ) -> Coroutine[Any, Any, Queue] ``` Asynchronous version of [`register_queue`](#register_queue). #### retrieve_queue ```python client.retrieve_queue(name: str) -> Optional[Queue] ``` Retrieve a queue by name from the system database, or `None` if no queue with that name has been registered. Similar to [`DBOS.retrieve_queue`](./contexts.md#retrieve_queue). #### retrieve_queue_async ```python client.retrieve_queue_async(name: str) -> Coroutine[Any, Any, Optional[Queue]] ``` Asynchronous version of [`retrieve_queue`](#retrieve_queue). #### list_queues ```python client.list_queues( *, application_name: Optional[Union[str, List[str]]] = None, ) -> List[Queue] ``` List all queues registered in the system database. Returns an empty list if no queues have been registered. Similar to [`DBOS.list_queues`](./contexts.md#list_queues), including the `application_name` filter. If the filter is unset, it defaults to the client's own [`application_name`](#constructor); a client with no application name lists every application's queues. #### list_queues_async ```python client.list_queues_async( *, application_name: Optional[Union[str, List[str]]] = None, ) -> Coroutine[Any, Any, List[Queue]] ``` Asynchronous version of [`list_queues`](#list_queues). #### delete_queue ```python client.delete_queue(name: str) -> None ``` Delete a queue from the system database. No-op if no queue with that name exists. Similar to [`DBOS.delete_queue`](./contexts.md#delete_queue). :::warning Workflows already enqueued on a deleted queue can no longer be dequeued, executed, or recovered. However, if a queue with the same name is later registered, it will dequeue the leftover workflows. Do not rely on this: stale workflows unexpectedly resuming on a future queue is rarely the intended behavior. Instead, cancel or drain pending workflows on the queue before deleting it. ::: #### delete_queue_async ```python client.delete_queue_async(name: str) -> Coroutine[Any, Any, None] ``` Asynchronous version of [`delete_queue`](#delete_queue). ### Workflow Management Methods #### list_workflows ```python client.list_workflows( *, workflow_ids: Optional[List[str]] = None, status: Optional[Union[str, List[str]]] = None, start_time: Optional[str] = None, end_time: Optional[str] = None, completed_after: Optional[str] = None, completed_before: Optional[str] = None, dequeued_after: Optional[str] = None, dequeued_before: Optional[str] = None, name: Optional[Union[str, List[str]]] = None, app_version: Optional[Union[str, List[str]]] = None, forked_from: Optional[Union[str, List[str]]] = None, parent_workflow_id: Optional[Union[str, List[str]]] = None, user: Optional[Union[str, List[str]]] = None, queue_name: Optional[Union[str, List[str]]] = None, limit: Optional[int] = None, offset: Optional[int] = None, sort_desc: bool = False, workflow_id_prefix: Optional[Union[str, List[str]]] = None, load_input: bool = True, load_output: bool = True, executor_id: Optional[Union[str, List[str]]] = None, queues_only: bool = False, was_forked_from: Optional[bool] = None, has_parent: Optional[bool] = None, attributes: Optional[Dict[str, Any]] = None, schedule_name: Optional[Union[str, List[str]]] = None, application_name: Optional[Union[str, List[str]]] = None, ) -> List[WorkflowStatus]: ``` Retrieve a list of [`WorkflowStatus`](./contexts#workflow-status) of all workflows matching specified criteria. Similar to [`DBOS.list_workflows`](./contexts#list_workflows). **Parameters:** - **workflow_ids**: Retrieve workflows with these IDs. - **status**: Retrieve workflows with this status (or one of these statuses) (Must be `ENQUEUED`, `DELAYED`, `PENDING`, `SUCCESS`, `ERROR`, `CANCELLED`, or `MAX_RECOVERY_ATTEMPTS_EXCEEDED`) - **start_time**: Retrieve workflows started after this (RFC 3339-compliant) timestamp. - **end_time**: Retrieve workflows started before this (RFC 3339-compliant) timestamp. - **completed_after**: Retrieve workflows that completed after this (RFC 3339-compliant) timestamp. - **completed_before**: Retrieve workflows that completed before this (RFC 3339-compliant) timestamp. - **dequeued_after**: Retrieve workflows that were dequeued after this (RFC 3339-compliant) timestamp. - **dequeued_before**: Retrieve workflows that were dequeued before this (RFC 3339-compliant) timestamp. - **name**: Retrieve workflows with this name (or one of these names). - **app_version**: Retrieve workflows tagged with this application version (or one of these versions). - **forked_from**: Retrieve workflows forked from this workflow ID (or one of these IDs). - **parent_workflow_id**: Retrieve workflows that were started as children of this workflow (or one of these workflows). - **user**: Retrieve workflows run by this authenticated user (or one of these users). - **queue_name**: Retrieve workflows that were enqueued on this queue (or one of these queues). - **limit**: Retrieve up to this many workflows. - **offset**: Skip this many workflows from the results returned (for pagination). - **sort_desc**: Whether to sort the results in descending (`True`) or ascending (`False`) order by workflow start time. - **workflow_id_prefix**: Retrieve workflows whose IDs start with the specified string (or one of these strings). - **load_input**: Whether to load and deserialize workflow inputs. Set to `False` to improve performance when inputs are not needed. - **load_output**: Whether to load and deserialize workflow outputs. Set to `False` to improve performance when outputs are not needed. - **executor_id**: Retrieve workflows with this executor ID (or one of these IDs). - **queues_only**: If `True`, only retrieve workflows that are currently queued (status `DELAYED`, `ENQUEUED`, or `PENDING` and `queue_name` not null). Equivalent to using [`list_queued_workflows`](#list_queued_workflows). - **was_forked_from**: If `True`, only retrieve workflows that have been forked from. If `False`, only retrieve workflows that have not been forked from. - **has_parent**: If `True`, only retrieve workflows that have a parent workflow. If `False`, only retrieve workflows without a parent. - **attributes**: Retrieve workflows whose [custom attributes](./contexts.md#setworkflowattributes) contain all the given key-value pairs (nested values are matched exactly). Only supported when using a Postgres system database; raises `DBOSException` on SQLite. - **schedule_name**: Retrieve workflows that were enqueued by this [scheduled workflow](../tutorials/scheduled-workflows.md) (or one of these schedule names). - **application_name**: Retrieve workflows owned by this application (or one of these applications). Workflows owned by no application are always included. If unset, defaults to the client's own [`application_name`](#constructor) unless `workflow_ids` is set; a client with no application name retrieves every application's workflows. #### list_workflows_async ```python client.list_workflows_async( *, workflow_ids: Optional[List[str]] = None, status: Optional[Union[str, List[str]]] = None, start_time: Optional[str] = None, end_time: Optional[str] = None, completed_after: Optional[str] = None, completed_before: Optional[str] = None, dequeued_after: Optional[str] = None, dequeued_before: Optional[str] = None, name: Optional[Union[str, List[str]]] = None, app_version: Optional[Union[str, List[str]]] = None, forked_from: Optional[Union[str, List[str]]] = None, parent_workflow_id: Optional[Union[str, List[str]]] = None, user: Optional[Union[str, List[str]]] = None, queue_name: Optional[Union[str, List[str]]] = None, limit: Optional[int] = None, offset: Optional[int] = None, sort_desc: bool = False, workflow_id_prefix: Optional[Union[str, List[str]]] = None, load_input: bool = True, load_output: bool = True, executor_id: Optional[Union[str, List[str]]] = None, queues_only: bool = False, was_forked_from: Optional[bool] = None, has_parent: Optional[bool] = None, attributes: Optional[Dict[str, Any]] = None, schedule_name: Optional[Union[str, List[str]]] = None, application_name: Optional[Union[str, List[str]]] = None, ) -> List[WorkflowStatus]: ``` Asynchronous version of [`DBOSClient.list_workflows`](#list_workflows). #### list_queued_workflows ```python client.list_queued_workflows( *, workflow_ids: Optional[List[str]] = None, status: Optional[Union[str, List[str]]] = None, start_time: Optional[str] = None, end_time: Optional[str] = None, completed_after: Optional[str] = None, completed_before: Optional[str] = None, dequeued_after: Optional[str] = None, dequeued_before: Optional[str] = None, name: Optional[Union[str, List[str]]] = None, app_version: Optional[Union[str, List[str]]] = None, forked_from: Optional[Union[str, List[str]]] = None, parent_workflow_id: Optional[Union[str, List[str]]] = None, user: Optional[Union[str, List[str]]] = None, queue_name: Optional[Union[str, List[str]]] = None, limit: Optional[int] = None, offset: Optional[int] = None, sort_desc: bool = False, workflow_id_prefix: Optional[Union[str, List[str]]] = None, load_input: bool = True, load_output: bool = True, executor_id: Optional[Union[str, List[str]]] = None, has_parent: Optional[bool] = None, attributes: Optional[Dict[str, Any]] = None, application_name: Optional[Union[str, List[str]]] = None, ) -> List[WorkflowStatus]: ``` Retrieve a list of [`WorkflowStatus`](./contexts#workflow-status) of all **queued** workflows (status `DELAYED`, `ENQUEUED`, or `PENDING` and `queue_name` not null) matching specified criteria. Similar to [`DBOS.list_queued_workflows`](./contexts.md#list_queued_workflows). **Parameters:** - **workflow_ids**: Retrieve workflows with these IDs. - **status**: Retrieve workflows with this status (or one of these statuses) (Must be `DELAYED`, `ENQUEUED`, or `PENDING`) - **start_time**: Retrieve workflows enqueued after this (RFC 3339-compliant) timestamp. - **end_time**: Retrieve workflows enqueued before this (RFC 3339-compliant) timestamp. - **completed_after**: Retrieve workflows that completed after this (RFC 3339-compliant) timestamp. - **completed_before**: Retrieve workflows that completed before this (RFC 3339-compliant) timestamp. - **dequeued_after**: Retrieve workflows that were dequeued after this (RFC 3339-compliant) timestamp. - **dequeued_before**: Retrieve workflows that were dequeued before this (RFC 3339-compliant) timestamp. - **name**: Retrieve workflows with this name (or one of these names). - **app_version**: Retrieve workflows tagged with this application version (or one of these versions). - **forked_from**: Retrieve workflows forked from this workflow ID (or one of these IDs). - **parent_workflow_id**: Retrieve workflows that were started as children of this workflow (or one of these workflows). - **user**: Retrieve workflows run by this authenticated user (or one of these users). - **queue_name**: Retrieve workflows running on this queue (or one of these queues). - **limit**: Retrieve up to this many workflows. - **offset**: Skip this many workflows from the results returned (for pagination). - **sort_desc**: Whether to sort the results in descending (`True`) or ascending (`False`) order by workflow start time. - **workflow_id_prefix**: Retrieve workflows whose IDs start with the specified string (or one of these strings). - **load_input**: Whether to load and deserialize workflow inputs. Set to `False` to improve performance when inputs are not needed. - **load_output**: Whether to load and deserialize workflow outputs. Set to `False` to improve performance when outputs are not needed. - **executor_id**: Retrieve workflows with this executor ID (or one of these IDs). - **has_parent**: If `True`, only retrieve workflows that have a parent workflow. If `False`, only retrieve workflows without a parent. - **attributes**: Retrieve workflows whose [custom attributes](./contexts.md#setworkflowattributes) contain all the given key-value pairs (nested values are matched exactly). Only supported when using a Postgres system database; raises `DBOSException` on SQLite. - **application_name**: Retrieve workflows owned by this application (or one of these applications). Workflows owned by no application are always included. If unset, defaults to the client's own [`application_name`](#constructor) unless `workflow_ids` is set; a client with no application name retrieves every application's workflows. #### list_queued_workflows_async ```python client.list_queued_workflows_async( *, workflow_ids: Optional[List[str]] = None, status: Optional[Union[str, List[str]]] = None, start_time: Optional[str] = None, end_time: Optional[str] = None, completed_after: Optional[str] = None, completed_before: Optional[str] = None, dequeued_after: Optional[str] = None, dequeued_before: Optional[str] = None, name: Optional[Union[str, List[str]]] = None, app_version: Optional[Union[str, List[str]]] = None, forked_from: Optional[Union[str, List[str]]] = None, parent_workflow_id: Optional[Union[str, List[str]]] = None, user: Optional[Union[str, List[str]]] = None, queue_name: Optional[Union[str, List[str]]] = None, limit: Optional[int] = None, offset: Optional[int] = None, sort_desc: bool = False, workflow_id_prefix: Optional[Union[str, List[str]]] = None, load_input: bool = True, load_output: bool = True, executor_id: Optional[Union[str, List[str]]] = None, has_parent: Optional[bool] = None, attributes: Optional[Dict[str, Any]] = None, application_name: Optional[Union[str, List[str]]] = None, ) -> List[WorkflowStatus]: ``` Asynchronous version of [`DBOSClient.list_queued_workflows`](#list_queued_workflows). #### list_workflow_steps ```python client.list_workflow_steps( workflow_id: str, *, load_output: bool = True, limit: Optional[int] = None, offset: Optional[int] = None, ) -> List[StepInfo] ``` Similar to [`DBOS.list_workflow_steps`](./contexts.md#list_workflow_steps). **Parameters:** - **workflow_id**: The ID of the workflow whose steps to list. - **load_output**: Whether to load and deserialize step outputs and errors. Set to `False` to improve performance when they are not needed. - **limit**: The maximum number of steps to return. - **offset**: The number of steps to skip, for pagination. #### list_workflow_steps_async ```python client.list_workflow_steps_async( workflow_id: str, *, load_output: bool = True, limit: Optional[int] = None, offset: Optional[int] = None, ) -> List[StepInfo] ``` Asynchronous version of [`list_workflow_steps`](#list_workflow_steps). #### cancel_workflow ```python client.cancel_workflow( workflow_id: str, *, cancel_children: bool = False, ) -> None ``` Cancel a workflow. This sets its status to `CANCELLED`, removes it from its queue (if it is enqueued) and preempts its execution (interrupting it at the beginning of its next step). Similar to [`DBOS.cancel_workflow`](./contexts.md#cancel_workflow). **Parameters:** - **workflow_id**: The ID of the workflow to cancel. - **cancel_children**: If `True`, also recursively cancels all child workflows started by this workflow. #### cancel_workflow_async ```python client.cancel_workflow_async( workflow_id: str, *, cancel_children: bool = False, ) -> None ``` Asynchronous version of [`DBOSClient.cancel_workflow`](#cancel_workflow). #### cancel_workflows ```python client.cancel_workflows( workflow_ids: List[str], *, cancel_children: bool = False, ) -> None ``` Cancel multiple workflows. Behaves like [`cancel_workflow`](#cancel_workflow) but operates on a list of workflow IDs. Similar to [`DBOS.cancel_workflows`](./contexts.md#cancel_workflows). #### cancel_workflows_async Asynchronous version of [`DBOSClient.cancel_workflows`](#cancel_workflows). #### update_workflow_attributes ```python client.update_workflow_attributes( workflow_id: str, attributes: Optional[Dict[str, Any]], ) -> None ``` Replace the custom [attributes](./contexts.md#setworkflowattributes) attached to a workflow, identified by `workflow_id`. This overwrites the workflow's attributes dictionary; it is not a merge. Pass `None` to clear all attributes. Attributes must be a dictionary of JSON-serializable values. Similar to [`DBOS.update_workflow_attributes`](./contexts.md#update_workflow_attributes). **Parameters:** - `workflow_id`: The ID of the workflow whose attributes to replace. - `attributes`: The new attributes dictionary, or `None` to clear all attributes. #### update_workflow_attributes_async Asynchronous version of [`DBOSClient.update_workflow_attributes`](#update_workflow_attributes). #### resume_workflow ```python client.resume_workflow( workflow_id: str, *, queue_name: Optional[str] = None, ) -> WorkflowHandle[Any] ``` Resume a workflow. This immediately starts it from its last completed step. You can use this to resume workflows that are cancelled or have exceeded their maximum recovery attempts. You can also use this to start an enqueued workflow immediately, bypassing its queue. If `queue_name` is provided, the resumed workflow is enqueued on the specified queue instead of starting immediately. Raises `DBOSNonExistentWorkflowError` if no workflow with ID `workflow_id` exists. Similar to [`DBOS.resume_workflow`](./contexts.md#resume_workflow). #### resume_workflow_async ```python client.resume_workflow_async( workflow_id: str, *, queue_name: Optional[str] = None, ) -> WorkflowHandleAsync[Any] ``` Asynchronous version of [`DBOSClient.resume_workflow`](#resume_workflow). #### resume_workflows ```python client.resume_workflows( workflow_ids: List[str], *, queue_name: Optional[str] = None, ) -> List[WorkflowHandle[Any]] ``` Resume multiple workflows. Behaves like [`resume_workflow`](#resume_workflow) but operates on a list of workflow IDs and returns a list of handles. If any of the workflows does not exist, raises `DBOSNonExistentWorkflowError` and resumes none of them. Similar to [`DBOS.resume_workflows`](./contexts.md#resume_workflows). #### resume_workflows_async Asynchronous version of [`DBOSClient.resume_workflows`](#resume_workflows). Returns `List[WorkflowHandleAsync[Any]]`. #### fork_workflow ```python client.fork_workflow( workflow_id: str, start_step: int, *, application_version: Optional[str] = None, queue_name: Optional[str] = None, queue_partition_key: Optional[str] = None, replacement_children: Optional[dict[str, str]] = None, timeout_seconds: Optional[float] = None, ) -> WorkflowHandle[Any] ``` Similar to [`DBOS.fork_workflow`](./contexts.md#fork_workflow). Raises `DBOSNonExistentWorkflowError` if no workflow with ID `workflow_id` exists. #### fork_workflow_async ```python client.fork_workflow_async( workflow_id: str, start_step: int, *, application_version: Optional[str] = None, queue_name: Optional[str] = None, queue_partition_key: Optional[str] = None, replacement_children: Optional[dict[str, str]] = None, timeout_seconds: Optional[float] = None, ) -> WorkflowHandleAsync[Any] ``` Asynchronous version of [`DBOSClient.fork_workflow`](#fork_workflow). #### rewind_workflow ```python client.rewind_workflow( workflow_id: str, *, start_step: Optional[int] = None, application_version: Optional[str] = None, queue_name: Optional[str] = None, queue_partition_key: Optional[str] = None, ) -> WorkflowHandle[Any] ``` Similar to [`DBOS.rewind_workflow`](./contexts.md#rewind_workflow). Only a workflow in a terminal state can be rewound. Raises `DBOSNonExistentWorkflowError` if no workflow with ID `workflow_id` exists. :::warning Client rewind does not delete [datasource](./datasources.md) transaction checkpoints, so rewound transactions are not re-executed. To rewind workflows that use datasources, use [`DBOS.rewind_workflow`](./contexts.md#rewind_workflow). ::: #### rewind_workflow_async ```python client.rewind_workflow_async( workflow_id: str, *, start_step: Optional[int] = None, application_version: Optional[str] = None, queue_name: Optional[str] = None, queue_partition_key: Optional[str] = None, ) -> WorkflowHandleAsync[Any] ``` Asynchronous version of [`DBOSClient.rewind_workflow`](#rewind_workflow). #### delete_workflow ```python client.delete_workflow( workflow_id: str, *, delete_children: bool = False, ) -> None ``` Delete a workflow and all its associated data from the system database. Similar to [`DBOS.delete_workflow`](./contexts.md#delete_workflow). **Parameters:** - **workflow_id**: The ID of the workflow to delete. - **delete_children**: If `True`, also recursively deletes all child workflows started by this workflow. :::warning This operation is irreversible. Once a workflow is deleted, it cannot be recovered, resumed, or forked. ::: #### delete_workflow_async Asynchronous version of [`DBOSClient.delete_workflow`](#delete_workflow). #### delete_workflows ```python client.delete_workflows( workflow_ids: List[str], *, delete_children: bool = False, ) -> None ``` Delete multiple workflows and all their associated data. Behaves like [`delete_workflow`](#delete_workflow) but operates on a list of workflow IDs. Similar to [`DBOS.delete_workflows`](./contexts.md#delete_workflows). #### delete_workflows_async Asynchronous version of [`DBOSClient.delete_workflows`](#delete_workflows). ### Debouncing Workflows can be [debounced](./contexts.md#debouncing) with the DBOSClient. #### DebouncerClient ```python DebouncerClient( client: DBOSClient, workflow_options: EnqueueOptions, *, debounce_timeout_sec: Optional[float] = None, queue: Optional[Union[Queue, str]] = None, application_name: Optional[str] = None, ) ``` Similar to [`Debouncer.create`](./contexts.md#debouncercreate) but takes in a DBOSClient and `EnqueueOptions` instead of a workflow function. If `queue` (a queue or queue name) is set, it overrides the `queue_name` in `workflow_options`. `workflow_options` must not set `deduplication_id`, `delay_seconds`, `priority`, `queue_partition_key`, `duplication_policy="return-existing"`, or `workflow_id_reuse_policy="reject"`: `debounce` raises `DBOSException` if they are set. `application_name` debounces on behalf of that application; it defaults to the `application_name` in `workflow_options`, then to the client's own. #### debounce ```python debouncerClient.debounce( debounce_key: str, debounce_period_sec: float, *args: Any, **kwargs: Any, ) -> WorkflowHandle[R] ``` Similar to [`Debouncer.debounce`](./contexts.md#debounce). **Example Syntax**: ```python client: DBOSClient = ... workflow_options: EnqueueOptions = { "workflow_name": "process_input", "queue_name": "process_input_queue", } debouncer = DebouncerClient(client, workflow_options) # Each time a user submits a new input, debounce the process_input workflow. # The workflow will wait until 60 seconds after the user stops submitting new inputs, # then process the last input submitted. def on_user_input_submit(user_id, user_input): debounce_key = user_id debounce_period_sec = 60 debouncer.debounce(debounce_key, debounce_period_sec, user_input) ``` #### debounce_async ```python debouncerClient.debounce_async( debounce_key: str, debounce_period_sec: float, *args: Any, **kwargs: Any, ) -> WorkflowHandleAsync[R]: ``` Similar to [`Debouncer.debounce_async`](./contexts.md#debounce_async). ### Workflow Schedules `DBOSClient` provides methods to manage [workflow schedules](./contexts.md#workflow-schedules) from outside a DBOS application. Unlike the `DBOS` class methods which accept workflow functions directly, client schedule methods accept workflow names as strings. #### create_schedule ```python client.create_schedule( *, schedule_name: str, workflow_name: str, schedule: str, context: Any = None, workflow_class_name: Optional[str] = None, automatic_backfill: bool = False, cron_timezone: Optional[str] = None, queue_name: Optional[str] = None, application_name: Optional[str] = None, ) -> None ``` Create a cron schedule that periodically invokes a workflow. Similar to [`DBOS.create_schedule`](./contexts.md#create_schedule), but takes a `workflow_name` string instead of a workflow function. **Parameters:** - **schedule_name**: Unique name identifying this schedule. - **workflow_name**: Registered name of the workflow function to invoke. - **schedule**: A cron expression. Supports seconds as the first field with 6-field format. - **context**: An optional context object passed to the workflow function on each invocation. Must be serializable. - **workflow_class_name**: The registered class name if the workflow is a class method (`@classmethod`) on a [DBOS class](../tutorials/classes.md). - **automatic_backfill**: If `True`, on startup the scheduler will automatically backfill missed executions since the last time the schedule fired. Defaults to `False`. - **cron_timezone**: [IANA timezone name](https://en.wikipedia.org/wiki/List_of_tz_database_time_zones) (e.g. `"America/New_York"`) in which to evaluate the cron expression. Defaults to `None` (UTC). - **queue_name**: Optional name of a declared queue to enqueue scheduled workflows to. If `None`, uses an internal queue. Defaults to `None`. - **application_name**: The application that owns this schedule and runs its workflows. Defaults to the client's own [`application_name`](#constructor). Always set `application_name` either here or in the client constructor if multiple applications share a system database. #### create_schedule_async Coroutine version of [`create_schedule`](#create_schedule). #### list_schedules ```python client.list_schedules( *, status: Optional[Union[str, List[str]]] = None, workflow_name: Optional[Union[str, List[str]]] = None, schedule_name_prefix: Optional[Union[str, List[str]]] = None, application_name: Optional[Union[str, List[str]]] = None, ) -> List[WorkflowSchedule] ``` Return all registered workflow schedules, optionally filtered. Returns a list of [`WorkflowSchedule`](./contexts.md#workflowschedule). Similar to [`DBOS.list_schedules`](./contexts.md#list_schedules). **Parameters:** - **status**: Filter by status (e.g. `"ACTIVE"`) or a list of statuses. - **workflow_name**: Filter by workflow name or a list of names. - **schedule_name_prefix**: Filter by schedule name prefix or a list of prefixes. - **application_name**: List only schedules owned by this application (or one of these applications). Schedules owned by no application are always included. If unset, defaults to the client's own [`application_name`](#constructor); a client with no application name lists every application's schedules. #### list_schedules_async Coroutine version of [`list_schedules`](#list_schedules). #### get_schedule ```python client.get_schedule(name: str) -> Optional[WorkflowSchedule] ``` Return the [`WorkflowSchedule`](./contexts.md#workflowschedule) with the given name, or `None` if it does not exist. Similar to [`DBOS.get_schedule`](./contexts.md#get_schedule). #### get_schedule_async Coroutine version of [`get_schedule`](#get_schedule). #### delete_schedule ```python client.delete_schedule(name: str) -> None ``` Delete the schedule with the given name. No-op if it does not exist. Similar to [`DBOS.delete_schedule`](./contexts.md#delete_schedule). #### delete_schedule_async Coroutine version of [`delete_schedule`](#delete_schedule). #### pause_schedule ```python client.pause_schedule(name: str) -> None ``` Pause the schedule with the given name. A paused schedule does not fire. Similar to [`DBOS.pause_schedule`](./contexts.md#pause_schedule). #### resume_schedule ```python client.resume_schedule(name: str) -> None ``` Resume a paused schedule so it begins firing again. Similar to [`DBOS.resume_schedule`](./contexts.md#resume_schedule). #### apply_schedules ```python client.apply_schedules( schedules: List[ClientScheduleInput], ) -> None class ClientScheduleInput(TypedDict): schedule_name: str workflow_name: str schedule: str context: Any # Optional, defaults to None workflow_class_name: Optional[str] # Optional, defaults to None automatic_backfill: bool # Optional, defaults to False cron_timezone: Optional[str] # Optional, defaults to None (UTC) queue_name: Optional[str] # Optional, defaults to None (internal queue) application_name: Optional[str] # Optional, defaults to the client's application_name ``` Atomically apply a set of schedules. Useful for declaratively defining all your static schedules in one place. #### apply_schedules_async Asynchronous version of [`apply_schedules`](#apply_schedules). #### backfill_schedule ```python client.backfill_schedule( schedule_name: str, start: datetime, end: datetime, ) -> List[WorkflowHandle[None]] ``` Enqueue (on the schedule's `queue_name`, or an internal queue if it has none) all executions of a schedule that would have run between `start` and `end`. Each execution uses the same deterministic workflow ID as the live scheduler, so already-executed times are skipped. Similar to [`DBOS.backfill_schedule`](./contexts.md#backfill_schedule). #### trigger_schedule ```python client.trigger_schedule(schedule_name: str) -> WorkflowHandle[None] ``` Immediately enqueue (on the schedule's `queue_name`, or an internal queue if it has none) the scheduled workflow at the current time. Similar to [`DBOS.trigger_schedule`](./contexts.md#trigger_schedule). ### Version Management #### list_application_versions ```python client.list_application_versions() -> List[VersionInfo] ``` Return all registered application versions, ordered by timestamp descending (newest first). Similar to [`DBOS.list_application_versions`](./contexts.md#list_application_versions). If the client has an [`application_name`](#constructor), only versions registered by that application (plus versions owned by no application) are returned; otherwise, every application's versions are returned. #### list_application_versions_async ```python await client.list_application_versions_async() -> List[VersionInfo] ``` Coroutine version of [`list_application_versions`](#list_application_versions). #### get_latest_application_version ```python client.get_latest_application_version() -> VersionInfo ``` Return the latest application version (the one with the highest timestamp). Like [`list_application_versions`](#list_application_versions), if the client has an [`application_name`](#constructor), only versions registered by that application (plus versions owned by no application) are considered. Raises `DBOSException` if no versions are registered. Similar to [`DBOS.get_latest_application_version`](./contexts.md#get_latest_application_version). #### get_latest_application_version_async ```python await client.get_latest_application_version_async() -> VersionInfo ``` Coroutine version of [`get_latest_application_version`](#get_latest_application_version). #### set_latest_application_version ```python client.set_latest_application_version( version_name: str, *, application_name: Optional[str] = None, ) -> None ``` Promote a version to latest by updating its timestamp to the current time. This is useful when rolling back to a previous application version. Similar to [`DBOS.set_latest_application_version`](./contexts.md#set_latest_application_version). **Parameters:** - `version_name`: The name of the version to promote. - `application_name`: The application to act as. Defaults to the client's own [`application_name`](#constructor). Promoting a version registered by a different application raises an error. #### set_latest_application_version_async ```python await client.set_latest_application_version_async( version_name: str, *, application_name: Optional[str] = None, ) -> None ``` Coroutine version of [`set_latest_application_version`](#set_latest_application_version). ### Application Rename #### rename_application ```python client.rename_application( old_name: Optional[str], new_name: str, *, batch_size: Optional[int] = 10_000, adopt_unclaimed_rows: bool = False, ) -> ApplicationRowCounts class ApplicationRowCounts(TypedDict): queues: int schedules: int versions: int workflows: int steps: int ``` Every workflow, step, queue, schedule, and application version is owned by the application (identified by its configured [`name`](./configuration.md#application-settings)) that created it. After renaming an application, use this method (or the [`dbos rename-application`](./cli.md#dbos-rename-application) CLI command) to transfer everything owned by the old name to the new name. Returns the number of rows transferred, by table. Queues, schedules, versions, and in-flight workflows are transferred in a single transaction; completed workflows and then all workflow steps are transferred in batches of `batch_size` workflows. The operation is idempotent: if interrupted, running it again resumes where it left off. :::warning Stop the application being renamed before running this. A running application would race the rename, creating new work under its old name. ::: **Parameters:** - `old_name`: The application's previous name. If `None`, nothing is transferred except rows owned by no application, so `adopt_unclaimed_rows` must be set. - `new_name`: The application that ends up owning the rows. Must be a valid application name (between 3 and 256 characters, containing only lowercase letters, numbers, dashes, and underscores). - `batch_size`: The number of workflows per batch when transferring completed workflows and steps. Pass `None` to transfer them without batching. - `adopt_unclaimed_rows`: Also transfer rows owned by no application, such as rows created before upgrading to a DBOS version supporting application ownership. Defaults to `False`. #### rename_application_async Coroutine version of [`rename_application`](#rename_application). --- ## Configuration(Reference) ### Configuring DBOS To configure DBOS, pass a `DBOSConfig` object to its constructor. For example: ```python config: DBOSConfig = { "name": "dbos-example", "application_version": "0.1.0", "system_database_url": os.environ["DBOS_SYSTEM_DATABASE_URL"], } DBOS(config=config) ``` The `DBOSConfig` object has the following fields. All fields except `name` are optional. ```python class DBOSConfig(TypedDict): name: str enable_patching: Optional[bool] application_version: Optional[str] executor_id: Optional[str] system_database_url: Optional[str] sys_db_pool_size: Optional[int] sys_db_polling_concurrency: Optional[int] db_engine_kwargs: Optional[Dict[str, Any]] dbos_system_schema: Optional[str] system_database_engine: Optional[sqlalchemy.Engine] use_listen_notify: Optional[bool] run_migrations: Optional[bool] notification_listener_polling_interval_sec: Optional[float] notification_coalesce_sec: Optional[float] observability_query_timeout_sec: Optional[float] sys_db_idle_transaction_timeout_sec: Optional[float] conductor_key: Optional[str] conductor_url: Optional[str] conductor_executor_metadata: Optional[Dict[str, Any]] conductor_metadata_only_mode: Optional[bool] enable_otlp: Optional[bool] otlp_traces_endpoints: Optional[List[str]] otlp_logs_endpoints: Optional[List[str]] otlp_attributes: Optional[dict[str, str]] otel_attribute_format: Optional[Literal["legacy", "semconv"]] log_level: Optional[str] otlp_log_level: Optional[str] console_log_level: Optional[str] max_executor_threads: Optional[int] scheduler_polling_interval_sec: Optional[float] kafka_queue_polling_interval_sec: Optional[float] serializer: Optional[Serializer] ``` #### Application Settings - **name**: Your application's name. It must be between 3 and 256 characters long and contain only lowercase letters, numbers, dashes, and underscores. Multiple applications (potentially in different languages) may [share a system database](../../explanations/sharing-a-system-database.md), in which case each must have a distinct name: the name identifies which application owns each workflow, queue, schedule, and application version, and applications only run their own workflows. If you rename an application, transfer ownership of its data with [`dbos rename-application`](./cli.md#dbos-rename-application). - **enable_patching**: Enable the [patching](../tutorials/upgrading-workflows.md#patching) strategy for safely upgrading workflow code. Required to use [`DBOS.patch`](./contexts.md#patch) and [`DBOS.deprecate_patch`](./contexts.md#deprecate_patch), which otherwise raise a `DBOSException`. - **application_version**: If using the [versioning](../tutorials/upgrading-workflows.md#versioning) strategy for safely upgrading workflow code, the code version for this application and its workflows. - **executor_id**: A unique process ID used to identify the application instance in distributed environments. If using DBOS Conductor or Cloud, this is set automatically. #### Database Connection Settings - **system_database_url**: A connection string to your system database. This is the database in which DBOS stores workflow and step state; its schema is documented [here](../../explanations/system-tables.md). This may be either Postgres or SQLite, though Postgres is recommended for production. DBOS uses this connection string to create a [SQLAlchemy engine](https://docs.sqlalchemy.org/en/20/core/engines.html). For Postgres, DBOS always connects with the `psycopg` (version 3) driver, replacing any driver specified in the connection string. A valid connection string looks like: ``` postgresql://[username]:[password]@[hostname]:[port]/[database name] ``` Or with SQLite: ``` sqlite:///[path to database file] ``` :::info Passwords in connection strings must be escaped (for example with [urllib](https://docs.python.org/3/library/urllib.parse.html#urllib.parse.quote)) if they contain special characters. ::: If no connection string is provided, DBOS uses a SQLite database (with any dashes in the application name replaced by underscores, and an underscore prepended if the name starts with a digit): ```shell sqlite:///[application_name].sqlite ``` - **sys_db_pool_size**: The size of the connection pool used for the [DBOS system database](../../explanations/system-tables). Defaults to 20. - **sys_db_polling_concurrency**: The maximum number of database-backed polling reads from wait operations (such as [`get_result`](./contexts.md#get_result), [`recv`](./contexts.md#recv), [`get_event`](./contexts.md#get_event), and [`read_stream`](./contexts.md#read_stream)) that may run concurrently against the system database pool. This prevents high-fan-out polling from checking out every connection in the pool and starving control-plane operations (such as enqueue/dequeue, status writes, recovery, and cancellation). Defaults to half the `sys_db_pool_size` (minimum 1). Set to a non-positive value to disable the limit. - **db_engine_kwargs**: A dictionary of additional keyword arguments passed to the SQLAlchemy [create_engine](https://docs.sqlalchemy.org/en/20/core/engines.html#sqlalchemy.create_engine) call. Can be used to customize connection pool settings, timeouts, and other engine parameters. - **dbos_system_schema**: Postgres schema name for DBOS system tables. Defaults to `dbos`. - **system_database_engine**: A custom SQLAlchemy engine to use to connect to your system database. If provided, DBOS will not create an engine but use this instead. - **use_listen_notify**: Whether to use PostgreSQL LISTEN/NOTIFY (`True`) or polling (`False`) to await notifications and events. Defaults to `True`. Ignored in SQLite, which always uses polling. On Postgres, this setting determines which notification triggers are created with the system database, so do not change it after the system database is first created. - **run_migrations**: Whether to create and migrate the system database on launch. Defaults to `True`. Set to `False` for a process that must not alter the schema, such as one whose database role cannot run DDL, or a deployment that migrates out of band with [`dbos migrate`](./cli.md#dbos-migrate) or [`DBOS.migrate`](./dbos-class.md#migrate). Launch then verifies the schema instead of changing it: a system database whose DBOS tables are missing (including a SQLite file that does not exist) or behind the version this build of DBOS requires fails launch with a `DBOSInitializationError`, and a Postgres database that does not exist fails launch with a connection error. A system database ahead of the required version is accepted, so a process with migrations disabled can run alongside newer peers. - **notification_listener_polling_interval_sec**: Polling interval in seconds for the notification listener background process. Defaults to `1.0`; the minimum is `0.001`. Used when polling (when `use_listen_notify` is `False` or the system database is SQLite), and as the default `polling_interval_sec` of [`read_stream`](./contexts.md#read_stream) and [`read_stream_offset`](./contexts.md#read_stream_offset). - **notification_coalesce_sec**: Interval in seconds at which DBOS batches and sends the LISTEN/NOTIFY notifications that wake readers of [events](./contexts.md#get_event) and [streams](./contexts.md#read_stream). This bounds how long a waiting reader may be delayed and caps the rate of notifying commits regardless of write throughput. Defaults to `0.01`; the minimum is `0.001`. Only used on Postgres when `use_listen_notify` is `True`. - **observability_query_timeout_sec**: The statement timeout, in seconds, applied to observability queries (such as listing workflows, queued workflows, and workflow steps) on a Postgres system database, so a slow query on a large database does not hold resources indefinitely. A query that exceeds the timeout raises `DBOSQueryTimeoutError`. Defaults to 30 seconds. Set to zero or a negative value to disable the timeout. - **sys_db_idle_transaction_timeout_sec**: The Postgres `idle_in_transaction_session_timeout`, in seconds, set on the system database connections DBOS creates. Defaults to 60 seconds. #### Conductor Settings - **conductor_key**: An API key for [DBOS Conductor](../../conductor/overview.md). If provided, application connects to Conductor. API keys can be created from the [DBOS Console](https://console.dbos.dev). - **conductor_url**: The URL of the Conductor service to connect to. Only set if you are self-hosting Conductor. - **conductor_executor_metadata**: A JSON-serializable dictionary of metadata to associate with this executor. This metadata is sent to Conductor and displayed on the dashboard, making it easier to identify executors (e.g., by region, instance type, or deployment environment). - **conductor_metadata_only_mode**: If `True`, this process sends only workflow metadata to Conductor, never workflow data (inputs, outputs, errors, step outputs, events, messages, streams, or schedule context), regardless of the [metadata-only mode](../../conductor/overview.md#metadata-only-mode) setting in the Conductor console. Defaults to `False`. #### Logging and Tracing Settings - **enable_otlp**: Enable DBOS OpenTelemetry [tracing](../tutorials/logging-and-tracing.md), which makes DBOS create spans for all workflows and steps. To export those spans, either set `otlp_traces_endpoints`/`otlp_logs_endpoints` (DBOS runs its own built-in `TracerProvider`) or [register your own `TracerProvider`](../tutorials/logging-and-tracing.md#connecting-dbos-to-your-observability-provider) before launch and let your observability provider export them. Defaults to False. - **otlp_traces_endpoints**: If using the built-in DBOS OpenTelemetry `TracerProvider`, a list of receivers to which to send traces. - **otlp_logs_endpoints**: If using the built-in DBOS OpenTelemetry `TracerProvider`, a list of receivers to which to send logs. - **otlp_attributes**: A set of attributes (key-value pairs) to apply to all OTLP-exported logs and traces. - **otel_attribute_format**: Naming convention for DBOS-emitted span attributes. Defaults to `"legacy"`, which emits the original camelCase names (`operationUUID`, `executorID`, …) for backward compatibility. Set to `"semconv"` to emit OTel-style names under the `dbos.*` namespace (`dbos.operation.workflow_id`, `dbos.executor.id`, …), which follow the [OTel attribute naming spec](https://opentelemetry.io/docs/specs/semconv/general/attribute-naming/) and avoid colliding with attributes set by other instrumentation. The flag is process-wide; user-supplied `otlp_attributes` are passed through verbatim either way. - **log_level**: Configure the [DBOS logger](../tutorials/logging-and-tracing#logging) severity. Defaults to `INFO`. - **otlp_log_level**: Log level specifically for OTLP logging (if enabled). Must be no less severe than `log_level`. Defaults to the value of `log_level`. - **console_log_level**: Log level specifically for console logging. Must be no less severe than `log_level`. Defaults to the value of `log_level`. #### Execution Settings - **max_executor_threads**: The maximum number of threads in the executor thread pool used for running synchronous workflow and step functions. If unset, the pool is unbounded. #### Scheduler Settings - **scheduler_polling_interval_sec**: Polling interval in seconds for the scheduler thread to detect new [workflow schedules](./contexts.md#workflow-schedules). Defaults to `30.0`. #### Kafka Settings - **kafka_queue_polling_interval_sec**: Polling interval in seconds for the internal queues on which [Kafka consumer](../tutorials/kafka-integration.md) workflows run. Defaults to `1.0`; the minimum is `0.001`. Lowering it reduces the latency between a message being enqueued and its workflow starting, but requires more frequent database polling. #### Serialization Settings - **serializer**: A custom serializer for the system database. See the [custom serialization reference](./contexts.md#custom-serialization) for details. ### DBOS Configuration File Some tools in the DBOS ecosystem, including [DBOS Cloud](../../conductor/reference/dbos-cloud/deploying-to-cloud.md) and the [DBOS CLI](./cli.md), are configured by a `dbos-config.yaml` file. You can create a `dbos-config.yaml` with default parameters with: ```shell dbos init --config ``` #### Configuration File Fields ::::info You can use environment variables for configuration values through the syntax `field: ${VALUE}`. :::: Each `dbos-config.yaml` file has the following fields and sections: - **name**: Your application's name. Must match the name supplied to the DBOS constructor. - **language**: The application language. Must be set to `python` for Python applications. - **system_database_url**: The connection string to your DBOS system database. This connection string is used by the DBOS [CLI](cli.md). It has the same format as the `system_database_url` you pass to the DBOS constructor. - **runtimeConfig**: - **start**: (required only in DBOS Cloud) The command(s) with which to start your app. Called from [`dbos start`](../reference/cli.md#dbos-start), which is used to start your app in DBOS Cloud. - **setup**: Setup commands to run before your application is built in DBOS Cloud. Used only in DBOS Cloud. Documentation [here](../../conductor/reference/dbos-cloud/application-management.md#customizing-microvm-setup). #### Configuration Schema File There is a schema file available for the DBOS configuration file schema [on GitHub](https://github.com/dbos-inc/dbos-transact-py/blob/main/dbos/dbos-config.schema.json). This schema file can be used to provide an improved YAML editing experience for developer tools that leverage it. For example, the Visual Studio Code [RedHat YAML extension](https://marketplace.visualstudio.com/items?itemName=redhat.vscode-yaml) provides tooltips, statement completion and real-time validation for editing DBOS config files. This extension provides [multiple ways](https://github.com/redhat-developer/vscode-yaml#associating-schemas) to associate a YAML file with its schema. The easiest is to simply add a comment with a link to the schema at the top of the config file: ```yaml # yaml-language-server: $schema=https://raw.githubusercontent.com/dbos-inc/dbos-transact-py/main/dbos/dbos-config.schema.json ``` --- ## DBOS Methods & Variables(3) DBOS provides a number of useful context methods and variables. All are accessed through the syntax `DBOS.` and can only be used once a DBOS class object has been initialized. ### Context Methods #### start_workflow ```python DBOS.start_workflow( func: Callable[P, R], *args: P.args, **kwargs: P.kwargs, ) -> WorkflowHandle[R] ``` Start a workflow in the background and return a [handle](./workflow_handles.md) to it. The `DBOS.start_workflow` method resolves after the handle is durably created; at this point the workflow is guaranteed to run to completion even if the app is interrupted. **Example syntax:** ```python @DBOS.workflow() def example_workflow(var1: str, var2: str): DBOS.logger.info("I am a workflow") # Start example_workflow in the background handle: WorkflowHandle = DBOS.start_workflow(example_workflow, "var1", "var2") ``` #### start_workflow_async ```python DBOS.start_workflow_async( func: Callable[P, Coroutine[Any, Any, R]], *args: P.args, **kwargs: P.kwargs, ) -> Coroutine[Any, Any, WorkflowHandleAsync[R]] ``` Start an asynchronous workflow and return a [handle](./workflow_handles.md) to it. The `DBOS.start_workflow_async` method resolves after the handle is durably created; at this point the workflow is guaranteed to run to completion even if the app is interrupted. The workflow started with `DBOS.start_workflow_async` runs in the same event loop as its caller. **Example syntax:** ```python @DBOS.workflow() async def example_workflow(var1: str, var2: str): DBOS.logger.info("I am a workflow") # Start example_workflow handle: WorkflowHandleAsync = await DBOS.start_workflow_async(example_workflow, "var1", "var2") ``` #### wait_first ```python DBOS.wait_first( handles: List[WorkflowHandle[Any]], *, polling_interval_sec: float = 1.0, ) -> WorkflowHandle[Any] ``` Wait for any one of the given workflow handles to complete and return the first completed handle. This is useful when you have multiple concurrent workflows and want to process results as they complete. **Parameters:** - **handles**: A non-empty list of workflow handles to wait on. Raises `ValueError` if the list is empty or contains duplicate workflow IDs. - **polling_interval_sec**: The interval (in seconds) at which DBOS polls the database. Defaults to `1.0`. See the [queue tutorial](../tutorials/queue-tutorial.md#queue-example) for an example. #### wait_first_async ```python DBOS.wait_first_async( handles: List[WorkflowHandleAsync[Any]], *, polling_interval_sec: float = 1.0, ) -> Coroutine[Any, Any, WorkflowHandleAsync[Any]] ``` Async version of [`wait_first`](#wait_first). Wait for any one of the given async workflow handles to complete and return the first completed handle. #### send ```python DBOS.send( destination_id: str, message: Any, topic: Optional[str] = None, *, idempotency_key: Optional[str] = None, serialization_type: Optional[WorkflowSerializationFormat] = WorkflowSerializationFormat.DEFAULT, send_to_forks: bool = False, ) -> None ``` Send a message to the workflow identified by `destination_id`. Messages can optionally be associated with a topic. The `send` function should not be used in [coroutine workflows](../tutorials/workflow-tutorial.md#coroutine-async-workflows), [`send_async`](#send_async) should be used instead. **Parameters:** - `destination_id`: The workflow to which to send the message. - `message`: The message to send. Must be serializable. - `topic`: A topic with which to associate the message. Messages are enqueued per-topic on the receiver. - `idempotency_key`: If an idempotency key is set, the message will only be sent once to each destination no matter how many times `DBOS.send` is called with this key. The key is scoped per destination workflow. - `serialization_type`: The [serialization format](#serialization-strategy) to use for this message. Defaults to `WorkflowSerializationFormat.DEFAULT`. - `send_to_forks`: If `True`, also deliver the message to every workflow recursively [forked](#fork_workflow) from `destination_id` (forks, forks of forks, and so on) that exists at send time. Defaults to `False`. #### send_async ```python DBOS.send_async( destination_id: str, message: Any, topic: Optional[str] = None, *, idempotency_key: Optional[str] = None, serialization_type: Optional[WorkflowSerializationFormat] = WorkflowSerializationFormat.DEFAULT, send_to_forks: bool = False, ) -> Coroutine[Any, Any, None] ``` Coroutine version of [`send`](#send) #### send_bulk ```python DBOS.send_bulk( messages: List[SendMessage], *, serialization_type: Optional[WorkflowSerializationFormat] = WorkflowSerializationFormat.DEFAULT, send_to_forks: bool = False, ) -> None ``` Send many messages to workflow executions in a single transaction. Each message is described by a `SendMessage` object specifying its destination, payload, and optional topic and idempotency key: ```python @dataclass class SendMessage: # The workflow to which to send the message destination_id: str # The message to send. Must be serializable. message: Any # A topic with which to associate the message. Messages are enqueued per-topic on the receiver. topic: Optional[str] = None # If set, the message is sent only once per destination no matter how many times it is submitted with this key. idempotency_key: Optional[str] = None ``` The send is atomic: if any message cannot be delivered (for example, its destination workflow does not exist), the entire batch is rolled back and no messages are sent. The `send_bulk` function should not be used in [coroutine workflows](../tutorials/workflow-tutorial.md#coroutine-async-workflows), [`send_bulk_async`](#send_bulk_async) should be used instead. **Parameters:** - `messages`: The list of `SendMessage` objects to send. Two messages in the same call may not share an idempotency key. - `serialization_type`: The [serialization format](#serialization-strategy) to use for these messages. Defaults to `WorkflowSerializationFormat.DEFAULT`. - `send_to_forks`: If `True`, every message is also delivered to all workflows recursively [forked](#fork_workflow) from its destination. Defaults to `False`. #### send_bulk_async ```python DBOS.send_bulk_async( messages: List[SendMessage], *, serialization_type: Optional[WorkflowSerializationFormat] = WorkflowSerializationFormat.DEFAULT, send_to_forks: bool = False, ) -> Coroutine[Any, Any, None] ``` Coroutine version of [`send_bulk`](#send_bulk) #### recv ```python DBOS.recv( topic: Optional[str] = None, timeout_seconds: float = 60, ) -> Any ``` Receive and return a message sent to this workflow. Can only be called from within a workflow. Messages are dequeued first-in, first-out from a queue associated with the topic. Calls to `recv` wait for the next message in the queue, returning `None` if the wait times out. If no topic is specified, `recv` can only access messages sent without a topic. The `recv` function should not be used in [coroutine workflows](../tutorials/workflow-tutorial.md#coroutine-async-workflows), [`recv_async`](#recv_async) should be used instead. **Parameters:** - `topic`: A topic queue on which to wait. - `timeout_seconds`: A timeout in seconds. If the wait times out, return `None`. **Returns:** - The first message enqueued on the input topic, or `None` if the wait times out. #### recv_async ```python DBOS.recv_async( topic: Optional[str] = None, timeout_seconds: float = 60, ) -> Coroutine[Any, Any, Any] ``` Coroutine version of [`recv`](#recv) #### set_event ```python DBOS.set_event( key: str, value: Any, *, serialization_type: WorkflowSerializationFormat = WorkflowSerializationFormat.DEFAULT, ) -> None ``` Create and associate with this workflow an event with key `key` and value `value`. If the event already exists, update its value. Can only be called from within a workflow or its steps. The `set_event` function should not be used in [coroutine workflows](../tutorials/workflow-tutorial.md#coroutine-async-workflows), `set_event_async` should be used instead. **Parameters:** - `key`: The key of the event. - `value`: The value of the event. Must be serializable. - `serialization_type`: The [serialization format](#serialization-strategy) to use for this event. Defaults to `WorkflowSerializationFormat.DEFAULT`. #### set_event_async ```python DBOS.set_event_async( key: str, value: Any, *, serialization_type: WorkflowSerializationFormat = WorkflowSerializationFormat.DEFAULT, ) -> Coroutine[Any, Any, None] ``` Coroutine version of [`set_event`](#set_event) #### get_event ```python DBOS.get_event( workflow_id: str, key: str, timeout_seconds: float = 60, ) -> Any ``` Retrieve the latest value of an event published by the workflow identified by `workflow_id` to the key `key`. If the event does not yet exist, wait for it to be published, returning `None` if the wait times out. The `get_event` function should not be used in [coroutine workflows](../tutorials/workflow-tutorial.md#coroutine-async-workflows), [`get_event_async`](#get_event_async) should be used instead. **Parameters:** - `workflow_id`: The identifier of the workflow whose events to retrieve. - `key`: The key of the event to retrieve. - `timeout_seconds`: A timeout in seconds. If the wait times out, return `None`. **Returns:** - The value of the event published by `workflow_id` with name `key`, or `None` if the wait times out. #### get_event_async ```python DBOS.get_event_async( workflow_id: str, key: str, timeout_seconds: float = 60, ) -> Coroutine[Any, Any, Any] ``` Coroutine version of [`get_event`](#get_event) #### get_all_events ```python DBOS.get_all_events( workflow_id: str ) -> Dict[str, Any] ``` Retrieve the latest values of all events published by `workflow_id`. - `workflow_id`: The identifier of the workflow whose events to retrieve. #### get_all_events_async ```python DBOS.get_all_events_async( workflow_id: str ) -> Coroutine[Any, Any, Dict[str, Any]] ``` Coroutine version of [`get_all_events`](#get_all_events). #### sleep ```python DBOS.sleep( seconds: float ) -> None ``` Sleep for the given number of seconds. May only be called from within a workflow. This sleep is durable—it records its intended wake-up time in the database so if it is interrupted and recovers, it still wakes up at the intended time. The `sleep` function should not be used in [coroutine workflows](../tutorials/workflow-tutorial.md#coroutine-async-workflows), [`sleep_async`](#sleep_async) should be used instead. **Parameters:** - `seconds`: The number of seconds to sleep. #### sleep_async ```python DBOS.sleep_async( seconds: float ) -> Coroutine[Any, Any, None] ``` Coroutine version of [`sleep`](#sleep) #### asyncio_wait ```python DBOS.asyncio_wait( fs: List[Awaitable[Any]], *, timeout: Optional[float] = None, return_when: str = asyncio.ALL_COMPLETED, ) -> Coroutine[Any, Any, tuple[set[asyncio.Task[Any]], set[asyncio.Task[Any]]]] ``` A durable wrapper around [`asyncio.wait`](https://docs.python.org/3/library/asyncio-task.html#asyncio.wait) with the same interface and semantics. It checkpoints which futures are done vs. pending so the result is deterministic during workflow recovery. When called outside a workflow, it falls back to regular `asyncio.wait`. **Parameters:** - **fs**: A list of awaitables (coroutines, tasks, or futures) to wait on. - **timeout**: Maximum number of seconds to wait. If `None` (the default), wait until the `return_when` condition is met. - **return_when**: Controls when the function returns. Must be one of the following constants: - `asyncio.FIRST_COMPLETED`: The function will return when any future finishes or is cancelled. - `asyncio.FIRST_EXCEPTION`: The function will return when any future finishes by raising an exception. If no future raises an exception then it is equivalent to `ALL_COMPLETED`. - `asyncio.ALL_COMPLETED`: The function will return when all futures finish or are cancelled. This is the default. **Returns:** Two sets of Tasks/Futures: `(done, pending)`. The `done` set contains futures that completed (finished or were cancelled) before the function returned. The `pending` set contains futures that are still running. See the [`asyncio.wait` documentation](https://docs.python.org/3/library/asyncio-task.html#asyncio.wait) for full details. #### run_step ```python DBOS.run_step( dbos_step_options: Optional[StepOptions], func: Callable[P, R], *args: P.args, **kwargs: P.kwargs, ) -> R: ``` Runs the provided `func` function (or lambda) as a checkpointed DBOS [step](../tutorials/step-tutorial.md). `args` and `kwargs` will be passed to `func`. The `StepOptions` object has the following fields. All fields are optional. ```python class StepOptions(TypedDict, total=False): """ Configuration options for steps. Attributes: name: Optional name for the step. If not provided, the function's name will be used. retries_allowed: Whether the step should be retried on failure. interval_seconds: Initial delay (in seconds) between retry attempts. max_attempts: Maximum number of attempts before the step is considered failed. backoff_rate: Multiplier applied to `interval_seconds` after each failed attempt (e.g. 2.0 = exponential backoff). should_retry: Optional predicate called with a raised exception to decide whether the step should be retried. If it returns False (or an awaitable resolving to False), the exception is re-raised immediately without further retries. Async validators are only supported for async steps. preemptible: If True, cancel the (async) step if its workflow is cancelled. Only supported for async steps. timeout_seconds: If set, cancel the (async) step and raise DBOSStepTimeoutError if it runs for longer than this many seconds. Only supported for async steps. Each retry attempt gets a fresh timeout. Inert outside a workflow, where the step runs as a normal function call. """ name: Optional[str] retries_allowed: bool interval_seconds: float max_attempts: int backoff_rate: float should_retry: Optional[Callable[[BaseException], Union[bool, Awaitable[bool]]]] preemptible: bool timeout_seconds: Optional[float] ``` #### run_step_async Version of [`run_step`](#run_step) to be called from `async` contexts. #### get_result ```python DBOS.get_result( workflow_id: str, ) -> Optional[Any] ``` Wait for the workflow identified by `workflow_id` to complete, and return its result. This is similar to calling [`get_result`](./workflow_handles.md#get_result) on a [WorkflowHandle](./workflow_handles.md), but is a single step that does not require a handle. **Parameters:** - `workflow_id`: The identifier of the workflow whose result to return. **Returns:** - The result of the workflow, or throws an exception if the workflow threw an exception. #### get_result_async ```python DBOS.get_result_async( workflow_id: str, ) -> Coroutine[Any, Any, Optional[Any]] ``` Coroutine version of [`get_result`](#get_result). #### get_workflow_status ```python DBOS.get_workflow_status( workflow_id: str, ) -> Optional[WorkflowStatus] ``` Retrieve the status of a workflow by its ID. Returns `None` if no workflow with the given ID exists. **Parameters:** - `workflow_id`: The identifier of the workflow whose status to retrieve. **Returns:** - The [`WorkflowStatus`](#workflow-status) of the workflow, or `None` if not found. #### get_workflow_status_async ```python DBOS.get_workflow_status_async( workflow_id: str, ) -> Coroutine[Any, Any, Optional[WorkflowStatus]] ``` Coroutine version of [`get_workflow_status`](#get_workflow_status). #### retrieve_workflow ```python DBOS.retrieve_workflow( workflow_id: str, existing_workflow: bool = True, ) -> WorkflowHandle[R] ``` Retrieve the [handle](./workflow_handles.md) of a workflow with identity `workflow_id`. **Parameters:** - `workflow_id`: The identifier of the workflow whose handle to retrieve. - `existing_workflow`: Whether to throw an exception (`DBOSNonExistentWorkflowError`) if the workflow does not yet exist. If set to `False`, return a handle immediately without checking whether the workflow exists; calling `get_result` on the handle waits for the workflow to be created and complete. **Returns:** - The [handle](./workflow_handles.md) of the workflow whose ID is `workflow_id`. #### retrieve_workflow_async ```python DBOS.retrieve_workflow_async( workflow_id: str, existing_workflow: bool = True, ) -> Coroutine[Any, Any, WorkflowHandleAsync[R]] ``` Coroutine version of [`DBOS.retrieve_workflow`](#retrieve_workflow), retrieving an async workflow handle. #### write_stream ```python DBOS.write_stream( key: str, value: Any, *, serialization_type: WorkflowSerializationFormat = WorkflowSerializationFormat.DEFAULT, ) -> None ``` Write a value to a stream. Can only be called from within a workflow or its steps. The `write_stream` function should not be used in [coroutine workflows](../tutorials/workflow-tutorial.md#coroutine-async-workflows), [`write_stream_async`](#write_stream_async) should be used instead. **Parameters:** - `key`: The stream key / name within the workflow - `value`: A serializable value to write to the stream - `serialization_type`: The [serialization format](#serialization-strategy) to use for this value. Defaults to `WorkflowSerializationFormat.DEFAULT`. #### write_stream_async ```python DBOS.write_stream_async( key: str, value: Any, *, serialization_type: WorkflowSerializationFormat = WorkflowSerializationFormat.DEFAULT, ) -> Coroutine[Any, Any, None] ``` Coroutine version of [`write_stream`](#write_stream) #### close_stream ```python DBOS.close_stream( key: str ) -> None ``` Close a stream identified by a key. After this is called, readers stop at the close, so any value written to the stream afterward is never read. Can only be called from within a workflow or its steps. The `close_stream` function should not be used in [coroutine workflows](../tutorials/workflow-tutorial.md#coroutine-async-workflows), [`close_stream_async`](#close_stream_async) should be used instead. **Parameters:** - `key`: The stream key / name within the workflow #### close_stream_async ```python DBOS.close_stream_async( key: str ) -> Coroutine[Any, Any, None] ``` Coroutine version of [`close_stream`](#close_stream) #### read_stream ```python DBOS.read_stream( workflow_id: str, key: str, *, offset: int = 0, polling_interval_sec: Optional[float] = None, timeout_seconds: Optional[float] = None, ) -> Generator[Any, Any, None] ``` Read values from a stream as a generator. This function reads values from a stream identified by the workflow_id and key, yielding each value in order until the stream is closed or the workflow terminates. **Parameters:** - `workflow_id`: The workflow instance ID that owns the stream - `key`: The stream key / name within the workflow - `offset`: The offset to start reading from. Defaults to `0`, the start of the stream. A higher offset skips that many values from the beginning of the stream. - `polling_interval_sec`: Polling interval in seconds when waiting for new values when not using LISTEN/NOTIFY. Must be at least `0.001`. Defaults to the configured [`notification_listener_polling_interval_sec`](./configuration.md#database-connection-settings) (`1.0` if not configured). - `timeout_seconds`: How long to wait for **each** value before raising `DBOSStreamTimeoutError`. The clock restarts every time a value is delivered, so this bounds the gap between values, not the total duration of the read. Defaults to `None`, waiting indefinitely. **Yields:** - Each value in the stream until the stream is closed **Raises:** - `DBOSStreamTimeoutError`: If `timeout_seconds` passes without a value arriving. - `DBOSNonExistentWorkflowError`: If no workflow with ID `workflow_id` exists. **Example syntax:** ```python for value in DBOS.read_stream(workflow_id, example_key): print(f"Received: {value}") ``` When called from workflow code, each value read is checkpointed to the database as a step, so a replayed workflow re-yields the values it originally read instead of re-reading a stream that may have advanced since. #### read_stream_async ```python DBOS.read_stream_async( workflow_id: str, key: str, *, offset: int = 0, polling_interval_sec: Optional[float] = None, timeout_seconds: Optional[float] = None, ) -> AsyncGenerator[Any, None] ``` Coroutine version of [`read_stream`](#read_stream), returning an async generator. **Example syntax:** ```python async for value in DBOS.read_stream_async(workflow_id, example_key): print(f"Received: {value}") ``` #### read_stream_offset ```python DBOS.read_stream_offset( workflow_id: str, key: str, offset: int, *, polling_interval_sec: Optional[float] = None, timeout_seconds: Optional[float] = None, ) -> Any ``` Read the single value at one offset of a stream, waiting for it to be written. Use this when you want one specific value instead of iterating the whole stream—for example, to resume from where a previous reader left off. **Parameters:** - `workflow_id`: The workflow instance ID that owns the stream - `key`: The stream key / name within the workflow - `offset`: The offset to read - `polling_interval_sec`: Polling interval in seconds when waiting for the value when not using LISTEN/NOTIFY. Must be at least `0.001`. Defaults to the configured [`notification_listener_polling_interval_sec`](./configuration.md#database-connection-settings) (`1.0` if not configured). - `timeout_seconds`: How long to wait for the value before raising `DBOSStreamTimeoutError`. Defaults to `None`, waiting indefinitely. **Returns:** - The value at the offset **Raises:** - `DBOSStreamTimeoutError`: If `timeout_seconds` passes, or if the stream ends before reaching `offset` (no value will ever arrive at that offset). - `DBOSNonExistentWorkflowError`: If no workflow with ID `workflow_id` exists. **Example syntax:** ```python from dbos import error as dboserror try: value = DBOS.read_stream_offset(workflow_id, example_key, 5, timeout_seconds=30) except dboserror.DBOSStreamTimeoutError: ... ``` Like [`read_stream`](#read_stream), a value read from workflow code is checkpointed to the database as a step. #### read_stream_offset_async ```python DBOS.read_stream_offset_async( workflow_id: str, key: str, offset: int, *, polling_interval_sec: Optional[float] = None, timeout_seconds: Optional[float] = None, ) -> Coroutine[Any, Any, Any] ``` Coroutine version of [`read_stream_offset`](#read_stream_offset). #### patch ```python DBOS.patch( patch_name: str ) -> bool ``` Insert a patch marker at the current point in workflow history, returning `True` if it was successfully inserted (or this patch marker is already present) and `False` if a different checkpoint is already present at this point in history. Used to safely upgrade workflow code, see the [patching tutorial](../tutorials/upgrading-workflows.md#patching) for more detail. Requires [`enable_patching`](./configuration.md#application-settings) to be set in your DBOS configuration. The `patch` function should not be used in [coroutine workflows](../tutorials/workflow-tutorial.md#coroutine-async-workflows), [`patch_async`](#patch_async) should be used instead. **Parameters:** - `patch_name`: The name to give the patch marker that will be inserted into workflow history. #### patch_async ```python DBOS.patch_async( patch_name: str ) -> Coroutine[Any, Any, bool] ``` Coroutine version of [`DBOS.patch()`](#patch). #### deprecate_patch ```python DBOS.deprecate_patch( patch_name: str ) -> bool ``` Safely bypass a patch marker at the current point in workflow history if present. Always returns `True`. Used to safely deprecate patches, see the [patching tutorial](../tutorials/upgrading-workflows.md#patching) for more detail. Requires [`enable_patching`](./configuration.md#application-settings) to be set in your DBOS configuration. The `deprecate_patch` function should not be used in [coroutine workflows](../tutorials/workflow-tutorial.md#coroutine-async-workflows), [`deprecate_patch_async`](#deprecate_patch_async) should be used instead. **Parameters:** - `patch_name`: The name of the patch marker to be bypassed. #### deprecate_patch_async ```python DBOS.deprecate_patch_async( patch_name: str ) -> Coroutine[Any, Any, bool] ``` Coroutine version of [`DBOS.deprecate_patch()`](#deprecate_patch) ### Queue Management Queues are persisted to the system database, so any DBOS process or [`DBOSClient`](./client.md) connected to the same system database can register, retrieve, and enqueue workflows on them. #### register_queue ```python DBOS.register_queue( name: str, *, # Applied to the queue as a whole global_concurrency: Optional[int] = None, worker_concurrency: Optional[int] = None, limiter: Optional[QueueRateLimit] = None, # Applied to each partition separately partition_concurrency: Optional[int] = None, partition_worker_concurrency: Optional[int] = None, partition_limiter: Optional[QueueRateLimit] = None, polling_interval_sec: float = 1.0, on_conflict: QueueConflictResolution = "update_if_latest_version", ) -> Queue QueueConflictResolution = Literal[ "update_if_latest_version", "always_update", "never_update" ] ``` Register a [queue](./queues.md) and persist its configuration to the system database, returning the [`Queue`](./queues.md#class-dbosqueue). DBOS must be launched before calling `register_queue`. If the queue already exists in the database, the `on_conflict` parameter controls whether its configuration is overwritten. **Parameters:** - `name`: The name of the queue. Must be unique among all queues in the system database, including those of other applications sharing it. Names starting with `_dbos_` are reserved for DBOS. - `global_concurrency`: The maximum number of functions from this queue that may run concurrently across all DBOS processes. If not provided, any number of functions may run concurrently. - `worker_concurrency`: The maximum number of functions from this queue that may run concurrently on a single DBOS process. Must be less than or equal to `global_concurrency`. - `limiter`: A limit on the maximum number of functions which may be started in a given period. - `partition_concurrency`: The maximum number of functions from any one [partition](../tutorials/queue-tutorial.md#partitioning-queues) of this queue that may run concurrently across all DBOS processes. Must be at least 1 and less than or equal to `global_concurrency`. - `partition_worker_concurrency`: The maximum number of functions from any one partition of this queue that may run concurrently on a single DBOS process. Must be at least 1 and less than or equal to `partition_concurrency`, `worker_concurrency`, and `global_concurrency`. - `partition_limiter`: A limit on the maximum number of functions which may be started from any one partition in a given period. - `polling_interval_sec`: The minimum interval at which DBOS polls the database for new workflows on this queue. The actual interval includes random jitter and increases with backoff under contention, then scales back down when contention clears. - `on_conflict`: How to behave when a queue with this name already exists in the system database: - `"update_if_latest_version"` (default): overwrite the existing configuration only if the running application is the latest registered [application version](#version-management). This prevents older versions in a rolling deploy from overwriting a newer configuration. - `"always_update"`: always overwrite the existing configuration. - `"never_update"`: leave the existing configuration unchanged. Setting any `partition_*` limit makes the queue [partitioned](../tutorials/queue-tutorial.md#partitioning-queues): every enqueue must supply a [`queue_partition_key`](./queues.md#setenqueueoptions), and [deduplication](../tutorials/queue-tutorial.md#deduplication) is not supported. The queue-wide limits (`global_concurrency`, `worker_concurrency`, and `limiter`) continue to apply across all partitions. Queues are owned by the application (identified by its configured [`name`](./configuration.md#application-settings)) that registers them, and queue names are globally unique across all applications sharing a system database, so registering a queue whose name is owned by a different application raises an error regardless of `on_conflict`. **Example syntax:** ```python DBOS.register_queue("email", global_concurrency=10, limiter={"limit": 100, "period": 60}) ``` #### register_queue_async ```python DBOS.register_queue_async( name: str, *, global_concurrency: Optional[int] = None, worker_concurrency: Optional[int] = None, limiter: Optional[QueueRateLimit] = None, partition_concurrency: Optional[int] = None, partition_worker_concurrency: Optional[int] = None, partition_limiter: Optional[QueueRateLimit] = None, polling_interval_sec: float = 1.0, on_conflict: QueueConflictResolution = "update_if_latest_version", ) -> Coroutine[Any, Any, Queue] ``` Coroutine version of [`register_queue`](#register_queue). #### retrieve_queue ```python DBOS.retrieve_queue(name: str) -> Optional[Queue] ``` Retrieve a queue by name from the system database, or `None` if no queue with that name has been registered. **Example syntax:** ```python queue = DBOS.retrieve_queue("email") if queue is not None: print(queue.global_concurrency) ``` #### retrieve_queue_async ```python DBOS.retrieve_queue_async(name: str) -> Coroutine[Any, Any, Optional[Queue]] ``` Coroutine version of [`retrieve_queue`](#retrieve_queue). #### list_queues ```python DBOS.list_queues( *, application_name: Optional[Union[str, List[str]]] = None, ) -> List[Queue] ``` List all database-backed queues registered in the system database. Returns an empty list if no queues have been registered. **Parameters:** - `application_name`: List only queues owned by this application (or one of these applications). Queues owned by no application are always included. If unset, list only this application's queues. **Example syntax:** ```python for queue in DBOS.list_queues(): print(queue.name, queue.global_concurrency) ``` #### list_queues_async ```python DBOS.list_queues_async( *, application_name: Optional[Union[str, List[str]]] = None, ) -> Coroutine[Any, Any, List[Queue]] ``` Coroutine version of [`list_queues`](#list_queues). #### enqueue_workflow ```python DBOS.enqueue_workflow( queue_name: str, func: Callable[P, R], *args: P.args, **kwargs: P.kwargs, ) -> WorkflowHandle[R] ``` Enqueue a workflow on a queue and return a [handle](./workflow_handles.md) to it. Equivalent to retrieving the queue by name and calling [`Queue.enqueue`](./queues.md#enqueue) on it. The queue does not need to be registered at the time of the call: if no queue with `queue_name` exists yet, the workflow is durably recorded as `ENQUEUED` and starts running once the queue is registered and a worker becomes available. **Example syntax:** ```python @DBOS.workflow() def send_email(to: str) -> None: ... DBOS.register_queue("email", global_concurrency=10) handle = DBOS.enqueue_workflow("email", send_email, "alice@example.com") handle.get_result() ``` #### enqueue_workflow_async ```python DBOS.enqueue_workflow_async( queue_name: str, func: Callable[P, Coroutine[Any, Any, R]], *args: P.args, **kwargs: P.kwargs, ) -> Coroutine[Any, Any, WorkflowHandleAsync[R]] ``` Coroutine version of [`enqueue_workflow`](#enqueue_workflow). #### enqueue_workflow_with_options ```python DBOS.enqueue_workflow_with_options( options: EnqueueOptions, *args: Any, **kwargs: Any, ) -> WorkflowHandle[Any] ``` Enqueue a workflow by name, without a reference to its function. Takes the same [`EnqueueOptions`](./client.md#enqueue) as `DBOSClient.enqueue`, so the workflow may be implemented by another process or another application, as long as it shares this system database. Can safely be called from inside a workflow: the enqueued workflow is recorded as a child of the calling workflow. Unlike [`enqueue_workflow`](#enqueue_workflow), options are not validated against the local registry, and `app_version` is left unset unless given (an unset `app_version` is only dequeued by an executor running the latest registered application version). The enqueued workflow is owned by this application unless the `application_name` option names another one, in which case that application dequeues and runs it. #### enqueue_workflow_with_options_async ```python DBOS.enqueue_workflow_with_options_async( options: EnqueueOptions, *args: Any, **kwargs: Any, ) -> Coroutine[Any, Any, WorkflowHandleAsync[Any]] ``` Coroutine version of [`enqueue_workflow_with_options`](#enqueue_workflow_with_options). #### delete_queue ```python DBOS.delete_queue(name: str) -> None ``` Delete a queue from the system database. No-op if no queue with that name exists. :::warning Workflows already enqueued on a deleted queue can no longer be dequeued, executed, or recovered. However, if a queue with the same name is later registered, it will dequeue the leftover workflows. Do not rely on this: stale workflows unexpectedly resuming on a future queue is rarely the intended behavior. Instead, cancel or drain pending workflows on the queue before deleting it. ::: #### delete_queue_async ```python DBOS.delete_queue_async(name: str) -> Coroutine[Any, Any, None] ``` Coroutine version of [`delete_queue`](#delete_queue). ### Workflow Management Methods #### list_workflows ```python def list_workflows( *, workflow_ids: Optional[List[str]] = None, status: Optional[Union[str, List[str]]] = None, start_time: Optional[str] = None, end_time: Optional[str] = None, completed_after: Optional[str] = None, completed_before: Optional[str] = None, dequeued_after: Optional[str] = None, dequeued_before: Optional[str] = None, name: Optional[Union[str, List[str]]] = None, app_version: Optional[Union[str, List[str]]] = None, forked_from: Optional[Union[str, List[str]]] = None, parent_workflow_id: Optional[Union[str, List[str]]] = None, user: Optional[Union[str, List[str]]] = None, queue_name: Optional[Union[str, List[str]]] = None, limit: Optional[int] = None, offset: Optional[int] = None, sort_desc: bool = False, workflow_id_prefix: Optional[Union[str, List[str]]] = None, load_input: bool = True, load_output: bool = True, executor_id: Optional[Union[str, List[str]]] = None, queues_only: bool = False, was_forked_from: Optional[bool] = None, has_parent: Optional[bool] = None, attributes: Optional[Dict[str, Any]] = None, schedule_name: Optional[Union[str, List[str]]] = None, application_name: Optional[Union[str, List[str]]] = None, ) -> List[WorkflowStatus]: ``` Retrieve a list of [`WorkflowStatus`](#workflow-status) of all workflows matching specified criteria. **Parameters:** - **workflow_ids**: Retrieve workflows with these IDs. - **status**: Retrieve workflows with this status (or one of these statuses) (Must be `ENQUEUED`, `DELAYED`, `PENDING`, `SUCCESS`, `ERROR`, `CANCELLED`, or `MAX_RECOVERY_ATTEMPTS_EXCEEDED`) - **start_time**: Retrieve workflows started after this (RFC 3339-compliant) timestamp. - **end_time**: Retrieve workflows started before this (RFC 3339-compliant) timestamp. - **completed_after**: Retrieve workflows that completed after this (RFC 3339-compliant) timestamp. - **completed_before**: Retrieve workflows that completed before this (RFC 3339-compliant) timestamp. - **dequeued_after**: Retrieve workflows that were dequeued after this (RFC 3339-compliant) timestamp. - **dequeued_before**: Retrieve workflows that were dequeued before this (RFC 3339-compliant) timestamp. - **name**: Retrieve workflows with this name (or one of these names). - **app_version**: Retrieve workflows tagged with this application version (or one of these versions). - **forked_from**: Retrieve workflows forked from this workflow ID (or one of these IDs). - **parent_workflow_id**: Retrieve workflows that were started as children of this workflow (or one of these workflows). - **user**: Retrieve workflows run by this authenticated user (or one of these users). - **queue_name**: Retrieve workflows that were enqueued on this queue (or one of these queues). - **limit**: Retrieve up to this many workflows. - **offset**: Skip this many workflows from the results returned (for pagination). - **sort_desc**: Whether to sort the results in descending (`True`) or ascending (`False`) order by workflow start time. - **workflow_id_prefix**: Retrieve workflows whose IDs start with the specified string (or one of these strings). - **load_input**: Whether to load and deserialize workflow inputs. Set to `False` to improve performance when inputs are not needed. - **load_output**: Whether to load and deserialize workflow outputs. Set to `False` to improve performance when outputs are not needed. - **executor_id**: Retrieve workflows with this executor ID (or one of these IDs). - **queues_only**: If `True`, only retrieve workflows that are currently queued (status `DELAYED`, `ENQUEUED`, or `PENDING` and `queue_name` not null). Equivalent to using [`list_queued_workflows`](#list_queued_workflows). - **was_forked_from**: If `True`, only retrieve workflows that have been forked from. If `False`, only retrieve workflows that have not been forked from. - **has_parent**: If `True`, only retrieve workflows that have a parent workflow. If `False`, only retrieve workflows without a parent. - **attributes**: Retrieve workflows whose [custom attributes](#setworkflowattributes) contain all the given key-value pairs (nested values are matched exactly). Only supported when using a Postgres system database; raises `DBOSException` on SQLite. - **schedule_name**: Retrieve workflows that were enqueued by this [scheduled workflow](../tutorials/scheduled-workflows.md) (or one of these schedule names). - **application_name**: Retrieve workflows owned by this application (or one of these applications). Workflows owned by no application are always included. If unset, retrieve only this application's workflows (or, if `workflow_ids` is set, workflows owned by any application). On a Postgres system database, this query is subject to the [`observability_query_timeout_sec`](./configuration.md#database-connection-settings) statement timeout (unless `workflow_ids` is set) and raises `DBOSQueryTimeoutError` if it exceeds it. #### list_workflows_async Coroutine version of [`list_workflows`](#list_workflows). #### list_queued_workflows ```python def list_queued_workflows( *, workflow_ids: Optional[List[str]] = None, status: Optional[Union[str, List[str]]] = None, start_time: Optional[str] = None, end_time: Optional[str] = None, completed_after: Optional[str] = None, completed_before: Optional[str] = None, dequeued_after: Optional[str] = None, dequeued_before: Optional[str] = None, name: Optional[Union[str, List[str]]] = None, app_version: Optional[Union[str, List[str]]] = None, forked_from: Optional[Union[str, List[str]]] = None, parent_workflow_id: Optional[Union[str, List[str]]] = None, user: Optional[Union[str, List[str]]] = None, queue_name: Optional[Union[str, List[str]]] = None, limit: Optional[int] = None, offset: Optional[int] = None, sort_desc: bool = False, workflow_id_prefix: Optional[Union[str, List[str]]] = None, load_input: bool = True, load_output: bool = True, executor_id: Optional[Union[str, List[str]]] = None, has_parent: Optional[bool] = None, attributes: Optional[Dict[str, Any]] = None, application_name: Optional[Union[str, List[str]]] = None, ) -> List[WorkflowStatus]: ``` Retrieve a list of [`WorkflowStatus`](#workflow-status) of all **queued** workflows (status `DELAYED`, `ENQUEUED`, or `PENDING` and `queue_name` not null) matching specified criteria. **Parameters:** - **workflow_ids**: Retrieve workflows with these IDs. - **status**: Retrieve workflows with this status (or one of these statuses) (Must be `DELAYED`, `ENQUEUED`, or `PENDING`) - **start_time**: Retrieve workflows enqueued after this (RFC 3339-compliant) timestamp. - **end_time**: Retrieve workflows enqueued before this (RFC 3339-compliant) timestamp. - **completed_after**: Retrieve workflows that completed after this (RFC 3339-compliant) timestamp. - **completed_before**: Retrieve workflows that completed before this (RFC 3339-compliant) timestamp. - **dequeued_after**: Retrieve workflows that were dequeued after this (RFC 3339-compliant) timestamp. - **dequeued_before**: Retrieve workflows that were dequeued before this (RFC 3339-compliant) timestamp. - **name**: Retrieve workflows with this name (or one of these names). - **app_version**: Retrieve workflows tagged with this application version (or one of these versions). - **forked_from**: Retrieve workflows forked from this workflow ID (or one of these IDs). - **parent_workflow_id**: Retrieve workflows that were started as children of this workflow (or one of these workflows). - **user**: Retrieve workflows run by this authenticated user (or one of these users). - **queue_name**: Retrieve workflows running on this queue (or one of these queues). - **limit**: Retrieve up to this many workflows. - **offset**: Skip this many workflows from the results returned (for pagination). - **sort_desc**: Whether to sort the results in descending (`True`) or ascending (`False`) order by workflow start time. - **workflow_id_prefix**: Retrieve workflows whose IDs start with the specified string (or one of these strings). - **load_input**: Whether to load and deserialize workflow inputs. Set to `False` to improve performance when inputs are not needed. - **load_output**: Whether to load and deserialize workflow outputs. Set to `False` to improve performance when outputs are not needed. - **executor_id**: Retrieve workflows with this executor ID (or one of these IDs). - **has_parent**: If `True`, only retrieve workflows that have a parent workflow. If `False`, only retrieve workflows without a parent. - **attributes**: Retrieve workflows whose [custom attributes](#setworkflowattributes) contain all the given key-value pairs (nested values are matched exactly). Only supported when using a Postgres system database; raises `DBOSException` on SQLite. - **application_name**: Retrieve workflows owned by this application (or one of these applications). Workflows owned by no application are always included. If unset, retrieve only this application's workflows (or, if `workflow_ids` is set, workflows owned by any application). On a Postgres system database, this query is subject to the [`observability_query_timeout_sec`](./configuration.md#database-connection-settings) statement timeout (unless `workflow_ids` is set) and raises `DBOSQueryTimeoutError` if it exceeds it. #### list_queued_workflows_async Coroutine version of [`list_queued_workflows`](#list_queued_workflows). #### list_workflow_steps ```python def list_workflow_steps( workflow_id: str, *, load_output: bool = True, limit: Optional[int] = None, offset: Optional[int] = None, ) -> List[StepInfo] ``` Retrieve the steps of a workflow. Steps are ordered by `function_id`. Use `limit` and `offset` to paginate results. Set `load_output` to `False` to improve performance when step outputs and errors are not needed; the `output` and `error` fields are then always `None`. On a Postgres system database, this query is subject to the [`observability_query_timeout_sec`](./configuration.md#database-connection-settings) statement timeout and raises `DBOSQueryTimeoutError` if it exceeds it. This is a list of `StepInfo` objects, with the following structure: ```python class StepInfo(TypedDict): # The unique ID of the step in the workflow. One-indexed. function_id: int # The name of the step function_name: str # The step's output, if any output: Optional[Any] # The error the step threw, if any error: Optional[Exception] # If the step starts or retrieves the result of a workflow, its ID child_workflow_id: Optional[str] # The Unix epoch timestamp at which this step started started_at_epoch_ms: Optional[int] # The Unix epoch timestamp at which this step completed completed_at_epoch_ms: Optional[int] ``` #### list_workflow_steps_async Coroutine version of [`list_workflow_steps`](#list_workflow_steps). #### set_workflow_delay ```python DBOS.set_workflow_delay( workflow_id: str, *, delay_seconds: Optional[float] = None, delay_until_epoch_ms: Optional[int] = None, ) -> None ``` Set or update the delay on a workflow. Only affects workflows with `DELAYED` status. Provide exactly one of `delay_seconds` (relative) or `delay_until_epoch_ms` (absolute). **Parameters:** - `workflow_id`: The ID of the workflow whose delay to set. - `delay_seconds`: Delay the workflow by this many seconds from now. Must be non-negative. - `delay_until_epoch_ms`: Delay the workflow until this absolute time, specified as a Unix epoch timestamp in milliseconds. Must be non-negative. #### set_workflow_delay_async Coroutine version of [`set_workflow_delay`](#set_workflow_delay). #### update_workflow_attributes ```python DBOS.update_workflow_attributes( workflow_id: str, attributes: Optional[Dict[str, Any]], ) -> None ``` Replace the custom [attributes](#setworkflowattributes) attached to a workflow, identified by `workflow_id`. This overwrites the workflow's attributes dictionary; it is not a merge. Pass `None` to clear all attributes. Attributes must be a dictionary of JSON-serializable values. You can use this to attach attributes to a workflow that started without them, or to update attributes as a workflow progresses. This method is safe to call from within a workflow (including to update the calling workflow's own attributes). **Parameters:** - `workflow_id`: The ID of the workflow whose attributes to replace. - `attributes`: The new attributes dictionary, or `None` to clear all attributes. #### update_workflow_attributes_async Coroutine version of [`update_workflow_attributes`](#update_workflow_attributes). #### cancel_workflow ```python DBOS.cancel_workflow( workflow_id: str, *, cancel_children: bool = False, ) -> None ``` Cancel a workflow. This sets its status to `CANCELLED`, removes it from its queue (if it is enqueued) and preempts its execution (interrupting it at the beginning of its next step) **Parameters:** - **workflow_id**: The ID of the workflow to cancel. - **cancel_children**: If `True`, also recursively cancels all child workflows started by this workflow. #### cancel_workflow_async Coroutine version of [`cancel_workflow`](#cancel_workflow). #### cancel_workflows ```python DBOS.cancel_workflows( workflow_ids: List[str], *, cancel_children: bool = False, ) -> None ``` Cancel multiple workflows. Behaves like [`cancel_workflow`](#cancel_workflow) but operates on a list of workflow IDs. #### cancel_workflows_async Coroutine version of [`cancel_workflows`](#cancel_workflows). #### resume_workflow ```python DBOS.resume_workflow( workflow_id: str, *, queue_name: Optional[str] = None, ) -> WorkflowHandle[Any] ``` Resume a workflow. This immediately starts it from its last completed step. You can use this to resume workflows that are cancelled or have exceeded their maximum recovery attempts. You can also use this to start an enqueued workflow immediately, bypassing its queue. If `queue_name` is provided, the resumed workflow is enqueued on the specified queue instead of starting immediately. Raises `DBOSNonExistentWorkflowError` if the workflow does not exist. #### resume_workflow_async Coroutine version of [`resume_workflow`](#resume_workflow). #### resume_workflows ```python DBOS.resume_workflows( workflow_ids: List[str], *, queue_name: Optional[str] = None, ) -> List[WorkflowHandle[Any]] ``` Resume multiple workflows. Behaves like [`resume_workflow`](#resume_workflow) but operates on a list of workflow IDs and returns a list of handles. If any of the workflows does not exist, raises `DBOSNonExistentWorkflowError` without resuming any of them. #### resume_workflows_async Coroutine version of [`resume_workflows`](#resume_workflows). Returns `List[WorkflowHandleAsync[Any]]`. #### fork_workflow ```python DBOS.fork_workflow( workflow_id: str, start_step: int, *, application_version: Optional[str] = None, queue_name: Optional[str] = None, queue_partition_key: Optional[str] = None, replacement_children: Optional[dict[str, str]] = None, timeout_seconds: Optional[float] = None, ) -> WorkflowHandle[Any] ``` Start a new execution of a workflow from a specific step. The input step ID must match the `function_id` of the step returned by `list_workflow_steps`. The specified `start_step` is the step from which the new workflow will start, so any steps whose ID is less than `start_step` will not be re-executed. Raises `DBOSNonExistentWorkflowError` if the workflow identified by `workflow_id` does not exist. The forked workflow will have a new workflow ID, which can be set with [`SetWorkflowID`](#setworkflowid). It is possible to specify the application version on which the forked workflow will run by setting `application_version`, this is useful for "patching" workflows that failed due to a bug in a previous application version. If `queue_name` is provided, the forked workflow is enqueued on the specified queue instead of starting immediately. If the queue is partitioned, you can also specify `queue_partition_key`. If `replacement_children` is provided, it maps original child workflow IDs to replacement child workflow IDs. When the forked workflow encounters a step that started a child workflow matching an original ID, it substitutes the replacement ID instead. This is useful when you need to fork a parent workflow that depends on the results of child workflows that have also been forked. If `timeout_seconds` is provided, it sets a [timeout](#setworkflowtimeout) for the forked workflow, in seconds. #### fork_workflow_async Coroutine version of [`fork_workflow`](#fork_workflow). #### rewind_workflow ```python DBOS.rewind_workflow( workflow_id: str, *, start_step: Optional[int] = None, application_version: Optional[str] = None, queue_name: Optional[str] = None, queue_partition_key: Optional[str] = None, ) -> WorkflowHandle[Any] ``` Rewind a workflow to a specific step and run it again from that step, keeping its workflow ID. DBOS discards the workflow's recorded steps with IDs greater than or equal to `start_step`, then re-enqueues the workflow so it re-executes from that step. Steps with an ID less than `start_step` are not re-executed; their recorded outputs are replayed. The input step ID must match the `function_id` of the step returned by [`list_workflow_steps`](#list_workflow_steps). If `start_step` is not provided, the workflow's entire history is discarded and it re-executes from the beginning. Unlike [`fork_workflow`](#fork_workflow), which creates a new workflow with a new ID, rewind replays the workflow in place. Other workflows and clients can keep sending messages to, reading events from, and reading streams from the same workflow ID, and its child workflows keep the same IDs. Only a workflow in a terminal state (`SUCCESS`, `ERROR`, `CANCELLED`, or `MAX_RECOVERY_ATTEMPTS_EXCEEDED`) can be rewound. To rewind a workflow that is still running, [cancel](#cancel_workflow) it first. Rewinding a workflow: - Clears its recorded output or error. - Discards events it set at or after `start_step`, restoring each such event to the last value it set before `start_step` (or removing it if there is none). - Deletes messages it consumed at or after `start_step`, as well as any unconsumed messages. - Does not modify streamed values. Streams are append-only. New values will be appended to the end of existing streams. However, it does "un-close" any streams closed at or after `start_step`. - Does not modify child workflows, even those started after `start_step`. If you want to rerun child workflows, delete them or rewind them separately. **Parameters:** - **workflow_id**: The ID of the workflow to rewind. - **start_step**: The step to rewind to. Steps with IDs greater than or equal to this value are discarded and re-executed. Defaults to the first step. - **application_version**: The [application version](../tutorials/upgrading-workflows.md) on which the rewound workflow runs. Useful for "patching" a workflow that failed due to a bug in a previous application version. Defaults to the workflow's current version. - **queue_name**: The queue on which to enqueue the rewound workflow. Defaults to an internal queue, which dequeues it immediately. - **queue_partition_key**: The partition key to enqueue the rewound workflow under, if `queue_name` is a partitioned queue. #### rewind_workflow_async Coroutine version of [`rewind_workflow`](#rewind_workflow). #### delete_workflow ```python DBOS.delete_workflow( workflow_id: str, *, delete_children: bool = False, ) -> None ``` Delete a workflow and all its associated data (inputs, outputs, step results, etc.) from the system database. **Parameters:** - **workflow_id**: The ID of the workflow to delete. - **delete_children**: If `True`, also recursively deletes all child workflows started by this workflow. :::warning This operation is irreversible. Once a workflow is deleted, it cannot be recovered, resumed, or forked. ::: #### delete_workflow_async ```python DBOS.delete_workflow_async( workflow_id: str, *, delete_children: bool = False, ) -> Coroutine[Any, Any, None] ``` Coroutine version of [`delete_workflow`](#delete_workflow). #### delete_workflows ```python DBOS.delete_workflows( workflow_ids: List[str], *, delete_children: bool = False, ) -> None ``` Delete multiple workflows and all their associated data. Behaves like [`delete_workflow`](#delete_workflow) but operates on a list of workflow IDs. #### delete_workflows_async Coroutine version of [`delete_workflows`](#delete_workflows). ### Workflow Schedules #### create_schedule ```python DBOS.create_schedule( *, schedule_name: str, workflow_fn: Union[Callable[[datetime, Any], None], Callable[[datetime, Any], Coroutine[Any, Any, None]]], schedule: str, context: Any = None, automatic_backfill: bool = False, cron_timezone: Optional[str] = None, queue_name: Optional[str] = None, ) -> None ``` Create a cron schedule that periodically invokes a workflow function. **Parameters:** - **schedule_name**: Unique name identifying this schedule. - **workflow_fn**: The workflow function to invoke. Must take two arguments: a `datetime` (the scheduled execution time) and a context object. - **schedule**: A cron expression. Supports seconds as the first field with 6-field format. - **context**: An optional context object passed to the workflow function on each invocation. Must be serializable. - **automatic_backfill**: If `True`, on startup the scheduler will automatically backfill missed executions since the last time the schedule fired. Defaults to `False`. - **cron_timezone**: [IANA timezone name](https://en.wikipedia.org/wiki/List_of_tz_database_time_zones) (e.g. `"America/New_York"`) in which to evaluate the cron expression. Defaults to `None` (UTC). - **queue_name**: Optional name of a declared queue to enqueue scheduled workflows to. If `None`, uses an internal queue. This is useful for managing the concurrency of scheduled workflows. Defaults to `None`. Schedules are owned by the application that creates them: only that application's processes fire the schedule, and its workflows run on that application. Schedule names are globally unique across all applications sharing a system database, so creating a schedule whose name already exists (including one owned by a different application) raises an error. DBOS uses [croniter](https://pypi.org/project/croniter/) to parse cron schedules, using seconds as an optional first field ([`second_at_beginning=True`](https://pypi.org/project/croniter/#about-second-repeats)). Valid cron schedules contain 5 or 6 items, separated by spaces: ``` ┌────────────── second (optional) │ ┌──────────── minute │ │ ┌────────── hour │ │ │ ┌──────── day of month │ │ │ │ ┌────── month │ │ │ │ │ ┌──── day of week │ │ │ │ │ │ │ │ │ │ │ │ * * * * * * ``` **Example:** ```python from datetime import datetime from typing import Any from dbos import DBOS @DBOS.workflow() def my_periodic_task(scheduled_time: datetime, context: Any): DBOS.logger.info(f"Running task scheduled for {scheduled_time} with context {context}") # Create a schedule that runs every 5 minutes DBOS.create_schedule( schedule_name="my-task-schedule", # The schedule name is a unique identifier of the schedule workflow_fn=my_periodic_task, schedule="*/5 * * * *", # Every 5 minutes context="my context", # The context is passed into every iteration of the workflow ) ``` #### create_schedule_async Coroutine version of [`create_schedule`](#create_schedule). #### list_schedules ```python DBOS.list_schedules( *, status: Optional[Union[str, List[str]]] = None, workflow_name: Optional[Union[str, List[str]]] = None, schedule_name_prefix: Optional[Union[str, List[str]]] = None, application_name: Optional[Union[str, List[str]]] = None, ) -> List[WorkflowSchedule] ``` Return all registered workflow schedules, optionally filtered. Returns a list of [`WorkflowSchedule`](#workflowschedule). **Parameters:** - **status**: Filter by status (e.g. `"ACTIVE"`) or a list of statuses. - **workflow_name**: Filter by workflow name or a list of names. - **schedule_name_prefix**: Filter by schedule name prefix or a list of prefixes. - **application_name**: List only schedules owned by this application (or one of these applications). Schedules owned by no application are always included. If unset, list only this application's schedules. #### list_schedules_async Coroutine version of [`list_schedules`](#list_schedules). #### get_schedule ```python DBOS.get_schedule(name: str) -> Optional[WorkflowSchedule] ``` Return the [`WorkflowSchedule`](#workflowschedule) with the given name, or `None` if it does not exist. #### get_schedule_async Coroutine version of [`get_schedule`](#get_schedule). #### delete_schedule ```python DBOS.delete_schedule(name: str) -> None ``` Delete the schedule with the given name. No-op if it does not exist. #### delete_schedule_async Coroutine version of [`delete_schedule`](#delete_schedule). #### pause_schedule ```python DBOS.pause_schedule(name: str) -> None ``` Pause the schedule with the given name. A paused schedule does not fire. #### resume_schedule ```python DBOS.resume_schedule(name: str) -> None ``` Resume a paused schedule so it begins firing again. #### apply_schedules ```python DBOS.apply_schedules( schedules: List[ScheduleInput], ) -> None class ScheduleInput(TypedDict): schedule_name: str workflow_fn: Union[Callable[[datetime, Any], None], Callable[[datetime, Any], Coroutine[Any, Any, None]]] schedule: str context: Any # Optional, defaults to None automatic_backfill: bool # Optional, defaults to False cron_timezone: Optional[str] # Optional, defaults to None (UTC) queue_name: Optional[str] # Optional, defaults to None (internal queue) ``` Atomically apply a set of schedules. Useful for declaratively defining all your static schedules in one place. May not be called from within a workflow. Existing schedules are upserted by name: all definition fields are replaced with the new entry's values (so any optional field left unset is cleared, e.g. an omitted `queue_name` reverts the schedule to the internal queue), while the schedule's ID, status, and last-fired time are preserved. **Example:** ```python DBOS.apply_schedules([ {"schedule_name": "schedule-a", "workflow_fn": workflow_a, "schedule": "*/10 * * * *", "context": None}, # Every 10 minutes {"schedule_name": "schedule-b", "workflow_fn": workflow_b, "schedule": "0 0 * * *", "context": None}, # Every day at midnight ]) ``` #### apply_schedules_async Coroutine version of [`apply_schedules`](#apply_schedules). #### backfill_schedule ```python DBOS.backfill_schedule( schedule_name: str, start: datetime, end: datetime, ) -> List[WorkflowHandle[None]] ``` Enqueue (on the schedule's `queue_name`, or an internal queue if it has none) all executions of a schedule that would have run between `start` and `end` (both exclusive). Each execution uses the same deterministic workflow ID as the live scheduler, so already-executed times are skipped. May not be called from within a workflow. #### trigger_schedule ```python DBOS.trigger_schedule(schedule_name: str) -> WorkflowHandle[None] ``` Immediately enqueue (on the schedule's `queue_name`, or an internal queue if it has none) the scheduled workflow at the current time. May not be called from within a workflow. #### WorkflowSchedule Some schedule management methods return the `WorkflowSchedule` type: ```python class WorkflowSchedule(TypedDict): # The unique identifier of the schedule schedule_id: str # The human-readable name of the schedule schedule_name: str # The name of the workflow function to execute workflow_name: str # The class name of the workflow function, if it is a class method workflow_class_name: Optional[str] # The cron expression defining the schedule schedule: str # The status of the schedule: "ACTIVE" or "PAUSED" status: str # The context object passed to each workflow invocation context: Any # The timestamp of when the schedule last fired, if ever last_fired_at: Optional[str] # Whether missed executions are automatically backfilled on startup automatic_backfill: bool # The IANA timezone in which the cron expression is evaluated, or None for UTC cron_timezone: Optional[str] # The name of the queue scheduled workflows are enqueued to, or None for the internal queue queue_name: Optional[str] # The application that owns this schedule, or None if it is owned by no application application_name: Optional[str] ``` #### Workflow Status Some workflow introspection and management methods return a `WorkflowStatus`. This object has the following definition: ```python class WorkflowStatus: # The workflow ID workflow_id: str # The workflow status. Must be one of ENQUEUED, DELAYED, PENDING, SUCCESS, ERROR, CANCELLED, or MAX_RECOVERY_ATTEMPTS_EXCEEDED status: str # The name of the workflow function name: str # The name of the workflow's class, if any class_name: Optional[str] # The name with which the workflow's class instance was configured, if any config_name: Optional[str] # The user who ran the workflow, if specified authenticated_user: Optional[str] # The role with which the workflow ran, if specified assumed_role: Optional[str] # All roles which the authenticated user could assume authenticated_roles: Optional[list[str]] # The deserialized workflow input object input: Optional[WorkflowInputs] # The workflow's output, if any output: Optional[Any] # The error the workflow threw, if any error: Optional[Exception] # Workflow start time, as a Unix epoch timestamp in ms created_at: Optional[int] # Last time the workflow status was updated, as a Unix epoch timestamp in ms updated_at: Optional[int] # If this workflow was enqueued, on which queue queue_name: Optional[str] # The executor to most recently execute this workflow executor_id: Optional[str] # The application version on which this workflow was started app_version: Optional[str] # The start-to-close timeout of the workflow in ms workflow_timeout_ms: Optional[int] # The deadline of a workflow, computed by adding its timeout to its start time. workflow_deadline_epoch_ms: Optional[int] # Unique ID for deduplication on a queue deduplication_id: Optional[str] # Priority of the workflow on the queue, starting from 1 ~ 2,147,483,647. Default 0 (highest priority). priority: Optional[int] # If this workflow is enqueued on a partitioned queue, its partition key queue_partition_key: Optional[str] # If this workflow was forked from another, that workflow's ID. forked_from: Optional[str] # Whether this workflow has ever been forked from by another workflow. was_forked_from: bool # If this workflow was started as a child of another workflow, that workflow's ID. parent_workflow_id: Optional[str] # The Unix epoch timestamp in ms at which the workflow was last dequeued, if it had been enqueued dequeued_at: Optional[int] # The Unix epoch timestamp in ms before which the workflow should not be dequeued, if it was delayed delay_until_epoch_ms: Optional[int] # The Unix epoch timestamp in ms at which the workflow completed (SUCCESS, ERROR, or CANCELLED), if it has completed completed_at: Optional[int] # Custom key-value attributes attached to the workflow with SetWorkflowAttributes # or update_workflow_attributes, if any attributes: Optional[Dict[str, Any]] # If this workflow was enqueued by a scheduled workflow, that schedule's name schedule_name: Optional[str] # The application that owns this workflow, or None if it is owned by no application application_name: Optional[str] ``` ### Context Variables #### logger ```python DBOS.logger: Logger ``` Retrieve the DBOS logger. This is a pre-configured Python logger provided as a convenience. #### workflow_id ```python DBOS.workflow_id: Optional[str] ``` Return the ID of the currently executing workflow. If a workflow is not executing, return None. #### step_id ```python DBOS.step_id: Optional[int] ``` Return the step ID for the currently executing step. This is a unique identifier of the current step within the workflow. If a step is not currently executing, return None. #### step_status ```python DBOS.step_status: Optional[StepStatus] ``` Return the status of the currently executing step. If a step is not currently executing, return None. The `StepStatus` object has the following properties: ```python class StepStatus: # The unique ID of this step in its workflow. step_id: int # For steps with automatic retries, which attempt number (zero-indexed) is currently executing. current_attempt: Optional[int] # For steps with automatic retries, the maximum number of attempts that will be made before the step fails. max_attempts: Optional[int] ``` #### span ```python DBOS.span: Optional[opentelemetry.trace.Span] ``` Retrieve the OpenTelemetry span associated with the current workflow or step, or `None` if there is no active span. You can use this to set custom attributes in your span. #### executor_id ```python DBOS.executor_id: str ``` Retrieve the current executor ID, a unique process ID used to identify the application instance in distributed environments. ### Version Management #### application_version ```python DBOS.application_version: str ``` Retrieve the application version of this process. #### list_application_versions ```python DBOS.list_application_versions() -> List[VersionInfo] class VersionInfo(TypedDict): # A unique ID for this version version_id: str # The unique name of this version version_name: str # The epoch timestamp (in milliseconds) of this version. This is used to determine the latest version. version_timestamp: int # The epoch timestamp (in milliseconds) when this version was first registered. created_at: int # The application that registered this version, or None if it is owned by no application application_name: Optional[str] ``` Return all registered application versions, ordered by timestamp descending (newest first). Versions are tracked per application: this returns only versions registered by this application, plus versions owned by no application. #### list_application_versions_async ```python await DBOS.list_application_versions_async() -> List[VersionInfo] ``` Coroutine version of [`list_application_versions`](#list_application_versions). #### get_latest_application_version ```python DBOS.get_latest_application_version() -> VersionInfo ``` Return the latest application version (the one with the highest timestamp) among versions registered by this application, plus versions owned by no application. Raises `DBOSException` if no versions are registered. #### get_latest_application_version_async ```python await DBOS.get_latest_application_version_async() -> VersionInfo ``` Coroutine version of [`get_latest_application_version`](#get_latest_application_version). #### set_latest_application_version ```python DBOS.set_latest_application_version( version_name: str, *, application_name: Optional[str] = None, ) -> None ``` Promote a version to latest by updating its timestamp to the current time. This is useful when rolling back to a previous application version. **Parameters:** - `version_name`: The name of the version to promote. - `application_name`: The application to act as. Defaults to this application. Version names are globally unique across applications sharing a system database, so promoting a version registered by a different application raises an error. #### set_latest_application_version_async ```python await DBOS.set_latest_application_version_async( version_name: str, *, application_name: Optional[str] = None, ) -> None ``` Coroutine version of [`set_latest_application_version`](#set_latest_application_version). ### Debouncing You can create a `Debouncer` to debounce your workflows. Debouncing delays workflow execution until some time has passed since the workflow has last been called. This is useful for preventing wasted work when a workflow may be triggered multiple times in quick succession. For example, if a user is editing an input field, you can debounce their changes to execute a processing workflow only after they haven't edited the field for some time: #### Debouncer.create ```python Debouncer.create( workflow: Callable[P, R], *, debounce_timeout_sec: Optional[float] = None, queue: Optional[Union[Queue, str]] = None, application_name: Optional[str] = None, ) -> Debouncer[P, R] ``` **Parameters:** - `workflow`: The workflow to debounce. Must be a function or static method: bound methods, including those of [configured instances](../tutorials/classes.md), cannot be debounced and raise a `TypeError`. - `debounce_timeout_sec`: After this time elapses since the first time a workflow is submitted from this debouncer, the workflow is started regardless of the debounce period. - `queue`: When starting a workflow after debouncing, enqueue it on this queue (a `Queue` or a queue name) instead of an internal queue. - `application_name`: Debounce on behalf of this application instead of your own: the debounced workflow is owned and run by that application. #### debounce ```python debouncer.debounce( debounce_key: str, debounce_period_sec: float, *args: P.args, **kwargs: P.kwargs, ) -> WorkflowHandle[R] ``` Submit a workflow for execution but delay it by `debounce_period_sec`. Returns a handle to the workflow. The workflow may be debounced again, which further delays its execution (up to `debounce_timeout_sec`). When the workflow eventually executes, it uses the **last** set of inputs passed into `debounce`. Once the debounce period expires and the workflow is released for execution, the next call to `debounce` starts the debouncing process again for a new workflow execution. **Parameters:** - `debounce_key`: A key used to group workflow executions that will be debounced together. For example, if the debounce key is set to customer ID, each customer's workflows would be debounced separately. - `debounce_period_sec`: Delay this workflow's execution by this period. - `*args`: Variadic workflow arguments. - `**kwargs`: Variadic workflow keyword arguments. **Example Syntax**: ```python @DBOS.workflow() def process_input(user_input): ... # Each time a user submits a new input, debounce the process_input workflow. # The debouncer will wait until 60 seconds after the user stops submitting new inputs, # then start the workflow processing the last input submitted. debouncer = Debouncer.create(process_input) def on_user_input_submit(user_id, user_input): debounce_key = user_id debounce_period_sec = 60 debouncer.debounce(debounce_key, debounce_period_sec, user_input) ``` #### Debouncer.create_async ```python Debouncer.create_async( workflow: Callable[P, Coroutine[Any, Any, R]], *, debounce_timeout_sec: Optional[float] = None, queue: Optional[Union[Queue, str]] = None, application_name: Optional[str] = None, ) -> Debouncer[P, R] ``` Async version of `Debouncer.create`. #### debounce_async ```python debouncer.debounce_async( debounce_key: str, debounce_period_sec: float, *args: P.args, **kwargs: P.kwargs, ) -> Coroutine[Any, Any, WorkflowHandleAsync[R]] ``` Async version of `debouncer.debounce`. ### Authentication #### authenticated_user ```python DBOS.authenticated_user: Optional[str] ``` Return the current authenticated user, if any, associated with the current context. #### authenticated_roles ```python DBOS.authenticated_roles: Optional[List[str]] ``` Return the roles granted to the current authenticated user, if any, associated with the current context. #### assumed_role ```python DBOS.assumed_role: Optional[str] ``` Return the role currently assumed by the authenticated user, if any, associated with the current context. #### set_authentication ```python DBOS.set_authentication( authenticated_user: Optional[str], authenticated_roles: Optional[List[str]] ) -> None ``` Set the current authenticated user and granted roles into the current context. This would generally be done by HTTP middleware ### Context Management #### SetWorkflowID ```python SetWorkflowID( wfid: str, *, workflow_id_reuse_policy: Optional[WorkflowIDReusePolicy] = None, ) ``` Set the [workflow ID](../tutorials/workflow-tutorial.md#workflow-ids-and-idempotency) of the next workflow to run. Should be used in a `with` statement. **Parameters:** - **wfid**: The workflow ID to assign. - **workflow_id_reuse_policy**: What to do if a workflow with this ID already exists, whatever its status. Defaults to `"return-existing"`. - `"return-existing"`: return a handle to the existing workflow (or, for a direct call, its result) without starting a new one. This is the idempotency behavior described in [Workflow IDs and Idempotency](../tutorials/workflow-tutorial.md#workflow-ids-and-idempotency). - `"reject"`: raise `DBOSWorkflowIDInUseError` without starting a new workflow or modifying the existing one. The exception has `workflow_id`, `workflow_status`, and `workflow_name` attributes describing the existing workflow. Cannot be used with a [debouncer](#debouncing). Example syntax: ```python @DBOS.workflow() def example_workflow(): DBOS.logger.info(f"I am a workflow with ID {DBOS.workflow_id}") # The workflow will run with the supplied ID with SetWorkflowID("very-unique-id"): example_workflow() ``` Example of rejecting a reused workflow ID: ```python from dbos import DBOS, SetWorkflowID from dbos import error as dboserror def submit_order(order_id: str, order: Order) -> str: try: with SetWorkflowID(f"order-{order_id}", workflow_id_reuse_policy="reject"): handle = DBOS.start_workflow(process_order, order) return handle.get_result() except dboserror.DBOSWorkflowIDInUseError: return "already started" ``` #### SetWorkflowTimeout ```python SetWorkflowTimeout( workflow_timeout_sec: Optional[float] ) ``` Set a timeout for all enclosed workflow invocations or enqueues. When the timeout expires, the workflow **and all its children** are cancelled. Cancelling a workflow sets its status to `CANCELLED` and preempts its execution at the beginning of its next step. Timeouts are **start-to-completion**: if a workflow is enqueued, the timeout does not begin until the workflow is dequeued and starts execution. Also, timeouts are **durable**: they are stored in the database and persist across restarts, so workflows can have very long timeouts. Timeout deadlines are propagated to child workflows by default, so when a workflow's deadline expires all of its child workflows (and their children, and so on) are also cancelled. If you want to detach a child workflow from its parent's timeout, you can start it with `SetWorkflowTimeout(custom_timeout)` to override the propagated timeout. You can use `SetWorkflowTimeout(None)` to start a child workflow with no timeout. Example syntax: ```python @DBOS.workflow() def example_workflow(): ... # If the workflow does not complete within 10 seconds, it times out and is cancelled with SetWorkflowTimeout(10): example_workflow() ``` #### SetWorkflowAttributes ```python SetWorkflowAttributes( attributes: Optional[Dict[str, Any]] ) ``` Attach a dictionary of custom key-value attributes to all workflows started or enqueued within the block. Attributes must be a dictionary of JSON-serializable values. Pass `None` to attach no attributes. Attributes are recorded in a workflow's [status](#workflow-status) at creation time and are **not inherited** by child workflows. You can later retrieve a workflow's attributes from its [`WorkflowStatus`](#workflow-status) and filter workflows by attribute with [`list_workflows`](#list_workflows) and [`list_queued_workflows`](#list_queued_workflows). To change a workflow's attributes after it is created, use [`update_workflow_attributes`](#update_workflow_attributes). Attributes are stored in Postgres as GIN-indexed JSONB, so they are efficiently searchable. Example syntax: ```python @DBOS.workflow() def example_workflow(): ... # example_workflow is recorded with these attributes with SetWorkflowAttributes({"customer": "acme", "region": "us-east-1"}): example_workflow() ``` #### PropagateOtelContext ```python PropagateOtelContext( context: Optional[opentelemetry.context.Context] = None ) ``` Propagate the current OpenTelemetry context (or optionally, a passed-in context) to all workflows started or enqueued within the block, so their spans join the caller's trace. See [the tracing tutorial](../tutorials/logging-and-tracing.md#keeping-enqueued-workflows-on-the-callers-trace) for an overview. The propagated trace context is durably recorded in the workflow's [attributes](#setworkflowattributes) (under the reserved `dbos.otelContext` key), so the workflow's span parents to the caller's trace no matter when or where the workflow runs: immediately, after a queue handoff (possibly in another process), or on recovery. Only the W3C trace context (`traceparent`/`tracestate`) is propagated, not baggage. If there is no active span, nothing is recorded. The propagated context is **not inherited** by child workflows; use `PropagateOtelContext` again inside a workflow to keep its children on the same trace. Requires the DBOS OpenTelemetry dependencies (`pip install dbos[otel]`); [tracing](../tutorials/logging-and-tracing.md#tracing) must be enabled for the propagated context to take effect. To propagate a trace context when enqueuing from outside a DBOS application, use the `otel_context` field of DBOS Client's [`EnqueueOptions`](./client.md#enqueue) instead. Example syntax: ```python with PropagateOtelContext(): handle = DBOS.enqueue_workflow("example_queue", workflow_function, ...) ``` #### DBOSContextEnsure ```python DBOSContextEnsure() # Code inside will run with a DBOS context available with DBOSContextEnsure(): # Call DBOS functions pass ``` Use of `DBOSContextEnsure` ensures that there is a DBOS context associated with the enclosed code prior to calling DBOS functions. `DBOSContextEnsure` is generally not used by applications directly, but used by event dispatchers, HTTP server middleware, etc., to set up the DBOS context prior to entry into function calls. #### DBOSContextSetAuth ```python DBOSContextSetAuth(user: Optional[str], roles: Optional[List[str]]) # Code inside will run with `curuser` and `curroles` with DBOSContextSetAuth(curuser, curroles): # Call DBOS functions pass ``` `with DBOSContextSetAuth` sets the current authorized user and roles for the code inside the `with` block. Similar to `DBOSContextEnsure`, `DBOSContextSetAuth` also ensures that there is a DBOS context associated with the enclosed code prior to calling DBOS functions. `DBOSContextSetAuth` is generally not used by applications directly, but used by event dispatchers, HTTP server middleware, etc., to set up the DBOS context prior to entry into function calls. ### Alerting #### alert_handler ```python @DBOS.alert_handler def my_handler(rule_type: str, message: str, metadata: Dict[str, str]) -> None: ... ``` Register a function to handle [alerts](../../conductor/alerting.md) received from Conductor. The handler function is called with three arguments: - **rule_type**: The type of alert rule. One of `WorkflowFailure`, `SlowQueue`, or `UnresponsiveApplication`. - **message**: The alert message. - **metadata**: A dictionary of string key-value pairs with additional alert information. Only one alert handler may be registered per application, and it must be defined after DBOS is initialized (`DBOS(config=...)`) and before `DBOS.launch()` is called. If no handler is registered, alerts are logged to the DBOS logger. **Example syntax:** ```python @DBOS.alert_handler def handle_alert(rule_type: str, message: str, metadata: dict[str, str]) -> None: DBOS.logger.warning(f"Alert received: {rule_type} - {message}") for key, value in metadata.items(): DBOS.logger.warning(f" {key}: {value}") ``` ### Custom Serialization DBOS must serialize data such as workflow inputs and outputs and step outputs to store it in the system database. By default, data is serialized with `pickle` then Base64-encoded, but you can optionally supply a custom serializer through DBOS configuration. A custom serializer must match this interface: ```python class Serializer(ABC): @abstractmethod def serialize(self, data: Any) -> str: pass @abstractmethod def deserialize(self, serialized_data: str) -> Any: pass def name(self) -> str: """The serializer's `name` is stored with serialized values and used to ensure that the correct deserializer is used.""" return "custom_serializer" ``` For example, here is how to configure DBOS to use a JSON serializer: ```python import json import os from typing import Any from dbos import DBOS, DBOSConfig, Serializer class JsonSerializer(Serializer): def serialize(self, data: Any) -> str: return json.dumps(data) def deserialize(self, serialized_data: str) -> Any: return json.loads(serialized_data) def name(self) -> str: return "basic_json" serializer = JsonSerializer() config: DBOSConfig = { "name": "dbos-starter", "application_version": "0.1.0", "system_database_url": os.environ.get("DBOS_SYSTEM_DATABASE_URL"), "serializer": serializer } DBOS(config=config) DBOS.launch() ``` ### Serialization Strategy Several DBOS methods accept an optional `serialization_type` parameter that controls how data is serialized. This is useful for cross-language interoperability—for example, if a TypeScript or Java DBOS application needs to read events or messages set by a Python application. ```python from dbos import WorkflowSerializationFormat ``` The available strategies are: - **`WorkflowSerializationFormat.DEFAULT`**: Uses the serializer configured in [`DBOSConfig`](./configuration.md) (defaults to pickle). When called from within a workflow, uses that workflow's serialization format instead (for example, the portable format for a workflow that uses portable serialization). - **`WorkflowSerializationFormat.PORTABLE`**: Uses a portable JSON format (`portable_json`) that can be deserialized by DBOS applications in any language. - **`WorkflowSerializationFormat.NATIVE`**: Explicitly uses the native Python pickle serializer (`py_pickle`). --- ## Datasources(Reference) Datasources wrap a SQLAlchemy engine so that database transactions run inside DBOS workflows are tracked and replayed with exactly-once guarantees. See the [Transactions & Datasources tutorial](../tutorials/transaction-tutorial.md) for a full walkthrough. ### SQLAlchemyDatasource A synchronous datasource backed by a SQLAlchemy `Engine`. Use this for non-async code. #### `SQLAlchemyDatasource.create` ```python SQLAlchemyDatasource.create( database_url: str, engine_kwargs: Optional[Dict[str, Any]] = None, engine: Optional[sa.Engine] = None, schema: Optional[str] = None, serializer: Optional[Serializer] = None, sessionmaker: Optional[sessionmaker] = None, run_migrations: bool = True, ) -> SQLAlchemyDatasource ``` Factory method. Creates (or reuses) a SQLAlchemy engine and creates or migrates the datasource's `datasource_outputs` tracking table, unless `run_migrations` is `False`. **Parameters:** - `database_url`: A SQLAlchemy-compatible connection URL (e.g., `"postgresql+psycopg://..."` or `"sqlite:///./my.db"`). - `engine_kwargs`: Optional keyword arguments forwarded verbatim to SQLAlchemy's `create_engine`. - `engine`: Provide an existing `sa.Engine` instead of creating one from `database_url`. When set, `engine_kwargs` is ignored. - `schema`: The PostgreSQL schema in which the `datasource_outputs` table is created. Defaults to `"dbos"`. Has no effect for SQLite. - `serializer`: A custom serializer for transaction outputs. Defaults to the DBOS default serializer (`pickle`, then Base64-encoded), not the `serializer` set in [`DBOSConfig`](./configuration.md#serialization-settings). - `sessionmaker`: A custom `sqlalchemy.orm.sessionmaker` used to create the session for each transaction, for example to use a custom `Session` subclass, session options, or event hooks. DBOS always binds sessions to the datasource's engine, so any `bind` you set is ignored, and a sessionmaker that sets `binds` raises a `DBOSException`. Defaults to `sessionmaker(expire_on_commit=False)`. - `run_migrations`: Whether to create and migrate the datasource's tables. Defaults to `True`. Set to `False` if your application's database role cannot run DDL: the datasource then only verifies that its tables are migrated, raising a `DBOSInitializationError` if they are not. Migrate them separately with [`SQLAlchemyDatasource.migrate`](#sqlalchemydatasourcemigrate). **Example:** ```python from dbos import SQLAlchemyDatasource ds = SQLAlchemyDatasource.create(os.environ["APP_DATABASE_URL"]) ``` #### `SQLAlchemyDatasource.migrate` ```python SQLAlchemyDatasource.migrate( database_url: str, *, schema: Optional[str] = None, application_role: Optional[str] = None, ) -> None ``` Create or migrate a datasource's tables without creating a datasource. Run this with a privileged database role (for example, as part of your deployment's migration step), then create your datasources with `run_migrations=False` so your application's role needs no DDL privileges. **Parameters:** - `database_url`: The connection URL of the datasource's database. - `schema`: The PostgreSQL schema holding the datasource's tables. Defaults to `"dbos"`. Has no effect for SQLite. - `application_role`: A PostgreSQL role to grant the minimal permissions a datasource needs at runtime: usage on the schema, `SELECT`, `INSERT`, and `DELETE` on the `datasource_outputs` table, and `SELECT` on the migration version table. Not supported for SQLite. **Example:** ```python from dbos import SQLAlchemyDatasource # In your migration script, run with a privileged role: SQLAlchemyDatasource.migrate(os.environ["ADMIN_DATABASE_URL"], application_role="my_app_role") # In your application, run with my_app_role: ds = SQLAlchemyDatasource.create(os.environ["APP_DATABASE_URL"], run_migrations=False) ``` #### `SQLAlchemyDatasource.transaction` ```python ds.transaction( func: Optional[Callable] = None, *, name: Optional[str] = None, isolation_level: IsolationLevel = "SERIALIZABLE", ) ``` Decorator that registers a synchronous function as a datasource transaction step. The decorated function must **not** be a coroutine (`async def`). Decorating an `async def` function raises `DBOSException` at decoration time. **Parameters:** - `name`: Step name recorded in the workflow log. Defaults to the function's qualified name (`__qualname__`). - `isolation_level`: SQL transaction isolation level. Must be one of `"SERIALIZABLE"` (default), `"REPEATABLE READ"`, or `"READ COMMITTED"`. SQLite supports only `"SERIALIZABLE"`. **Example:** ```python @ds.transaction() def insert_row(name: str, value: int) -> None: session = ds.sql_session() session.execute(text("INSERT INTO t VALUES (:n, :v)"), {"n": name, "v": value}) @ds.transaction(isolation_level="READ COMMITTED", name="read_row") def get_row(name: str) -> Optional[int]: session = ds.sql_session() row = session.execute(text("SELECT value FROM t WHERE name = :n"), {"n": name}).first() return row[0] if row else None ``` #### `SQLAlchemyDatasource.run_tx_step` ```python ds.run_tx_step( ds_options: Optional[DatasourceOptions], func: Callable[P, R], *args: P.args, **kwargs: P.kwargs, ) -> R ``` Runs `func` as a datasource transaction step without requiring the `@ds.transaction` decorator. Raises `DBOSException` if `func` is a coroutine. **Parameters:** - `ds_options`: A `DatasourceOptions` dict with optional keys `name` and `isolation_level`, or `None` to use the defaults. - `func`: The function to execute inside the transaction. - `*args`, `**kwargs`: Arguments forwarded to `func`. **Example:** ```python def insert_row(name: str, value: int) -> None: session = ds.sql_session() session.execute(text("INSERT INTO t VALUES (:n, :v)"), {"n": name, "v": value}) @DBOS.workflow() def my_workflow(name: str, value: int) -> None: ds.run_tx_step({"name": "insert_row"}, insert_row, name, value) ``` #### `SQLAlchemyDatasource.sql_session` ```python ds.sql_session() -> Session ``` Returns the SQLAlchemy `Session` for the current datasource transaction. Must be called from within a function that is executing inside a datasource transaction (i.e., decorated with `@ds.transaction` or called via `run_tx_step`). Raises `AssertionError` if called outside a transaction. --- ### AsyncSQLAlchemyDatasource An asynchronous datasource backed by a SQLAlchemy `AsyncEngine`. Use this for `async def` code. #### `AsyncSQLAlchemyDatasource.create` ```python await AsyncSQLAlchemyDatasource.create( database_url: str, engine_kwargs: Optional[Dict[str, Any]] = None, engine: Optional[AsyncEngine] = None, schema: Optional[str] = None, serializer: Optional[Serializer] = None, sessionmaker: Optional[async_sessionmaker] = None, run_migrations: bool = True, ) -> AsyncSQLAlchemyDatasource ``` Async factory method. Creates (or reuses) a SQLAlchemy `AsyncEngine` and creates or migrates the datasource's `datasource_outputs` tracking table, unless `run_migrations` is `False`. **Parameters:** - `database_url`: A SQLAlchemy-compatible async connection URL (e.g., `"postgresql+psycopg://..."` or `"sqlite+aiosqlite:///./my.db"`). - `engine_kwargs`: Optional keyword arguments forwarded verbatim to SQLAlchemy's `create_async_engine`. - `engine`: Provide an existing `AsyncEngine` instead of creating one from `database_url`. When set, `engine_kwargs` is ignored. - `schema`: The PostgreSQL schema in which the `datasource_outputs` table is created. Defaults to `"dbos"`. Has no effect for SQLite. - `serializer`: A custom serializer for transaction outputs. Defaults to the DBOS default serializer (`pickle`, then Base64-encoded), not the `serializer` set in [`DBOSConfig`](./configuration.md#serialization-settings). - `sessionmaker`: A custom `sqlalchemy.ext.asyncio.async_sessionmaker` used to create the session for each transaction, for example to use a custom `AsyncSession` subclass, session options, or event hooks. DBOS always binds sessions to the datasource's engine, so any `bind` you set is ignored, and a sessionmaker that sets `binds` raises a `DBOSException`. Defaults to `async_sessionmaker(expire_on_commit=False)`. - `run_migrations`: Whether to create and migrate the datasource's tables. Defaults to `True`. Set to `False` if your application's database role cannot run DDL: the datasource then only verifies that its tables are migrated, raising a `DBOSInitializationError` if they are not. Migrate them separately with [`AsyncSQLAlchemyDatasource.migrate`](#asyncsqlalchemydatasourcemigrate). **Example:** To safely create `AsyncSQLAlchemyDatasource` at module scope: ```python import asyncio import os from dbos import AsyncSQLAlchemyDatasource # The datasource is not tied to the event loop that created it, # so it can be used from the event loop that runs your application. ads = asyncio.run(AsyncSQLAlchemyDatasource.create(os.environ["APP_DATABASE_URL"])) ``` #### `AsyncSQLAlchemyDatasource.migrate` ```python await AsyncSQLAlchemyDatasource.migrate( database_url: str, *, schema: Optional[str] = None, application_role: Optional[str] = None, ) -> None ``` Coroutine version of [`SQLAlchemyDatasource.migrate`](#sqlalchemydatasourcemigrate). Create or migrate a datasource's tables without creating a datasource, then create your datasources with `run_migrations=False`. #### `AsyncSQLAlchemyDatasource.transaction` ```python ads.transaction( func: Optional[Callable] = None, *, name: Optional[str] = None, isolation_level: IsolationLevel = "SERIALIZABLE", ) ``` Decorator that registers an async coroutine function as a datasource transaction step. The decorated function **must** be a coroutine (`async def`). Decorating a non-coroutine raises `DBOSException` at decoration time. **Parameters:** - `name`: Step name recorded in the workflow log. Defaults to the function's qualified name (`__qualname__`). - `isolation_level`: SQL transaction isolation level. Must be one of `"SERIALIZABLE"` (default), `"REPEATABLE READ"`, or `"READ COMMITTED"`. SQLite supports only `"SERIALIZABLE"`. **Example:** ```python @ads.transaction() async def insert_row(name: str, value: int) -> None: session = ads.sql_session() await session.execute(text("INSERT INTO t VALUES (:n, :v)"), {"n": name, "v": value}) @ads.transaction(isolation_level="READ COMMITTED", name="read_row") async def get_row(name: str) -> Optional[int]: session = ads.sql_session() row = (await session.execute(text("SELECT value FROM t WHERE name = :n"), {"n": name})).first() return row[0] if row else None ``` #### `AsyncSQLAlchemyDatasource.run_tx_step_async` ```python await ads.run_tx_step_async( ds_options: Optional[DatasourceOptions], func: Callable[P, Coroutine[Any, Any, R]], *args: P.args, **kwargs: P.kwargs, ) -> R ``` Runs `func` as a datasource transaction step without requiring the `@ds.transaction` decorator. Raises `DBOSException` if `func` is not a coroutine. **Parameters:** - `ds_options`: A `DatasourceOptions` dict with optional keys `name` and `isolation_level`, or `None` to use the defaults. - `func`: The coroutine function to execute inside the transaction. - `*args`, `**kwargs`: Arguments forwarded to `func`. **Example:** ```python async def insert_row(name: str, value: int) -> None: session = ads.sql_session() await session.execute(text("INSERT INTO t VALUES (:n, :v)"), {"n": name, "v": value}) @DBOS.workflow() async def my_workflow(name: str, value: int) -> None: await ads.run_tx_step_async({"name": "insert_row"}, insert_row, name, value) ``` #### `AsyncSQLAlchemyDatasource.sql_session` ```python ads.sql_session() -> AsyncSession ``` Returns the SQLAlchemy `AsyncSession` for the current datasource transaction. Must be called from within a coroutine that is executing inside a datasource transaction (i.e., decorated with `@ds.transaction` or called via `run_tx_step_async`). Raises `AssertionError` if called outside a transaction. --- ### DatasourceOptions ```python class DatasourceOptions(TypedDict, total=False): name: Optional[str] isolation_level: Optional[IsolationLevel] ``` A `TypedDict` passed to `run_tx_step` / `run_tx_step_async` to configure the step. Both fields are optional; pass `None` instead of the dict to use all defaults. **Fields:** - `name`: Step name recorded in the workflow log. - `isolation_level`: One of `"SERIALIZABLE"`, `"REPEATABLE READ"`, or `"READ COMMITTED"` (SQLite supports only `"SERIALIZABLE"`). --- ## DBOS Class The DBOS class is a singleton—you must instantiate it (by calling its constructor) exactly once in a program's lifetime. Here, we document its constructor and lifecycle methods. Decorators are documented [here](./decorators.md) and context methods and variables [here](./contexts.md). ### class dbos.DBOS ```python DBOS( *, config: DBOSConfig, ) ``` **Parameters:** - `config`: Configuration parameters for DBOS. See the [configuration docs](./configuration.md). #### launch ```python DBOS.launch() ``` Launch DBOS, initializing database connections and starting queues and scheduled workflows. Should be called after all decorators run. **You should not run a DBOS workflow until after DBOS is launched.** **Example:** ```python import os from dbos import DBOS, DBOSConfig @DBOS.step() def step_one(): print("Step one completed!") @DBOS.step() def step_two(): print("Step two completed!") @DBOS.workflow() def dbos_workflow(): step_one() step_two() # Configure and launch DBOS, then run a workflow. if __name__ == "__main__": config: DBOSConfig = { "name": "dbos-starter", "application_version": "0.1.0", "system_database_url": os.environ.get("DBOS_SYSTEM_DATABASE_URL"), } DBOS(config=config) DBOS.launch() dbos_workflow() ``` #### listen_queues ```python DBOS.listen_queues( queues: Sequence[str] ) ``` Configure this DBOS process to only listen to (dequeue workflows from) specific queues. If this is not used, DBOS will listen to all registered queues. Must be called after DBOS is constructed and before it is launched, and may be called at most once. **Parameters:** - `queues`: The names of the queues to listen to. #### destroy ```python DBOS.destroy( *, destroy_registry: bool = False, workflow_completion_timeout_sec: int = 0, ) ``` Destroy the DBOS singleton, stopping its background threads (such as queue polling and the scheduler) and closing database connections. `destroy` waits up to `workflow_completion_timeout_sec` for active workflows to complete and cancels executor tasks that have not yet started; it does not interrupt workflows that are still running, but they can no longer checkpoint their progress. After this completes, the singleton can be re-initialized: construct a new instance with `DBOS(config=...)` before calling `DBOS.launch()` again. Useful for testing. **Parameters:** - `workflow_completion_timeout_sec`: Wait this many seconds for active workflows to complete before shutting down. - `destroy_registry`: Whether to destroy the global registry of decorated functions. If set to `True`, `destroy` will "un-register" all decorated functions. You probably want to leave this `False`. #### migrate ```python DBOS.migrate( system_database_url: str, *, schema: str = "dbos", application_role: Optional[str] = None, ) ``` Create or migrate the DBOS [system database](../../explanations/system-tables.md) without launching DBOS. This is the programmatic equivalent of the [`dbos migrate`](./cli.md#dbos-migrate) command. Run it with a privileged database role (for example, as part of your deployment's migration step), then configure your application with [`run_migrations=False`](./configuration.md#database-connection-settings) so its role needs no DDL privileges. It can be called before or without constructing a DBOS instance. **Parameters:** - `system_database_url`: The system database to create or migrate. The database is created if it does not exist. - `schema`: The Postgres schema containing the DBOS system tables. Defaults to `dbos`. - `application_role`: A Postgres role to grant access to the DBOS system schema once it is migrated. Not supported for SQLite. **Example:** ```python from dbos import DBOS # In your migration script, run with a privileged role: DBOS.migrate(os.environ["ADMIN_SYSTEM_DATABASE_URL"], application_role="my_app_role") ``` To migrate the tables used by [datasources](./datasources.md), use [`SQLAlchemyDatasource.migrate`](./datasources.md#sqlalchemydatasourcemigrate). #### reset_system_database ```python DBOS.reset_system_database( *, system_database_url: Optional[str] = None, schema: Optional[str] = None, truncate: bool = False, ) ``` Reset the DBOS [system database](../../explanations/system-tables.md), clearing DBOS's internal state. By default, this destroys the system database entirely; pass `truncate=True` to instead empty its tables, which is substantially faster. Useful when testing a DBOS application to reset the internal state of DBOS between tests. For example, see its use in the [testing tutorial](../tutorials/testing.md). It cannot be called after DBOS is launched. **This is a destructive operation and should only be used in a test environment.** **Parameters:** - `system_database_url`: The system database to reset. Defaults to the system database of the current DBOS instance. If no DBOS instance has been constructed, this must be supplied. - `schema`: The Postgres schema containing the DBOS system tables. Defaults to the configured [`dbos_system_schema`](./configuration.md#database-connection-settings), or `dbos`. Only used when truncating. - `truncate`: If `True`, empty the system tables instead of dropping the system database. **This is recommended**, as it is substantially faster. Because the migrated schema is left in place, DBOS does not need to re-create the system database on its next launch. --- ## Workflows & Steps(3) ### Function Decorators #### workflow ```python DBOS.workflow( *, name: Optional[str] = None, max_recovery_attempts: Optional[int] = 100, serialization_type: Optional[WorkflowSerializationFormat] = None, validate_args: Optional[ValidateArgsCallable] = None, ) ``` Durably execute this function as a [DBOS workflow](../tutorials/workflow-tutorial.md). **Example:** ```python @DBOS.workflow() def greeting_workflow(name: str, note: str): sign_guestbook(name) insert_greeting(name, note) ``` **Parameters:** - `name`: A name for this workflow. If not provided, the function's qualified name (`__qualname__`, which does not include its module) is used. Workflow names must be unique: registering workflows with the same name from different modules raises a `DBOSException`. - `max_recovery_attempts`: The maximum number of times execution of a workflow may be attempted. This acts as a [dead letter queue](https://en.wikipedia.org/wiki/Dead_letter_queue) so that a buggy workflow that crashes its application (for example, by running it out of memory) does not do so infinitely. If a workflow exceeds this limit, its status is set to `MAX_RECOVERY_ATTEMPTS_EXCEEDED` and it may no longer be executed. A workflow in this state may be [resumed](../tutorials/workflow-management.md#resuming-workflows), which will resume execution and reset the count of execution attempts. If this behavior is not desired, it may be disabled by setting `max_recovery_attempts=None`. - `serialization_type`: The default [serialization format](../../explanations/portable-workflows.md) to use for local invocations of this workflow. Set to `WorkflowSerializationFormat.PORTABLE` to test [cross-language interoperability](../../explanations/portable-workflows.md). - `validate_args`: An optional callable that validates and optionally coerces workflow arguments before execution. Pass the built-in `pydantic_args_validator` to automatically validate arguments against the function's type hints using [Pydantic](https://docs.pydantic.dev/). See [Input Validation and Coercion](#input-validation-and-coercion) below for details and examples. #### step ```python DBOS.step( *, name: Optional[str] = None, retries_allowed: bool = False, interval_seconds: float = 1.0, max_attempts: int = 3, backoff_rate: float = 2.0, should_retry: Optional[Callable[[BaseException], Union[bool, Awaitable[bool]]]] = None, preemptible: bool = False, timeout_seconds: Optional[float] = None, ) ``` Annotate a function as a step in a workflow. Workflows automatically checkpoint the outcomes of their steps. If a workflow is interrupted, it recovers from the last completed step. **Example:** ```python @DBOS.step(retries_allowed=True, max_attempts=10) def example_step(): return requests.get("https://example.com").text ``` **Parameters:** - `name`: A name for this step. If not provided, the function's qualified name (`__qualname__`) is used. - `retries_allowed`: Whether to retry the step if it throws an exception. - `interval_seconds`: How long to wait before the initial retry. - `max_attempts`: The maximum number of times to attempt a step that is throwing exceptions, including the first attempt. - `backoff_rate`: How much to multiplicatively increase `interval_seconds` between retries. - `should_retry`: Optional predicate called with the raised exception to decide whether the step should be retried. If it returns `False` (or an awaitable resolving to `False`), the exception is re-raised immediately without further retries. Ignored when `retries_allowed` is `False`. Async predicates are only supported for async steps. - `preemptible`: If `True`, the step is cancelled immediately when its workflow is cancelled, rather than running to completion. Only supported for async steps. - `timeout_seconds`: If set, cancel the step and raise `DBOSStepTimeoutError` if it runs for longer than this many seconds. Only supported for async steps, and must be positive and finite. Each retry attempt gets a fresh timeout. See [Step Timeouts](../tutorials/step-tutorial.md#step-timeouts). #### required_roles ```python DBOS.required_roles( roles: List[str] ) ``` The `@DBOS.required_roles` decorator applies role-based security to the decorated function. The authenticated user must have at least one of the roles on the `roles` list in order to access the function. **Parameters:** - `roles`: List of required roles applied to the decorated function. **Example:** ```python @DBOS.workflow() @DBOS.required_roles(["support","admin"]) def my_support_workflow(): pass # Function accessible only with "support" or "admin" role ``` #### kafka_consumer ```python DBOS.kafka_consumer( config: dict[str, Any], topics: list[str], *, ordering: Optional[Literal["none", "partition", "topic"]] = None, batch_size: int = 250, queue_name: Optional[str] = None, ) ``` Runs a function for each Kafka message received on the specified topic(s). Uses the Kafka message's topic, partition, and offset and the consumer group ID to create a unique [workflow id](../reference/contexts#setworkflowid) to ensure once and only once execution. Takes a configuration dictionary and a list of topics to consume. The decorated function must take a KafkaMessage as its only parameter. **Parameters:** - `config`: a dictionary of config settings. Information on key settings follows with full configuration setting details available in the [official Kafka documentation](https://kafka.apache.org/documentation/#consumerconfigs). - `bootstrap.servers`: A list of host/port pairs to use for establishing the initial connection to the Kafka cluster. This list should be in the form host1:port1,host2:port2,... - `group.id`: A unique string that identifies the consumer group this consumer belongs to. Setting it is recommended: if it is omitted, DBOS generates one from the function name and topics and logs a warning. - `topics`: a list of Kafka topics to subscribe to. A topic prefixed with `^` is treated as a regular expression. - `ordering`: Controls how messages are processed. See [In-Order Processing](../tutorials/kafka-integration.md#in-order-processing). - `"none"` (default): messages are processed in parallel. - `"partition"`: messages are processed serially per topic partition (preserving Kafka's per-partition delivery order) and in parallel across partitions. - `"topic"`: messages are processed serially per topic. - `batch_size`: The maximum number of messages consumed from Kafka and durably enqueued per batch. Defaults to 250. - `queue_name`: The name of an optional [queue](./queues.md) on which consumer workflows run, for example to configure concurrency or rate limits. Only valid with `ordering="none"`; ordered consumers share an internal partitioned queue. The named queue must not be a [partitioned queue](../tutorials/queue-tutorial.md#partitioning-queues). If you use [`DBOS.listen_queues`](./dbos-class.md#listen_queues), you must include this queue. **Example** ```python @DBOS.kafka_consumer( config={ "bootstrap.servers": "localhost:9092", "group.id": "dbos-kafka-group", }, topics=["example-topic"], ) @DBOS.workflow() def test_kafka_workflow(msg: KafkaMessage): DBOS.logger.info(f"Message received: {msg.value.decode()}") ``` ### Input Validation and Coercion Python workflows can specify a `validate_args` parameter on `@DBOS.workflow()`. The built-in `pydantic_args_validator` sentinel builds a [Pydantic](https://docs.pydantic.dev/) validator from the function's type hints at decoration time. This validates argument types and coerces compatible values (for example, ISO date strings to `datetime` objects). Validation runs when a workflow is dequeued for execution (for example, after being enqueued, recovered, or forked), not when the workflow function is called directly or started with `DBOS.start_workflow`. ```python from datetime import datetime from typing import Dict, Any, List from dbos import DBOS, WorkflowSerializationFormat, pydantic_args_validator @DBOS.workflow( serialization_type=WorkflowSerializationFormat.PORTABLE, validate_args=pydantic_args_validator, ) def process_order(name: str, count: int, tags: List[str]) -> str: # Pydantic validates types — "not_a_number" for count raises ValueError return f"{name}:{count}:{','.join(tags)}" @DBOS.workflow( serialization_type=WorkflowSerializationFormat.PORTABLE, validate_args=pydantic_args_validator, ) def schedule_task(name: str, due: datetime, tags: List[str]) -> str: # Pydantic coerces the ISO string "2025-06-15T10:30:00" to a datetime object return f"{name}@{due.isoformat()}#{','.join(tags)}" ``` You can also provide a custom validator function. It must accept `(positional_args_tuple, keyword_args_dict)` and return a validated `(positional_args_tuple, keyword_args_dict)`: ```python def my_validator(args, kwargs): # Custom validation or coercion logic return args, kwargs @DBOS.workflow( serialization_type=WorkflowSerializationFormat.PORTABLE, validate_args=my_validator, ) def my_workflow(x: int) -> str: return str(x) ``` For more context on why input validation matters for cross-language workflows, see [Input Validation and Coercion](../../explanations/portable-workflows.md#input-validation-and-coercion). ### Classes and Decorators Methods in classes can be decorated with any of the [function decorators](#function-decorators) above. Functions marked as `@classmethod` or `@staticmethod` are supported in the same way as regular functions. Classes with instance methods should extend from [`DBOSConfiguredInstance`](#dbosconfiguredinstance). #### dbos_class ```python DBOS.dbos_class( class_name: Optional[str] = None ) ``` The `@DBOS.dbos_class` decorator should be applied to all classes with DBOS workflow and step functions. This decorator assists in making sure all functions are properly registered with the class and provided with class-level configuration information. **Parameters** - `class_name` (Optional): A custom name to register the class with DBOS. By default, DBOS uses the class’s qualified name (`cls.__qualname__`) for identification. This can be overridden by providing a user-defined name, which may differ from the qualified name. All class names registered with DBOS must be globally unique. **Example:** ```python @DBOS.dbos_class() class MyClass: @staticmethod @DBOS.workflow() def my_class_wf(): pass ``` #### default_required_roles ```python DBOS.default_required_roles( roles: List[str] ) ``` The `@DBOS.default_required_roles` decorator can be applied to a class to set the default list of required access roles for all functions in the class. The list of required roles for individual functions can be overridden with [`required_roles`](#required_roles). **Parameters:** - `roles`: List of required roles to apply to all functions not individually decorated with [`required_roles`](#required_roles). **Example:** ```python @DBOS.default_required_roles(["user"]) class MyClass: @staticmethod @DBOS.workflow() def my_user_function() -> None: pass # Must have "user" role to access @staticmethod @DBOS.workflow() @DBOS.required_roles(["admin"]) def my_admin_function() -> None: pass # Must have "admin" role to access ``` #### DBOSConfiguredInstance ```python DBOSConfiguredInstance( config_name: str ) ``` `DBOSConfiguredInstance` should be used as a base for classes with decorated instance member functions. `DBOSConfiguredInstance` collects the instance name; this name is recorded in the database workflow records so that recovery can be targeted to the correct instance. `DBOSConfiguredInstance` also registers the class instance with the DBOS recovery system. **Parameters:** - `config_name`: The name of the instance, for recording in workflow database records **Example:** ```python @DBOS.dbos_class() class DBOSTestClass(DBOSConfiguredInstance): def __init__(self) -> None: super().__init__("instance1") ``` --- ## Queues(3) Queues let you submit functions for durable execution without starting them immediately. They are useful for controlling the number of functions run in parallel, or the rate at which functions are started. Queues are persisted to the system database. Register a queue with [`DBOS.register_queue`](./contexts.md#register_queue) and enqueue workflows on it with [`DBOS.enqueue_workflow`](./contexts.md#enqueue_workflow) or [`Queue.enqueue`](#enqueue). DBOS must be launched before you register a queue. ```python @DBOS.workflow() def send_email(to: str) -> None: ... DBOS.register_queue("email", global_concurrency=10, limiter={"limit": 100, "period": 60}) handle = DBOS.enqueue_workflow("email", send_email, "alice@example.com") ``` #### class dbos.Queue A `Queue` is returned by [`DBOS.register_queue`](./contexts.md#register_queue) and [`DBOS.retrieve_queue`](./contexts.md#retrieve_queue). You can inspect and modify the queue's properties and [`enqueue`](#enqueue)/[`enqueue_async`](#enqueue_async) workflows. ```python class QueueRateLimit(TypedDict): limit: int period: float # In seconds ``` **Properties:** - `name`: The name of the queue. - `global_concurrency`: The maximum number of functions from this queue that may run concurrently across all DBOS processes. If `None`, any number of functions may run concurrently. - `worker_concurrency`: The maximum number of functions from this queue that may run concurrently on a single DBOS process. - `limiter`: A limit on the maximum number of functions which may be started in a given period. - `partition_concurrency`: The maximum number of functions from any one [partition](../tutorials/queue-tutorial.md#partitioning-queues) of this queue that may run concurrently across all DBOS processes. - `partition_worker_concurrency`: The maximum number of functions from any one partition of this queue that may run concurrently on a single DBOS process. - `partition_limiter`: A limit on the maximum number of functions which may be started from any one partition in a given period. - `polling_interval_sec`: The minimum interval at which DBOS polls the database for new workflows on this queue. The actual interval includes random jitter and increases with backoff under contention, then scales back down when contention clears. - `application_name`: The application that owns this queue and dequeues workflows from it, or `None` if the queue is owned by no application. Unlike the other properties, ownership cannot be reconfigured. A queue is [partitioned](../tutorials/queue-tutorial.md#partitioning-queues) if any of its `partition_*` limits is set. Reading any property returns the latest value from the database, so changes made by other processes are reflected. In `async` code, use the [async getters](#async-property-accessors) (`get_global_concurrency_async`, `get_worker_concurrency_async`, etc.) instead of property access. Property access performs a synchronous database round-trip that blocks the event loop; reading a property from a running event loop logs a warning. #### enqueue ```python queue.enqueue( func: Callable[P, R], *args: P.args, **kwargs: P.kwargs, ) -> WorkflowHandle[R] ``` Enqueue a function for processing and return a [handle](./workflow_handles.md#workflowhandle) to it. You can enqueue any DBOS-annotated function. The `enqueue` method durably enqueues your function; after it returns your function is guaranteed to eventually execute even if your app is interrupted. [`DBOS.enqueue_workflow`](./contexts.md#enqueue_workflow) is a convenience wrapper that enqueues by queue name. **Example syntax:** ```python from dbos import DBOS @DBOS.step() def process_task(task): ... @DBOS.workflow() def process_tasks(tasks): task_handles = [] # Enqueue each task so all tasks are processed concurrently. for task in tasks: handle = queue.enqueue(process_task, task) task_handles.append(handle) # Wait for each task to complete and retrieve its result. # Return the results of all tasks. return [handle.get_result() for handle in task_handles] queue = DBOS.register_queue("example_queue") ``` #### enqueue_async ```python queue.enqueue_async( func: Callable[P, Coroutine[Any, Any, R]], *args: P.args, **kwargs: P.kwargs, ) -> Coroutine[Any, Any, WorkflowHandleAsync[R]] ``` Asynchronously enqueue an async function for processing and return an [async handle](./workflow_handles.md#workflowhandleasync) to it. You can enqueue any DBOS-annotated async function. The `enqueue_async` method durably enqueues your function; after it returns your function is guaranteed to eventually execute even if your app is interrupted. The enqueued function is launched into a different event loop than its caller. **Example syntax:** ```python from dbos import DBOS @DBOS.step() async def process_task_async(task): ... @DBOS.workflow() async def process_tasks(tasks): task_handles = [] # Enqueue each task so all tasks are processed concurrently. for task in tasks: handle = await queue.enqueue_async(process_task_async, task) task_handles.append(handle) # Wait for each task to complete and retrieve its result. # Return the results of all tasks. return [await handle.get_result() for handle in task_handles] queue = DBOS.register_queue("example_queue") ``` #### Async Property Accessors In `async` code, use these coroutine getters instead of reading the corresponding properties. Each one fetches the latest value from the database without blocking the event loop. Reading the synchronous property from a running event loop instead logs a warning. ```python queue.get_global_concurrency_async() -> Coroutine[Any, Any, Optional[int]] queue.get_worker_concurrency_async() -> Coroutine[Any, Any, Optional[int]] queue.get_limiter_async() -> Coroutine[Any, Any, Optional[QueueRateLimit]] queue.get_partition_concurrency_async() -> Coroutine[Any, Any, Optional[int]] queue.get_partition_worker_concurrency_async() -> Coroutine[Any, Any, Optional[int]] queue.get_partition_limiter_async() -> Coroutine[Any, Any, Optional[QueueRateLimit]] queue.get_polling_interval_sec_async() -> Coroutine[Any, Any, float] ``` #### Reconfiguring Queues You can reconfigure a queue at runtime by calling its `set_*` methods. Each setter writes the new value to the system database; workers pick up the new configuration on their next polling iteration without needing to restart. In `async` code, use the [`set_*_async`](#async-reconfiguration) variants instead so the event loop is not blocked on the database write. Calling a synchronous setter from a running event loop logs a warning. :::warning If your application calls [`DBOS.register_queue`](./contexts.md#register_queue) on startup, the next process to start can overwrite settings you applied at runtime via `set_*` methods. Either update the `register_queue` call to match the new configuration, or pass `on_conflict="never_update"` to preserve the runtime changes. ::: ##### set_global_concurrency ```python queue.set_global_concurrency(value: Optional[int]) -> None ``` Update the queue's global concurrency limit. Must be greater than or equal to the queue's `worker_concurrency`, `partition_concurrency`, and `partition_worker_concurrency`. Pass `None` to remove the limit. ##### set_worker_concurrency ```python queue.set_worker_concurrency(value: Optional[int]) -> None ``` Update the queue's per-worker concurrency limit. Must be less than or equal to the queue's `global_concurrency` and greater than or equal to its `partition_worker_concurrency`. Pass `None` to remove the limit. ##### set_limiter ```python queue.set_limiter(value: Optional[QueueRateLimit]) -> None ``` Update the queue's [rate limit](../tutorials/queue-tutorial.md#rate-limiting). Pass `None` to remove the limit. ##### set_partition_concurrency ```python queue.set_partition_concurrency(value: Optional[int]) -> None ``` Update the queue's per-partition concurrency limit. Must be at least 1, less than or equal to the queue's `global_concurrency`, and greater than or equal to its `partition_worker_concurrency`. Pass `None` to remove the limit. ##### set_partition_worker_concurrency ```python queue.set_partition_worker_concurrency(value: Optional[int]) -> None ``` Update the queue's per-partition, per-worker concurrency limit. Must be at least 1 and less than or equal to the queue's `partition_concurrency`, `worker_concurrency`, and `global_concurrency`. Pass `None` to remove the limit. ##### set_partition_limiter ```python queue.set_partition_limiter(value: Optional[QueueRateLimit]) -> None ``` Update the [rate limit](../tutorials/queue-tutorial.md#rate-limiting) applied to each partition separately. Pass `None` to remove the limit. :::info Setting any `partition_*` limit makes the queue [partitioned](../tutorials/queue-tutorial.md#partitioning-queues); clearing all of them makes it unpartitioned again. Take care when partitioning a queue at runtime: workflows already enqueued on it have no partition key and will not be dequeued until the queue is unpartitioned. ::: ##### set_polling_interval_sec ```python queue.set_polling_interval_sec(value: float) -> None ``` Update the queue's polling interval. Must be positive. ##### Async Reconfiguration Each setter has an async counterpart with the same parameters and validation behavior. Use these from `async` code to write to the database without blocking the event loop. ```python queue.set_global_concurrency_async(value: Optional[int]) -> Coroutine[Any, Any, None] queue.set_worker_concurrency_async(value: Optional[int]) -> Coroutine[Any, Any, None] queue.set_limiter_async(value: Optional[QueueRateLimit]) -> Coroutine[Any, Any, None] queue.set_partition_concurrency_async(value: Optional[int]) -> Coroutine[Any, Any, None] queue.set_partition_worker_concurrency_async(value: Optional[int]) -> Coroutine[Any, Any, None] queue.set_partition_limiter_async(value: Optional[QueueRateLimit]) -> Coroutine[Any, Any, None] queue.set_polling_interval_sec_async(value: float) -> Coroutine[Any, Any, None] ``` #### SetEnqueueOptions ```python SetEnqueueOptions( *, deduplication_id: Optional[str] = None, duplication_policy: Optional[DuplicationPolicy] = None, priority: Optional[int] = None, delay_seconds: Optional[float] = None, app_version: Optional[str] = None, queue_partition_key: Optional[str] = None, ) ``` Set options for enclosed workflow enqueue operations. These options are **not propagated** to child workflows. **Parameters:** - `deduplication_id`: At any given time, only one workflow with a specific deduplication ID can be enqueued in the specified queue. If a workflow with a deduplication ID is currently delayed, enqueued, or actively executing (status `DELAYED`, `ENQUEUED`, or `PENDING`), subsequent workflow enqueue attempt with the same deduplication ID in the same queue will raise a `DBOSQueueDeduplicatedError` exception. Defaults to `None`. - `duplication_policy`: How to handle a collision with another workflow that has the same `deduplication_id` on the same queue. Defaults to `"reject"`. - `"reject"`: raise `DBOSQueueDeduplicatedError`. - `"return-existing"`: return a handle to the existing workflow instead of raising. Requires a queue and a `deduplication_id`. Arguments passed by the colliding caller are discarded and the returned handle resolves with the original workflow's result. See [Singleton Workflows](../tutorials/queue-tutorial.md#singleton-workflows). - `priority`: The priority of the enqueued workflow in the specified queue. Workflows with the same priority are dequeued in **FIFO (first in, first out)** order. Priority values can range from `1` to `2,147,483,647`, where **a low number indicates a higher priority**. Defaults to `None`. Workflows without assigned priorities have the highest priority and are dequeued before workflows with assigned priorities. - `delay_seconds`: Delay the workflow by this many seconds before it becomes eligible for execution. The workflow is initially placed in `DELAYED` status and transitions to `ENQUEUED` after the delay expires. Defaults to `None` (no delay). - `app_version`: The application version of the workflow to enqueue. The workflow may only be dequeued by processes running that version. Defaults to the current application version. - `queue_partition_key`: The queue partition in which to enqueue this workflow. Use if and only if the queue is [partitioned](../tutorials/queue-tutorial.md#partitioning-queues) (registered with at least one `partition_*` limit). A partitioned queue applies its `partition_*` limits to each partition separately, while its `global_concurrency`, `worker_concurrency`, and `limiter` still apply across all partitions. Cannot be combined with `deduplication_id`. **Deduplication Example** ```python from dbos import DBOS, SetEnqueueOptions from dbos import error as dboserror DBOS.register_queue("example_queue") with SetEnqueueOptions(deduplication_id="my_dedup_id"): try: handle = DBOS.enqueue_workflow("example_queue", example_workflow, ...) except dboserror.DBOSQueueDeduplicatedError as e: # Handle deduplication error ... ``` **Singleton Workflow Example** ```python from dbos import DBOS, SetEnqueueOptions DBOS.register_queue("example_queue") # Only one workflow with deduplication ID "singleton" can be active on this queue # at a time. Subsequent callers attach to it and receive its result. with SetEnqueueOptions( deduplication_id="singleton", duplication_policy="return-existing" ): handle = DBOS.enqueue_workflow("example_queue", example_workflow, ...) result = handle.get_result() ``` **Priority Example** ```python DBOS.register_queue("priority_queue") with SetEnqueueOptions(priority=10): # All workflows are enqueued with priority set to 10 # They will be dequeued in FIFO order for task in tasks: DBOS.enqueue_workflow("priority_queue", task_workflow, task) # first_workflow (priority=1) will be dequeued before all task_workflows (priority=10) with SetEnqueueOptions(priority=1): DBOS.enqueue_workflow("priority_queue", first_workflow) ``` **Partitioned Queue Example** ```python @DBOS.workflow() def process_task(task: Task): ... DBOS.register_queue("partitioned_queue", partition_concurrency=1) def on_user_task_submission(user_id: str, task: Task): # Partition the task queue by user ID. As the queue has a # per-partition concurrency of 1, this means that at most one # task can run at once per user (but tasks from different # users can run concurrently). with SetEnqueueOptions(queue_partition_key=user_id): DBOS.enqueue_workflow("partitioned_queue", process_task, task) ``` --- ## Workflow Handles A workflow handle represents the state of a particular active or completed workflow execution. You obtain a workflow handle when using `DBOS.start_workflow` to start a workflow in the background or when enqueueing a workflow with [`DBOS.enqueue_workflow`](./contexts.md#enqueue_workflow) or [`Queue.enqueue`](./queues.md#enqueue). If you know a workflow's identity, you can also retrieve its handle using `DBOS.retrieve_workflow`. ### WorkflowHandle #### Methods ##### get_workflow_id ```python handle.get_workflow_id() -> str ``` Retrieve the ID of the workflow. ##### get_result ```python handle.get_result( *, polling_interval_sec: float = 1.0, ) -> R ``` Wait for the workflow to complete, then return its result. **Parameters:** - **polling_interval_sec**: The interval at which DBOS polls the database for the workflow's result. Only used for enqueued workflows or retrieved handles. ##### get_status ```python handle.get_status() -> WorkflowStatus ``` Retrieve the [`WorkflowStatus`](./contexts.md#workflow-status) of a workflow. ### WorkflowHandleAsync #### Methods ##### get_workflow_id ```python handle.get_workflow_id() -> str ``` Retrieve the ID of the workflow. Behaves identically to the [WorkflowHandle](#workflowhandle) version. ##### get_result ```python handle.get_result( *, polling_interval_sec: float = 1.0, ) -> Coroutine[Any, Any, R] ``` Asynchronously wait for the workflow to complete, then return its result. Similar to the [WorkflowHandle](#workflowhandle) version, except asynchronous. ##### get_status ```python handle.get_status() -> Coroutine[Any, Any, WorkflowStatus] ``` Asynchronously retrieve the [`WorkflowStatus`](./contexts.md#workflow-status) of a workflow. --- ## Authentication and Authorization DBOS Python supports modular, declarative security. This is a cooperative effort between the request framework (such as FastAPI) and DBOS Transact; the server framework performs authentication and forwards the authenticated user and roles on to DBOS functions, which then check for authorization. Authentication information is forwarded via the [DBOS context](../reference/contexts.md#set_authentication). The following tutorial shows use of FastAPI middleware to collect and authenticate user information, and shows how to set function- and class-level authorization in DBOS. ### Authentication Middleware Authentication information may arrive in various ways, depending on the approach, protocol, and framework used. It is common to use some sort of "middleware" that sits between the request handler and the application code that processes the authentication information in each request. A simple middleware for FastAPI may pass authentication information to DBOS in a manner similar to this: ```python @app.middleware("http") async def authMiddleware( request: Request, call_next: Callable[[Request], Awaitable[Response]] ) -> Response: with DBOSContextSetAuth("user1", ["user", "engineer"]): response = await call_next(request) return response ``` There are several things happening in this code snippet. * `authMiddleware` - The middleware function responsible for taking information from the `Request` and providing it to DBOS via the context * `@app.middleware("http")` - Registers the middleware function with `app`, which is a FastAPI instance, so that the function is called on inbound requests * `with DBOSContextSetAuth("user1", ["user", "engineer"]):` - Ensures that a DBOS context is associated with the request, and sets the authenticated user and roles into the context. This authentication information will be used to make authorization checks for any function calls within the scope of the `with`. Use of `with` also ensures that authentication information will be cleared when the code block finishes. * `response = await call_next(request)` - Proceeds down the handler call chain, with the authentication information in place ### Authorization Decorators DBOS Python uses [decorators](../reference/decorators.md) to declare the roles required to authorize access to functions and methods. * [`required_roles`](../reference/decorators.md#required_roles) is used at the function/method level to list roles required for function access. Users who were authenticated with any role on the list are allowed to access the function. * [`default_required_roles`](../reference/decorators.md#default_required_roles) can be used to set a list of roles that applies to each method in class, as a default. Any methods decorated with `required_roles` will use the list provided by `required_roles` instead of the default list provided to `default_required_roles`. For example, most methods in a class may require "user" access, with a few exceptions, such as "login", which requires no authentication/authorization, or a few administrative functions that require the "admin" role: ```python @DBOS.default_required_roles(["user"]) class DBOSTestClass: @staticmethod @DBOS.workflow() def user_function(var: str) -> str: # "user" role required due to default_required_roles return var @staticmethod @DBOS.workflow() @DBOS.required_roles(["admin"]) def admin_function(var: str) -> str: # "admin" role required due to method-level override return var @staticmethod @DBOS.required_roles([]) @DBOS.workflow() def login_function(var: str) -> str: # No role, or any authentication, required due to method-level override return var ``` ### Example In this example, we demonstrate how to use JWT tokens with DBOS declarative security. Here, the JWT tokens may be generated by a different part of the stack, separating out the concern of robust user management and authentication credentials, which can then be handled in specialized libraries or services. ```python @app.middleware("http") async def jwtAuthMiddleware( request: Request, call_next: Callable[[Request], Awaitable[Response]] ) -> Response: user: Optional[str] = None roles: Optional[List[str]] = None try: token = await oauth2_scheme(request) if token is not None: tdata = decode_jwt(token) user = tdata.username roles = tdata.roles except Exception as e: pass with DBOSContextSetAuth(user, roles): response = await call_next(request) return response ``` As with the simpler example in [Authentication Middleware](#authentication-middleware) above, `@app.middleware("http")` is used to insert the `jwtAuthMiddleware` function between the FastAPI `app` and DBOS. Use of `DBOSContextSetAuth` and `call_next` is also the same. What is different is that `oauth2_scheme` and `decode_jwt` are used to extract token contents. The resulting information is used to set the DBOS user and roles. If the user / roles are stored in different fields in the token, adjust the access to `tdata` accordingly. The user and roles can then be used in decorated DBOS workflow and step functions. ```python @app.get("/open/{var1}") @DBOS.required_roles([]) @DBOS.workflow() def test_open_endpoint(var1: str) -> str: # This function can be called with any user/role, or none at all # This is true because: # The `required_roles` list is empty # The middleware above allows the request to be processed even if no token is present pass @app.get("/user/{var1}") @DBOS.required_roles(["user"]) @DBOS.workflow() def test_user_endpoint(var1: str) -> str: # This function can be called only by a user with "user" role # The `required_roles` list contains "user" # Even though the middleware above allows the request to be processed # even if no token is present, DBOS will block the call because the # roles are not set. pass ``` When a call is not authorized, DBOS raises a `DBOSNotAuthorizedError`. --- ## Working With Python Classes You can add DBOS decorators to your Python class instance methods. You can add step decorators to any class methods, but to add a workflow decorator to a class method, its class must inherit from `DBOSConfiguredInstance` and must be decorated with `@DBOS.dbos_class`. For example: ```python @DBOS.dbos_class() class URLFetcher(DBOSConfiguredInstance): def __init__(self, url: str): self.url = url super().__init__(config_name=url) @DBOS.workflow() def fetch_workflow(self): return self.fetch_url() @DBOS.step() def fetch_url(self): return requests.get(self.url).text example_fetcher = URLFetcher("https://example.com") print(example_fetcher.fetch_workflow()) ``` When you create a new instance of a DBOS class, `DBOSConfiguredInstance` must be instantiated with a `config_name`. This `config_name` should be a unique identifier of the instance. Additionally, all DBOS-decorated classes must be instantiated before `DBOS.launch()` is called. The reason for these requirements is to enable workflow recovery. When you create a new instance of a DBOS class, DBOS stores it in a global registry indexed by its class name and `config_name`. When DBOS needs to recover a workflow belonging to that class, it looks up the class instance using `config_name` so it can run the workflow using the right instance of its class. If DBOS classes are dynamically instantiated after `DBOS.launch()`, then DBOS may not find the class instance it needs to recover a workflow. #### Static Methods and Class Methods You can add DBOS decorators to static methods and class methods of any class, even if it does not inherit from `DBOSConfiguredInstance`, because such methods do not access class instance variables. If the class contains a DBOS workflow, it must be decorated with `@DBOS.dbos_class()`. For example: ```python @DBOS.dbos_class() class ExampleClass(): @staticmethod @DBOS.workflow() def staticmethod_workflow(): return @classmethod @DBOS.workflow() def classmethod_workflow(cls): return ``` --- ## DBOS Database Connections DBOS uses a database to durably store workflow and step state. This database is called the **system database**. Its schema is documented [here](../../explanations/system-tables.md). You can use either a SQLite or Postgres database. A SQLite database is just a file on disk, while a Postgres database is a server that your application connects to. By default, DBOS uses SQLite. SQLite is excellent for prototyping and testing because it requires no configuration or server. However, because a SQLite database is just a file on disk, it can't be used in a distributed setting where an application runs on multiple servers. Therefore, **for production, we recommend using Postgres**. ### Configuring the System Database Connection You can configure the database DBOS connects to through the `system_database_url` field of `DBOSConfig`. For example: ```python config: DBOSConfig = { "name": "dbos-example", "application_version": "0.1.0", "system_database_url": os.environ["DBOS_SYSTEM_DATABASE_URL"], } DBOS(config=config) ``` A valid Postgres connection string looks like: ``` postgresql://[username]:[password]@[hostname]:[port]/[database name] ``` For example: ``` postgresql://postgres:dbos@localhost:5432/dbos_example ``` A valid SQLite connection string looks like: ``` sqlite:///[path to database file] ``` For example: ``` sqlite:///dbos_example.sqlite ``` For more information on DBOS configuration, see [the reference](../reference/configuration.md). ### Connecting to an Application Database To run durable database operations in workflows, use [datasources](./transaction-tutorial.md#datasources). You create a datasource with its own database URL, independent of the DBOS system database: ```python import asyncio import os from dbos import SQLAlchemyDatasource, AsyncSQLAlchemyDatasource # Sync ds = SQLAlchemyDatasource.create(os.environ["APP_DATABASE_URL"]) # Async. The datasource is not tied to the event loop that created it, # so it can be used from the event loop that runs your application ads = asyncio.run(AsyncSQLAlchemyDatasource.create(os.environ["APP_DATABASE_URL"])) ``` The datasource manages its own connection pool and can point to any PostgreSQL or SQLite database. To use an async datasource with SQLite, use an async driver URL such as `sqlite+aiosqlite:///app.sqlite`. Your application database does not need to be the same database (or even on the same server) as your system database, and no additional DBOS configuration is needed. See the [datasources tutorial](./transaction-tutorial.md#datasources) for full usage details. --- ## Integrating with Kafka In this guide, you'll learn how to use DBOS workflows to process Kafka messages with exactly-once semantics. First, install [Confluent Kafka](https://docs.confluent.io/kafka-clients/python/current/overview.html) in your application: ``` pip install confluent-kafka ``` Then, define your workflow. It must take in a Kafka message as an input parameter: ```python from dbos import DBOS, KafkaMessage @DBOS.workflow() def test_kafka_workflow(msg: KafkaMessage): DBOS.logger.info(f"Message received: {msg.value.decode()}") ``` Then, annotate your function with a [`@DBOS.kafka_consumer`](../reference/decorators#kafka_consumer) decorator specifying which brokers to connect to and which topics to consume from. Configuration setting details are available from the [Confluent Kafka API docs](https://docs.confluent.io/platform/current/clients/confluent-kafka-python/html/index.html#pythonclient-configuration) and the [official Kafka documentation](https://kafka.apache.org/documentation/#consumerconfigs). At a minimum, you must specify the [`bootstrap.servers`](https://kafka.apache.org/documentation/#consumerconfigs_bootstrap.servers) configuration setting. We also recommend setting [`group.id`](https://kafka.apache.org/documentation/#consumerconfigs_group.id); if you omit it, DBOS generates one from the function name and topics and logs a warning. ```python from dbos import DBOS, KafkaMessage @DBOS.kafka_consumer( config={ "bootstrap.servers": "localhost:9092", "group.id": "dbos-kafka-group", }, topics=["example-topic"], ) @DBOS.workflow() def test_kafka_workflow(msg: KafkaMessage): DBOS.logger.info(f"Message received: {msg.value.decode()}") ``` Under the hood, DBOS constructs an [idempotency key](../tutorials/workflow-tutorial.md#workflow-ids-and-idempotency) for each Kafka message from its topic, partition, consumer group, and offset and passes it into your workflow. This combination is guaranteed to be unique for each Kafka cluster. Thus, even if a message is delivered multiple times (e.g., due to transient network failures or application interruptions), your workflow processes it exactly once. ### In-Order Processing By default, DBOS processes Kafka messages in parallel. You can instead process messages in order using the `ordering` parameter of the `@DBOS.kafka_consumer` decorator: - `ordering="none"` (the default) processes messages in parallel. - `ordering="partition"` processes messages **serially per topic partition**, preserving Kafka's per-partition delivery order, while processing different partitions in parallel. This preserves ordering while still allowing your consumer to scale across partitions. - `ordering="topic"` processes messages **serially per topic**. Only one message from the topic is processed at a time: processing of the next message does not begin until the current one is fully processed. For example, to process each partition's messages in order: ```python from dbos import DBOS, KafkaMessage @DBOS.kafka_consumer( config=config, topics=["example-topic"], ordering="partition", ) @DBOS.workflow() def process_messages_in_order(msg: KafkaMessage): DBOS.logger.info(f"Messages within a partition are processed in order") ``` ### Batching and Throughput DBOS consumes and durably enqueues Kafka messages in batches for higher throughput. You can tune the maximum batch size with the `batch_size` parameter (default 250): ```python @DBOS.kafka_consumer( config=config, topics=["example-topic"], batch_size=500, ) @DBOS.workflow() def process_messages(msg: KafkaMessage): ... ``` For unordered (`ordering="none"`) consumers, you can also name a custom [queue](./queue-tutorial.md) on which to run your consumer workflows, for example to configure concurrency or rate limits: ```python from dbos import DBOS, KafkaMessage @DBOS.kafka_consumer( config=config, topics=["example-topic"], queue_name="kafka_processing_queue", ) @DBOS.workflow() def process_messages(msg: KafkaMessage): ... DBOS.register_queue("kafka_processing_queue", global_concurrency=10) ``` A custom queue is only supported with `ordering="none"`—ordered consumers share an internal partitioned queue—and it must not be a [partitioned queue](./queue-tutorial.md#partitioning-queues). Consumers that don't name a custom queue run on internal queues shared by the whole process. Those queues poll the database every second by default; you can tune that interval with the [`kafka_queue_polling_interval_sec`](../reference/configuration.md#kafka-settings) configuration parameter: ```python DBOS(config={ "name": "kafka-app", "kafka_queue_polling_interval_sec": 0.1, }) ``` Lowering the interval reduces the delay between a message being enqueued and its workflow starting, but requires more frequent database polling. ### Consumer Groups Each consumer's [`group.id`](https://kafka.apache.org/documentation/#consumerconfigs_group.id) determines how Kafka distributes messages among consumers. You can run multiple consumers on the same topics, including with ordering, by giving each a distinct `group.id`. Every consumer group receives its own copy of each message. Two consumers that share both a `group.id` and a topic would each receive only some of that topic's messages, so DBOS raises an error at startup if it detects this configuration. --- ## Logging & Tracing #### Logging For convenience, DBOS provides a pre-configured logger for you to use available at [`DBOS.logger`](../reference/contexts.md#logger). For example: ```python DBOS.logger.info("Welcome to DBOS!") ``` You can [configure](../reference/configuration.md) the log level of this built-in logger through the DBOS constructor. This also configures the log level of the DBOS library. ```python config: DBOSConfig = { "name": "my-app", "application_version": "0.1.0", "log_level": "INFO", } DBOS(config=config) ``` #### Tracing DBOS automatically constructs [OpenTelemetry](https://opentelemetry.io/) [spans](https://opentelemetry.io/docs/concepts/signals/traces/#spans) for every workflow and step. Spans are hierarchical: a step's span is a child of its workflow's span. If the workflow was started from an already-traced operation, such as an instrumented HTTP request, the workflow span is a child of that operation's span and shares its trace. Otherwise, DBOS starts a new trace. Enqueued workflows start a new trace by default; to keep them on the caller's trace, see [Keeping enqueued workflows on the caller's trace](#keeping-enqueued-workflows-on-the-callers-trace). DBOS emits spans on the **global OpenTelemetry tracer**. This means that if your application already sends telemetry to an observability provider through OpenTelemetry, DBOS spans automatically join your existing traces. You don't need to set up a separate export pipeline just for DBOS. :::info OpenTelemetry support in DBOS is optional. To use it, install the DBOS OpenTelemetry dependencies: ``` pip install dbos[otel] ``` You must also enable tracing through the `enable_otlp` flag (this is what tells DBOS to create spans): ```python config: DBOSConfig = { "name": "my-app", "application_version": "0.1.0", "enable_otlp": True, # DBOS creates spans; your provider exports them "otel_attribute_format": "semconv", # emit dbos.* attribute names (recommended) } DBOS(config=config) ``` ::: ##### Connecting DBOS to your observability provider This is the recommended way to use tracing with DBOS. If you already send telemetry to an observability provider (Datadog, Langfuse, Honeycomb, Grafana, Logfire, Jaeger, ...) through OpenTelemetry, you can have DBOS workflow and step spans join your existing traces with no extra setup. There are two steps: 1. Register your provider's OpenTelemetry `TracerProvider` (with `trace.set_tracer_provider(...)`) **before** configuring and launching DBOS. 2. Set `enable_otlp: True` in DBOS configuration. This will cause DBOS to emit spans onto your provider. Set up your provider, then configure and launch DBOS, using whichever option matches your platform: **OpenTelemetry (OTLP)** Most observability platforms, such as Honeycomb, Grafana, Logfire, a self-hosted [OpenTelemetry Collector](https://opentelemetry.io/docs/collector/), or local [Jaeger](https://www.jaegertracing.io/docs/latest/getting-started/), accept the OpenTelemetry Protocol (OTLP). Point an OTLP exporter at your endpoint: ```python from opentelemetry import trace from opentelemetry.sdk.trace import TracerProvider from opentelemetry.sdk.trace.export import BatchSpanProcessor from opentelemetry.exporter.otlp.proto.http.trace_exporter import OTLPSpanExporter from dbos import DBOS, DBOSConfig # Set up your provider provider = TracerProvider() provider.add_span_processor( BatchSpanProcessor(OTLPSpanExporter(endpoint="http://localhost:4318/v1/traces")) ) trace.set_tracer_provider(provider) # Configure and launch DBOS config: DBOSConfig = { "name": "my-app", "application_version": "0.1.0", "enable_otlp": True, # DBOS creates spans; your provider exports them "otel_attribute_format": "semconv", # emit dbos.* attribute names (recommended) } DBOS(config=config) DBOS.launch() ``` **Datadog** [`ddtrace`](https://docs.datadoghq.com/tracing/trace_collection/automatic_instrumentation/dd_libraries/python/) exposes an OpenTelemetry-compatible `TracerProvider`. Register it, and run your app under `ddtrace-run` (or `import ddtrace.auto` first) so `ddtrace` initializes: ```python from ddtrace.opentelemetry import TracerProvider from opentelemetry import trace from dbos import DBOS, DBOSConfig # Set up your provider trace.set_tracer_provider(TracerProvider()) # Configure and launch DBOS config: DBOSConfig = { "name": "my-app", "application_version": "0.1.0", "enable_otlp": True, # DBOS creates spans; your provider exports them "otel_attribute_format": "semconv", # emit dbos.* attribute names (recommended) } DBOS(config=config) DBOS.launch() ``` `ddtrace` then forwards all spans—including DBOS workflow and step spans—to your Datadog agent. **Langfuse** [Langfuse](https://langfuse.com/) is an LLM-observability platform that ingests OpenTelemetry. Point an OTLP exporter at its endpoint, authenticated with your project keys: ```python import base64, os from opentelemetry import trace from opentelemetry.sdk.trace import TracerProvider from opentelemetry.sdk.trace.export import BatchSpanProcessor from opentelemetry.exporter.otlp.proto.http.trace_exporter import OTLPSpanExporter from dbos import DBOS, DBOSConfig # Set up your provider auth = base64.b64encode( f"{os.environ['LANGFUSE_PUBLIC_KEY']}:{os.environ['LANGFUSE_SECRET_KEY']}".encode() ).decode() provider = TracerProvider() provider.add_span_processor( BatchSpanProcessor( OTLPSpanExporter( # Use https://us.cloud.langfuse.com/... for the US region, or your self-hosted host endpoint="https://cloud.langfuse.com/api/public/otel/v1/traces", headers={"Authorization": f"Basic {auth}"}, ) ) ) trace.set_tracer_provider(provider) # Configure and launch DBOS config: DBOSConfig = { "name": "my-app", "application_version": "0.1.0", "enable_otlp": True, # DBOS creates spans; your provider exports them "otel_attribute_format": "semconv", # emit dbos.* attribute names (recommended) } DBOS(config=config) DBOS.launch() ``` Each DBOS workflow becomes a Langfuse trace and each step a nested observation. To record LLM-specific data—model, token usage, prompts—attach attributes to the current span from inside your steps (see [Adding custom attributes and events](#adding-custom-attributes-and-events) below). For the current endpoint and credentials, see the [Langfuse OpenTelemetry docs](https://langfuse.com/docs/opentelemetry/get-started). :::tip Set up your provider before configuring DBOS with `DBOS(config=...)`: DBOS emits spans onto whichever global provider exists and checks for global providers when it is configured. If you enable OTLP without registering a provider for a signal (traces or logs) by then, DBOS logs a warning such as `OTLP is enabled but logger provider not set, skipping log exporter setup` and skips exporter setup for that signal. This is harmless; set a global `LoggerProvider` before configuring DBOS too if you also want to export logs. ::: ##### Adding custom attributes and events Within a workflow or step, you can access the current span and enrich it. This is useful for attaching domain data such as a user ID, a request size, or LLM token usage. Access it via [`DBOS.span`](../reference/contexts.md#span): ```python @DBOS.step() def call_model(prompt: str) -> str: span = DBOS.span # the current OpenTelemetry span if span is not None: span.set_attribute("gen_ai.request.model", "claude-opus-4-8") response = invoke_model(prompt) if span is not None: span.set_attribute("gen_ai.usage.input_tokens", response.usage.input_tokens) span.set_attribute("gen_ai.usage.output_tokens", response.usage.output_tokens) span.add_event("model call complete") return response.text @DBOS.workflow() def llm_workflow(prompt: str) -> str: return call_model(prompt) ``` Your provider exports these attributes on the span alongside DBOS's own, so they appear together in your dashboards. ##### Keeping enqueued workflows on the caller's trace A workflow called or started directly from a traced operation joins that operation's trace automatically. An enqueued workflow, however, may be dequeued much later, potentially by a different process, so by default it starts a new trace. To keep it on the caller's trace, enqueue it inside a [`PropagateOtelContext`](../reference/contexts.md#propagateotelcontext) block: ```python from dbos import PropagateOtelContext with PropagateOtelContext(): handle = DBOS.enqueue_workflow("example_queue", workflow_function, ...) ``` `PropagateOtelContext` durably records the current trace context (or optionally, a passed-in OpenTelemetry context) with every workflow started or enqueued in the block, so each workflow's span joins the caller's trace no matter when or where the workflow runs, even on recovery. The propagated context is not inherited by child workflows; use `PropagateOtelContext` again inside a workflow to keep its children on the same trace. When enqueuing from outside a DBOS application with [DBOS Client](../reference/client.md), pass an OpenTelemetry context through the `otel_context` field of [`EnqueueOptions`](../reference/client.md#enqueue) instead. ##### Letting DBOS export traces directly If you don't already run an observability provider, DBOS can export traces and logs itself to any OpenTelemetry Protocol (OTLP)-compliant receiver. Set `enable_otlp: True` and [configure](../reference/configuration.md) your export endpoints: ```python config: DBOSConfig = { "name": "my-app", "application_version": "0.1.0", "enable_otlp": True, "otlp_traces_endpoints": ["http://localhost:4318/v1/traces"], "otlp_logs_endpoints": ["http://localhost:4318/v1/logs"], } DBOS(config=config) ``` If `otlp_logs_endpoints` is configured, DBOS also exports your [`DBOS.logger`](#logging) logs over OTLP. For example, try using [Jaeger](https://www.jaegertracing.io/docs/latest/getting-started/) to visualize the traces of your local application, or export your logs and traces to [Logfire](../../integrations/logfire). #### Metrics Using [Conductor](../../conductor/overview.md), you can also scrape metrics about your applications' workflows, steps, and executors from a Prometheus-compatible endpoint. See [Metrics](../../conductor/metrics.md) for details. --- ## Queues & Concurrency(3) You can use queues to run many workflows at once with managed concurrency. Queues provide _flow control_, letting you manage how many workflows run at once or how often workflows are started. Register a queue with [`DBOS.register_queue`](../reference/contexts.md#register_queue), specifying its name: ```python from dbos import DBOS DBOS.register_queue("example_queue") ``` Queues are persisted to the system database, so they are visible to every DBOS process and [client](../reference/client.md) connected to that database. If multiple applications [share a system database](../../explanations/sharing-a-system-database.md), each queue is owned by the application that registers it, and only that application dequeues workflows from it. Register your queues after [`DBOS.launch()`](../reference/dbos-class.md#launch). You can then enqueue any DBOS workflow or step. If no queue with that name has been registered, the workflow stays `ENQUEUED` until one is. Enqueuing a function submits it for execution and returns a [handle](../reference/workflow_handles.md) to it. Queued tasks are started in first-in, first-out (FIFO) order. ```python DBOS.register_queue("example_queue") @DBOS.workflow() def process_task(task): ... task = ... handle = DBOS.enqueue_workflow("example_queue", process_task, task) ``` #### Queue Example Here's an example of a workflow using a queue to process tasks concurrently: ```python from dbos import DBOS DBOS.register_queue("example_queue") @DBOS.workflow() def process_task(task): ... @DBOS.workflow() def process_tasks(tasks): task_handles = [] # Enqueue each task so all tasks are processed concurrently. for task in tasks: handle = DBOS.enqueue_workflow("example_queue", process_task, task) task_handles.append(handle) # Wait for each task to complete and retrieve its result. # Return the results of all tasks. return [handle.get_result() for handle in task_handles] ``` Sometimes, you may wish to receive the result of each task as soon as it's ready instead of waiting for all tasks to complete. You can do this using [`DBOS.wait_first`](../reference/contexts.md#wait_first), which waits for any one of a list of workflow handles to complete and returns the first completed handle. ```python @DBOS.workflow() def process_task(task: Task): result = ... # Process the task return result @DBOS.workflow() def process_tasks(tasks: List[Task]): handles = [DBOS.enqueue_workflow("example_queue", process_task, task) for task in tasks] results = [] remaining = list(handles) while remaining: # Wait for any task to complete completed = DBOS.wait_first(remaining) result = completed.get_result() print(f"Task completed. Result: {result}") results.append(result) # Remove the completed handle remaining = [h for h in remaining if h.workflow_id != completed.workflow_id] return results ``` #### Enqueueing from Another Application Often, you want to enqueue a workflow from another DBOS application or from outside your DBOS application. For example, let's say you have an API server and a data processing service. You're using DBOS to build a durable data pipeline in the data processing service. When the API server receives a request, it should enqueue the data pipeline for execution on the data processing service. You can use the [DBOS Client](../reference/client.md) to register queues and enqueue workflows from outside your DBOS application by connecting directly to your system database. Since the DBOS Client is designed to be used from outside your DBOS application, workflow and queue metadata must be specified explicitly. ```python from dbos import DBOSClient, EnqueueOptions client = DBOSClient( system_database_url=os.environ["DBOS_SYSTEM_DATABASE_URL"], # The name of the application that runs the data pipeline application_name="data-processing-service", ) # Register the queue from the client. client.register_queue("pipeline_queue") options: EnqueueOptions = { "queue_name": "pipeline_queue", "workflow_name": "data_pipeline", } handle = client.enqueue(options, task) result = handle.get_result() ``` The [queue worker](../examples/queue-worker.md) example shows this design pattern in more detail. #### Enqueueing from PL/pgSQL You can also enqueue a workflow from a Postgres trigger or stored procedure. The DBOS System Database includes an [`enqueue_workflow`](../../explanations/system-tables.md#dbosenqueue_workflow) method for this scenario. For example, here is the previous example of enqueueing the `data_pipeline` workflow on the `pipeline_queue` queue with arguments, but using PL/pgSQL. ```sql DECLARE workflow_id text; workflow_id := dbos.enqueue_workflow( workflow_name => 'data_pipeline', queue_name => 'pipeline_queue', positional_args => ARRAY[ '"task-123"'::json, '"data"'::json ] ); ``` #### Managing Concurrency You can control how many workflows from a queue run simultaneously by configuring concurrency limits. This helps prevent resource exhaustion when workflows consume significant memory or processing power. ##### Worker Concurrency Worker concurrency sets the maximum number of workflows from a queue that can run concurrently on a single DBOS process. This is particularly useful for resource-intensive workflows to avoid exhausting the resources of any process. For example, this queue has a worker concurrency of 5, so each process will run at most 5 workflows from this queue simultaneously: ```python DBOS.register_queue("example_queue", worker_concurrency=5) ``` ##### Global Concurrency Global concurrency limits the total number of workflows from a queue that can run concurrently across all DBOS processes in your application. For example, this queue will have a maximum of 10 workflows running simultaneously across your entire application. :::warning Worker concurrency limits are recommended for most use cases. Take care when using a global concurrency limit as any `PENDING` workflow on the queue counts toward the limit, including workflows from previous application versions ::: ```python DBOS.register_queue("example_queue", global_concurrency=10) ``` ##### In-Order Processing You can use a queue with `global_concurrency=1` to guarantee sequential, in-order processing of events. Only a single event will be processed at a time. For example, this app processes events sequentially in the order of their arrival: ```python from dbos import DBOS DBOS.register_queue("in_order_queue", global_concurrency=1) @DBOS.step() def process_event(event: str): ... def event_endpoint(event: str): DBOS.enqueue_workflow("in_order_queue", process_event, event) ``` #### Rate Limiting You can set _rate limits_ for a queue, limiting the number of functions that it can start in a given period. Rate limits are global across all DBOS processes using this queue. For example, this queue has a limit of 50 with a period of 30 seconds, so it may not start more than 50 functions in 30 seconds: ```python DBOS.register_queue("example_queue", limiter={"limit": 50, "period": 30}) ``` Rate limits are especially useful when working with a rate-limited API, such as many LLM APIs. #### Reconfiguring Queues at Runtime Because queue configuration lives in the system database, you can change a queue's configuration at runtime without redeploying or restarting your workers. Use [`DBOS.retrieve_queue`](../reference/contexts.md#retrieve_queue) to fetch a queue, then call its [`set_*`](../reference/queues.md#reconfiguring-queues) methods. Workers pick up the new configuration on their next polling iteration. ```python queue = DBOS.retrieve_queue("example_queue") # Change the queue's global concurrency. queue.set_global_concurrency(20) # Change its rate limit. queue.set_limiter({"limit": 25, "period": 30}) ``` You can also do this from a [`DBOSClient`](../reference/client.md), which is useful for managing queues from an admin tool or another service. :::warning If your application calls [`DBOS.register_queue`](../reference/contexts.md#register_queue) on startup, the next process to start can overwrite settings you applied at runtime via `set_*` methods. Either update the `register_queue` call to match the new configuration, or pass `on_conflict="never_update"` to preserve the runtime changes. ::: ### Setting Timeouts You can set a timeout for an enqueued workflow with [`SetWorkflowTimeout`](../reference/contexts.md#setworkflowtimeout). When the timeout expires, the workflow **and all its children** are cancelled. Cancelling a workflow sets its status to `CANCELLED` and preempts its execution at the beginning of its next step. Timeouts are **start-to-completion**: a workflow's timeout does not begin until the workflow is dequeued and starts execution. Also, timeouts are **durable**: they are stored in the database and persist across restarts, so workflows can have very long timeouts. Example syntax: ```python @DBOS.workflow() def example_workflow(): ... DBOS.register_queue("example-queue") # If the workflow does not complete within 10 seconds after being dequeued, it times out and is cancelled with SetWorkflowTimeout(10): DBOS.enqueue_workflow("example-queue", example_workflow) ``` ### Partitioning Queues You can **partition** queues to distribute work across dynamically created queue partitions. A queue is partitioned if you register it with any per-partition flow control limit: | Parameter | Meaning | | --- | --- | | `partition_concurrency` | Maximum workflows from any one partition running at once across all processes. | | `partition_worker_concurrency` | Maximum workflows from any one partition running at once on a single process. | | `partition_limiter` | Maximum workflows that may be started from any one partition in a given period. | When you enqueue a workflow on a partitioned queue, you must supply a queue partition key. Essentially, you can think of each partition as a "subqueue" you dynamically create by enqueueing a workflow with a partition key. For example, suppose you want your users to each be able to run at most one task at a time. You can do this with a queue whose `partition_concurrency` is 1, where the partition key is user ID. **Example Syntax** ```python DBOS.register_queue("partitioned_queue", partition_concurrency=1) @DBOS.workflow() def process_task(task: Task): ... def on_user_task_submission(user_id: str, task: Task): # Partition the task queue by user ID. As the queue has a # per-partition concurrency of 1, this means that at most one # task can run at once per user (but tasks from different # users can run concurrently). with SetEnqueueOptions(queue_partition_key=user_id): DBOS.enqueue_workflow("partitioned_queue", process_task, task) ``` :::warning Every enqueue on a partitioned queue must supply a partition key. A workflow enqueued on a partitioned queue without a partition key stays `ENQUEUED` and is not dequeued. ::: #### Combining Queue-Wide and Per-Partition Limits A partitioned queue enforces its per-partition limits **and** its queue-wide limits ([`global_concurrency`](#global-concurrency), [`worker_concurrency`](#worker-concurrency), and [`limiter`](#rate-limiting)) at the same time. This lets you protect your workers from overload while still fairly distributing work between partitions. For example, this "fair queue" runs at most one task per user, but no more than 10 tasks on any single process: ```python DBOS.register_queue("fair_queue", partition_concurrency=1, worker_concurrency=10) ``` Each queue-wide limit has a per-partition counterpart, so you can mix and match them freely: ```python # At most 100 tasks running globally and 25 running per tenant, # at most 10 tasks running per process and 2 per tenant per process, # and at most 1000 tasks started per minute globally and 50 per tenant. DBOS.register_queue( "tenant_queue", global_concurrency=100, worker_concurrency=10, limiter={"limit": 1000, "period": 60}, partition_concurrency=25, partition_worker_concurrency=2, partition_limiter={"limit": 50, "period": 60}, ) ``` Each per-partition concurrency limit must be less than or equal to its queue-wide counterpart, and `partition_worker_concurrency` must be less than or equal to `partition_concurrency`. :::note [Deduplication](#deduplication) is not supported on partitioned queues. ::: ### Deduplication You can set a deduplication ID for an enqueued workflow with [`SetEnqueueOptions`](../reference/queues.md#setenqueueoptions). At any given time, only one workflow with a specific deduplication ID can be enqueued in the specified queue. If a workflow with a deduplication ID is currently enqueued, delayed, or actively executing (status `ENQUEUED`, `DELAYED`, or `PENDING`), subsequent workflow enqueue attempt with the same deduplication ID in the same queue will raise a `DBOSQueueDeduplicatedError` exception. For example, this is useful if you only want to have one workflow active at a time per user—set the deduplication ID to the user's ID. Example syntax: ```python from dbos import DBOS, SetEnqueueOptions from dbos import error as dboserror DBOS.register_queue("example_queue") with SetEnqueueOptions(deduplication_id="my_dedup_id"): try: handle = DBOS.enqueue_workflow("example_queue", example_workflow, ...) except dboserror.DBOSQueueDeduplicatedError as e: # Handle deduplication error ... ``` ### Singleton Workflows If you want only one instance of a workflow to be active at a time, you can set `duplication_policy="return-existing"` on [`SetEnqueueOptions`](../reference/queues.md#setenqueueoptions). When a workflow with the same `deduplication_id` is already enqueued, delayed, or executing on the queue, this returns a handle to that existing workflow instead of raising `DBOSQueueDeduplicatedError`. The arguments passed by the colliding caller are discarded, and the returned handle resolves with the original workflow's result. This requires both a queue and a `deduplication_id`. Once the original workflow completes, it releases its deduplication ID, so the next caller starts a new workflow. Example syntax: ```python from dbos import DBOS, SetEnqueueOptions DBOS.register_queue("example_queue") # Only one workflow with deduplication ID "singleton" can be active on this queue # at a time. Subsequent callers attach to it and receive its result. with SetEnqueueOptions( deduplication_id="singleton", duplication_policy="return-existing" ): handle = DBOS.enqueue_workflow("example_queue", example_workflow, ...) result = handle.get_result() ``` ### Priority You can set a priority for an enqueued workflow with [`SetEnqueueOptions`](../reference/queues.md#setenqueueoptions). Workflows with the same priority are dequeued in **FIFO (first in, first out)** order. Priority values can range from `1` to `2,147,483,647`, where **a low number indicates a higher priority**. :::tip Workflows without assigned priorities have the highest priority and are dequeued before workflows with assigned priorities. ::: Example syntax: ```python DBOS.register_queue("priority_queue") with SetEnqueueOptions(priority=10): # All workflows are enqueued with priority set to 10 # They will be dequeued in FIFO order for task in tasks: DBOS.enqueue_workflow("priority_queue", task_workflow, task) # first_workflow (priority=1) will be dequeued before all task_workflows (priority=10) with SetEnqueueOptions(priority=1): DBOS.enqueue_workflow("priority_queue", first_workflow) ``` ### Delayed Execution You can delay an enqueued workflow by a specified number of seconds using [`SetEnqueueOptions`](../reference/queues.md#setenqueueoptions). The workflow is initially placed in `DELAYED` status and does not execute. After the delay expires, it transitions to `ENQUEUED` status and may be dequeued and executed. This is useful for scheduling workflows to run at a future time. Example syntax: ```python from dbos import DBOS, SetEnqueueOptions DBOS.register_queue("example_queue") @DBOS.workflow() def send_reminder(user_id: str): ... # Send a reminder in one hour with SetEnqueueOptions(delay_seconds=3600): handle = DBOS.enqueue_workflow("example_queue", send_reminder, user_id) ``` You can also dynamically update the delay of a `DELAYED` workflow using [`DBOS.set_workflow_delay`](../reference/contexts.md#set_workflow_delay): ```python # Shorten the delay to 10 seconds from now DBOS.set_workflow_delay(handle.workflow_id, delay_seconds=10) # Or set an absolute deadline DBOS.set_workflow_delay(handle.workflow_id, delay_until_epoch_ms=int((time.time() + 60) * 1000)) ``` ### Explicit Queue Listening By default, a process running DBOS listens to (dequeues workflows from) all queues owned by its application in its system database. However, sometimes you only want a process to listen to a specific list of queues. You can use [`DBOS.listen_queues`](../reference/dbos-class.md#listen_queues) to explicitly tell a process running DBOS to only listen to a specific set of queues. You must call `DBOS.listen_queues` before DBOS is launched. This is particularly useful when managing heterogeneous workers, where specific tasks should execute on specific physical servers. For example, say you have a mix of CPU workers and GPU workers and you want CPU tasks to only execute on CPU workers and GPU tasks to only execute on GPU workers. You can configure each type of worker to only listen to the appropriate queue: ```python if __name__ == "__main__": worker_type = ... # "cpu" or "gpu" config: DBOSConfig = ... DBOS(config=config) if worker_type == "gpu": # GPU workers will only dequeue and execute workflows from the GPU queue DBOS.listen_queues(["gpu_queue"]) elif worker_type == "cpu": # CPU workers will only dequeue and execute workflows from the CPU queue DBOS.listen_queues(["cpu_queue"]) DBOS.launch() DBOS.register_queue("cpu_queue") DBOS.register_queue("gpu_queue") ``` Note that `DBOS.listen_queues` only controls what workflows are dequeued, not what workflows can be enqueued, so you can freely enqueue tasks onto the GPU queue from a CPU worker for execution on a GPU worker, and vice versa. --- ## Scheduling Workflows(Tutorials) You can schedule DBOS [workflows](./workflow-tutorial.md) to run on a cron schedule. Schedules are stored in the database and can be created, paused, resumed, and deleted at runtime. Each time a schedule fires, its workflow is executed by exactly one worker process. To schedule a workflow, first define a workflow that takes two arguments: a `datetime` (the scheduled execution time) and a context object: ```python from datetime import datetime from typing import Any from dbos import DBOS @DBOS.workflow() def my_periodic_task(scheduled_time: datetime, context: Any): DBOS.logger.info(f"Running task scheduled for {scheduled_time} with context {context}") ``` Then, create a schedule for it using [`DBOS.create_schedule`](../reference/contexts.md#create_schedule) with a [crontab](https://en.wikipedia.org/wiki/Cron) expression: ```python DBOS.create_schedule( schedule_name="my-task-schedule", # The schedule name is a unique identifier of the schedule workflow_fn=my_periodic_task, schedule="*/5 * * * *", # Every 5 minutes context="my context", # The context is passed into every iteration of the workflow ) ``` Because schedules are stored in the system database, `DBOS.create_schedule` and the other schedule management methods must be called after [`DBOS.launch()`](../reference/dbos-class.md#launch). Note that `DBOS.create_schedule` will fail if the schedule already exists. If you're defining a set of static schedules to be created on program start, you can instead use `DBOS.apply_schedules` to create them atomically, updating them if they already exist: ```python DBOS.apply_schedules([ { "schedule_name": "schedule-a", "workflow_fn": workflow_a, "schedule": "*/10 * * * *", # Every 10 minutes "context": "context-a", }, { "schedule_name": "schedule-b", "workflow_fn": workflow_b, "schedule": "0 0 * * *", # Every day at midnight "context": "context-b", }, ]) ``` When `DBOS.apply_schedules` updates an existing schedule, it replaces the entire definition with the new entry, so any optional field left unset is cleared. For example, if a schedule was routed to a named queue and you re-apply it without setting `queue_name`, it reverts to the internal queue. The schedule's status and last-fired time are preserved. To learn more about crontab syntax, see [this guide](https://docs.gitlab.com/ee/topics/cron/) or [this crontab editor](https://crontab.guru/). DBOS uses [croniter](https://pypi.org/project/croniter/) to parse cron schedules, using seconds as an optional first field ([`second_at_beginning=True`](https://pypi.org/project/croniter/#about-second-repeats)). Valid cron schedules contain 5 or 6 items, separated by spaces: ``` ┌────────────── second (optional) │ ┌──────────── minute │ │ ┌────────── hour │ │ │ ┌──────── day of month │ │ │ │ ┌────── month │ │ │ │ │ ┌──── day of week │ │ │ │ │ │ │ │ │ │ │ │ * * * * * * ``` Cron expressions are evaluated in UTC by default. You can set the `cron_timezone` parameter to an [IANA timezone name](https://en.wikipedia.org/wiki/List_of_tz_database_time_zones) (e.g. `"America/New_York"`) to evaluate the expression in a different timezone. You can dynamically create many schedules for the same workflow. For example, if you want to perform certain actions periodically for each of your customers, you can create one schedule per customer, using customer ID as context so each workflow knows which customer to act on: ```python from datetime import datetime from dbos import DBOS @DBOS.workflow() def customer_workflow(scheduled_time: datetime, customer_id: str): ... def on_customer_registration(customer_id: str): DBOS.create_schedule( schedule_name=f"customer-{customer_id}-sync", workflow_fn=customer_workflow, schedule="0 * * * *", # Every hour context=customer_id, ) ``` Note that scheduling is not supported for workflows that are methods on [configured instances](./classes.md). Scheduled workflows should be plain functions or `@staticmethod` or `@classmethod` class members. #### Managing Schedules You can pause, resume, and delete schedules at runtime: ```python # Pause a schedule so it stops firing DBOS.pause_schedule("my-task-schedule") # Resume a paused schedule DBOS.resume_schedule("my-task-schedule") # Delete a schedule DBOS.delete_schedule("my-task-schedule") ``` You can also list and inspect schedules: ```python # List all active schedules schedules = DBOS.list_schedules(status="ACTIVE") # Get a specific schedule by name schedule = DBOS.get_schedule("my-task-schedule") ``` Each workflow enqueued by a schedule is tagged with that schedule's name, which is recorded in its [`WorkflowStatus`](../reference/contexts.md#workflow-status) and is queryable. You can retrieve all runs of a given schedule by passing `schedule_name` to [`DBOS.list_workflows`](../reference/contexts.md#list_workflows): ```python # Retrieve all workflows enqueued by a schedule runs = DBOS.list_workflows(schedule_name="my-task-schedule") ``` #### Backfilling and Triggering If a schedule was paused or your application was offline, you can backfill missed executions using [`DBOS.backfill_schedule`](../reference/contexts.md#backfill_schedule). Already-executed times are automatically skipped: ```python from datetime import datetime, timezone DBOS.backfill_schedule( "my-task-schedule", start=datetime(2025, 1, 1, tzinfo=timezone.utc), end=datetime(2025, 1, 2, tzinfo=timezone.utc), ) ``` Alternatively, you can set `automatic_backfill=True` when creating a schedule so that missed executions are automatically backfilled whenever your application starts or a paused schedule is resumed. Backfills (manual or automatic) compute missed executions using the schedule's **current** cron expression. If you update a schedule's cron expression and then backfill, the backfill generates one execution per tick of the new expression over the requested window—including times the old expression would never have matched. For example, changing a daily schedule to an hourly one and then backfilling yesterday enqueues 24 executions, not 1. You can also immediately trigger a schedule using [`DBOS.trigger_schedule`](../reference/contexts.md#trigger_schedule): ```python handle = DBOS.trigger_schedule("my-task-schedule") ``` #### Scheduling to Queues By default, scheduled workflows are enqueued on an internal queue. You can instead enqueue them on a declared [queue](./queue-tutorial.md) to manage their concurrency or rate limits. Pass the `queue_name` parameter when creating the schedule: ```python from dbos import DBOS DBOS.register_queue("scheduled_queue", global_concurrency=1) DBOS.create_schedule( schedule_name="my-task-schedule", workflow_fn=my_periodic_task, schedule="*/5 * * * *", queue_name="scheduled_queue", ) ``` This ensures that scheduled workflow executions respect the queue's flow control settings. #### Managing Schedules from Another Application You can manage schedules from outside your DBOS application using the [DBOS Client](../reference/client.md#workflow-schedules). The client accepts workflow names as strings instead of function references: ```python from dbos import DBOSClient client = DBOSClient( system_database_url=os.environ["DBOS_SYSTEM_DATABASE_URL"], # The name of the application that owns and runs the schedule application_name="my-app", ) client.create_schedule( schedule_name="my-task-schedule", workflow_name="my_periodic_task", schedule="*/5 * * * *", context="my context", ) ``` #### How Scheduling Works Under the hood, DBOS constructs an [idempotency key](./workflow-tutorial.md#workflow-ids-and-idempotency) for each scheduled workflow execution. The key is a concatenation of the schedule name and the scheduled time, ensuring each scheduled invocation occurs exactly once while your application is active. For the full API reference, see [Workflow Schedules](../reference/contexts.md#workflow-schedules). --- ## Steps(3) When using DBOS workflows, you should annotate any function that performs complex operations or accesses external APIs or services as a _step_. If a workflow is interrupted, upon restart it automatically resumes execution from the **last completed step**. You can turn **any** Python function into a step by annotating it with the [`@DBOS.step`](../reference/decorators.md#step) decorator. The only requirement is that its outputs should be serializable. Here's a simple example: ```python @DBOS.step() def example_step(): return requests.get("https://example.com").text ``` You should make a function a step if you're using it in a DBOS workflow and it performs a [**nondeterministic**](../tutorials/workflow-tutorial.md#determinism) operation. A nondeterministic operation is one that may return different outputs given the same inputs. Common nondeterministic operations include: - Accessing an external API or service, like serving a file from [AWS S3](https://aws.amazon.com/s3/), calling an external API like [Stripe](https://stripe.com/), or accessing an external data store like [Elasticsearch](https://www.elastic.co/elasticsearch/). - Accessing files on disk. - Generating a random number. - Getting the current time. You **cannot** call, start, or enqueue workflows from within steps. These operations should be performed from workflow functions. You can call one step from another step, but the called step becomes part of the calling step's execution rather than functioning as a separate step. If you call a step from outside a workflow, it runs as an ordinary function, without checkpoints, retries, or timeouts. #### Configurable Retries You can optionally configure a step to automatically retry any exception a set number of times with exponential backoff. This is useful for automatically handling transient failures, like making requests to unreliable APIs. Retries are configurable through arguments to the [step decorator](../reference/decorators.md#step): ```python DBOS.step( retries_allowed: bool = False, interval_seconds: float = 1.0, max_attempts: int = 3, backoff_rate: float = 2.0, should_retry: Optional[Callable[[BaseException], Union[bool, Awaitable[bool]]]] = None, ) ``` For example, we configure this step to retry exceptions (such as if `example.com` is temporarily down) up to 10 times: ```python @DBOS.step(retries_allowed=True, max_attempts=10) def example_step(): return requests.get("https://example.com").text ``` If a step fails on all `max_attempts` attempts, it throws an exception (`DBOSMaxStepRetriesExceeded`) to the calling workflow. If that exception is not caught, the workflow [terminates](./workflow-tutorial.md). ##### Filtering Retries With `should_retry` By default, every exception raised by the step is retried until `max_attempts` is reached. If you only want to retry certain exceptions; for example, transient network errors but not validation failures, pass a `should_retry` predicate. The predicate receives the raised exception. If it returns `False`, the exception is re-raised immediately and no further retries are attempted. ```python @DBOS.step( retries_allowed=True, max_attempts=10, should_retry=lambda e: not isinstance(e, FatalError), ) def example_step(): return requests.get("https://example.com").text ``` For async steps, `should_retry` may itself be an `async` function: ```python async def is_retryable(e: BaseException) -> bool: return not isinstance(e, FatalError) @DBOS.step(retries_allowed=True, max_attempts=10, should_retry=is_retryable) async def example_step(): ... ``` Async predicates are only supported for async steps; pairing an async `should_retry` with a sync step raises an exception. #### Coroutine Steps You may also decorate coroutines (functions defined with `async def`, also known as async functions) with `@DBOS.step`. Coroutine steps can use Python's asynchronous language capabilities such as [await](https://docs.python.org/3/reference/expressions.html#await), [async for](https://docs.python.org/3/reference/compound_stmts.html#async-for) and [async with](https://docs.python.org/3/reference/compound_stmts.html#async-with). Like synchronous step functions, async steps support [configurable automatic retries](#configurable-retries) and require their inputs and outputs to be serializable. For example, here is an asynchronous version of the `example_step` function from above, using the [`aiohttp`](https://docs.aiohttp.org/en/stable/) library instead of [`requests`](https://requests.readthedocs.io/en/latest/). ```python @DBOS.step(retries_allowed=True, max_attempts=10) async def example_step(): async with aiohttp.ClientSession() as session: async with session.get("https://example.com") as response: return await response.text() ``` #### Step Timeouts You can set a timeout for an async step with `timeout_seconds`. If the step runs longer than that, it is cancelled and `DBOSStepTimeoutError` is raised to the calling workflow. This is useful for bounding a step that may hang, such as a request to an unresponsive service. ```python @DBOS.step(timeout_seconds=30) async def example_step(): async with aiohttp.ClientSession() as session: async with session.get("https://example.com") as response: return await response.text() ``` Step timeouts are only supported for [async steps](#coroutine-steps), because Python provides no way to preempt a running synchronous function. If the step also has [retries](#configurable-retries) enabled, **each attempt gets its own timeout**, and time spent waiting between retries is not counted against it. A step with `timeout_seconds=30, max_attempts=3` therefore allows up to three 30-second attempts, not 30 seconds total. You can configure this behavior with a `should_retry` predicate, not retrying `DBOSStepTimeoutError`. #### Running Steps In-Line With `run_step` If a function is not decorated with `@DBOS.step` and you would prefer not to wrap it, you can call the code as a step using [`DBOS.run_step`](../reference/contexts.md#run_step) (or `DBOS.run_step_async`). For example, if your code said: ```python res = send_email(user, msg) ``` It could be quickly changed to a checkpointed step: ```python res = DBOS.run_step(None, send_email, user, msg) ``` Or: ```python res = DBOS.run_step({"name": "send_email_to_user"}, lambda: send_email(user, msg)) ``` --- ## Testing Your App #### Testing DBOS Functions Because DBOS workflows, steps, and transactions are ordinary Python functions, you can unit test them using any Python testing framework, like [pytest](https://docs.pytest.org/en/stable/) or [unittest](https://docs.python.org/3/library/unittest.html). **You must reset the DBOS runtime between each test like this:** ```python def reset_dbos(): DBOS.destroy() config: DBOSConfig = { "name": "my-app", "application_version": "0.1.0", "system_database_url": os.environ.get("TESTING_DATABASE_URL"), } DBOS(config=config) DBOS.reset_system_database(truncate=True) DBOS.launch() ``` :::tip To minimize dependencies during testing, you may want to use a SQLite system database instead of Postgres. You can do this by setting a SQLite connection string in `reset_dbos`, for example: ```python def reset_dbos(): DBOS.destroy() config: DBOSConfig = { "name": "my-app", "application_version": "0.1.0", "system_database_url": "sqlite:///my_test_db.sqlite", } DBOS(config=config) DBOS.reset_system_database(truncate=True) DBOS.launch() ``` ::: First, destroy any existing DBOS instance. Then, create and configure a new DBOS instance (you may want to use a different database for testing). Next, reset the internal state of DBOS, cleaning up any state left over from previous tests. We recommend passing [`truncate=True`](../reference/dbos-class.md#reset_system_database), which empties the DBOS system tables instead of dropping and re-creating the system database, as this is substantially faster. Finally, launch a new DBOS instance. For example, if using pytest, declare `reset_dbos` as a fixture and require it from every test of a DBOS function: ```python title="conftest.py" import os import pytest from dbos import DBOS, DBOSConfig @pytest.fixture() def reset_dbos(): DBOS.destroy() config: DBOSConfig = { "name": "my-app", "application_version": "0.1.0", "system_database_url": os.environ.get("TESTING_DATABASE_URL"), } DBOS(config=config) DBOS.reset_system_database(truncate=True) DBOS.launch() ``` ```python title="test_example.py" from example_app.main import example_workflow def test_example_workflow(reset_dbos): example_input = ... example_output = ... assert example_workflow(example_input) == example_output ``` #### Mocking It is often useful in testing to mock your workflows and steps. Because workflows and steps are just Python functions, they can be mocked using popular mocking libraries like [unittest.mock](https://docs.python.org/3/library/unittest.mock.html). For example, the [widget store](../examples/widget-store.md) app has a `checkout_workflow` that creates an order, reserves inventory, waits for a payment notification, then either marks the order paid and starts a dispatch workflow or returns the reserved inventory and cancels the order: ```python @DBOS.workflow() def checkout_workflow(): # Create a new order order_id = ds.run_tx_step({"name": "create_order"}, create_order) # Attempt to reserve inventory, cancelling the order if no inventory remains. inventory_reserved = ds.run_tx_step( {"name": "reserve_inventory"}, reserve_inventory ) if not inventory_reserved: ds.run_tx_step( {"name": "update_order_status"}, update_order_status, order_id=order_id, status=OrderStatus.CANCELLED.value, ) DBOS.set_event(PAYMENT_ID, None) return # Send a unique payment ID to the checkout endpoint, then wait for # a message that the customer has completed payment. DBOS.set_event(PAYMENT_ID, DBOS.workflow_id) payment_status = DBOS.recv(PAYMENT_STATUS) if payment_status == "paid": ds.run_tx_step( {"name": "update_order_status"}, update_order_status, order_id=order_id, status=OrderStatus.PAID.value, ) DBOS.start_workflow(dispatch_order_workflow, order_id) else: ds.run_tx_step({"name": "undo_reserve_inventory"}, undo_reserve_inventory) ds.run_tx_step( {"name": "update_order_status"}, update_order_status, order_id=order_id, status=OrderStatus.CANCELLED.value, ) DBOS.set_event(ORDER_ID, str(order_id)) ``` We can test the workflow in isolation by mocking its database operations, the payment message it waits for, and the workflow it starts: ```python from unittest.mock import MagicMock, patch from dbos import DBOS import widget_store.main as widget_store from widget_store.schema import OrderStatus def test_checkout_workflow(reset_dbos): """ Use mocks to test that the main workflow function (checkout_workflow) correctly handles a checkout whose payment succeeds. """ order_id = 123 # Create a mock for each of the workflow's database transactions mock_create_order = MagicMock(return_value=order_id) mock_reserve_inventory = MagicMock(return_value=True) mock_undo_reserve_inventory = MagicMock() mock_update_order_status = MagicMock() mocks = { widget_store.create_order: mock_create_order, widget_store.reserve_inventory: mock_reserve_inventory, widget_store.undo_reserve_inventory: mock_undo_reserve_inventory, widget_store.update_order_status: mock_update_order_status, } # Run each transaction against its mock instead of against the database. # The app assigns `ds` at startup, so the test supplies its own datasource. def run_mocked_tx_step(ds_options, func, *args, **kwargs): return mocks[func](*args, **kwargs) mock_ds = MagicMock() mock_ds.run_tx_step.side_effect = run_mocked_tx_step # Also mock the payment message the workflow waits for and the # dispatch workflow it starts, then run the workflow. with ( patch.object(widget_store, "ds", mock_ds, create=True), patch.object(DBOS, "recv", return_value="paid") as mock_recv, patch.object(DBOS, "start_workflow") as mock_start_workflow, ): widget_store.checkout_workflow() # Verify an order was created and inventory was reserved for it mock_create_order.assert_called_once_with() mock_reserve_inventory.assert_called_once_with() # Verify the workflow waited for the payment webhook mock_recv.assert_called_once_with(widget_store.PAYMENT_STATUS) # Verify the paid order was marked paid and handed to the dispatch workflow mock_update_order_status.assert_called_once_with( order_id=order_id, status=OrderStatus.PAID.value ) mock_start_workflow.assert_called_once_with( widget_store.dispatch_order_workflow, order_id ) # Verify that because payment succeeded, inventory was never returned mock_undo_reserve_inventory.assert_not_called() ``` #### Example Test Suite To see a DBOS app tested using pytest, check out the [widget store](https://github.com/dbos-inc/dbos-demo-apps/tree/main/python/widget-store) example on GitHub. --- ## Transactions & Datasources(Tutorials) DBOS runs database operations durably inside workflows through _datasources_. Datasources connect to any PostgreSQL or SQLite database, support both sync and async transaction functions, and integrate with DBOS's exactly-once execution guarantees. ### Datasources Datasources wrap a SQLAlchemy engine with DBOS transaction tracking, ensuring that each database operation inside a workflow runs exactly once even if the workflow is interrupted and retried. #### Creating a Datasource Create a datasource by calling the `create` factory method with a database URL. By default, the factory creates or migrates the `datasource_outputs` tracking table in the target database (see [Running Datasource Migrations Separately](#running-datasource-migrations-separately) if your application's database role cannot run DDL). Use `SQLAlchemyDatasource` for synchronous (non-async) code and `AsyncSQLAlchemyDatasource` for async code: ```python import os from dbos import SQLAlchemyDatasource ds = SQLAlchemyDatasource.create(os.environ["APP_DATABASE_URL"]) ``` ```python import asyncio import os from dbos import AsyncSQLAlchemyDatasource # The datasource is not tied to the event loop that created it, # so it can be used from the event loop that runs your application. ads = asyncio.run(AsyncSQLAlchemyDatasource.create(os.environ["APP_DATABASE_URL"])) ``` Create all your datasources before calling [`DBOS.launch()`](../reference/dbos-class.md#launch): creating a datasource after launch raises a `DBOSException`. DBOS tracks every datasource created in the process so that [rewinding a workflow](./workflow-management.md#rewinding-workflows) also deletes the transaction checkpoints it holds. To use `AsyncSQLAlchemyDatasource` with SQLite, you must use an async driver URL such as `sqlite+aiosqlite:///app.sqlite` (install the driver with `pip install "dbos[aiosqlite]"`); a plain `sqlite:///` URL raises an error. :::warning Due to the nature of SQLAlchemy's object model, `AsyncSQLAlchemyDatasource` only supports coroutine functions (`async def`) and `SQLAlchemyDatasource` only supports regular synchronous functions. Decorating the wrong function type raises a `DBOSException` at decoration time. ::: Both `create` methods take a required `database_url` and accept optional arguments for advanced configuration: | Parameter | Type | Description | |---|---|---| | `database_url` | `str` | SQLAlchemy-compatible database URL (required). DBOS connects to Postgres with the psycopg driver. | | `engine_kwargs` | `dict` | Extra kwargs forwarded to SQLAlchemy's `create_engine` / `create_async_engine` | | `engine` | `Engine` / `AsyncEngine` | Provide your own SQLAlchemy engine instead of creating one | | `schema` | `str` | Postgres schema name for the `datasource_outputs` table (defaults to `"dbos"`; ignored for SQLite) | | `serializer` | `Serializer` | Custom serializer for transaction outputs | | `sessionmaker` | `sessionmaker` / `async_sessionmaker` | Custom SQLAlchemy sessionmaker for transaction sessions, for example to use a custom `Session` subclass or event hooks. Sessions are always bound to the datasource's engine. | | `run_migrations` | `bool` | Whether to create and migrate the `datasource_outputs` table (defaults to `True`). If `False`, only verify it is migrated. | #### Running Datasource Migrations Separately By default, creating a datasource creates its `datasource_outputs` table (and the schema containing it) if they do not exist, which requires DDL privileges. If your application's database role should not have those privileges, run the datasource migrations separately with a privileged role, using [`SQLAlchemyDatasource.migrate`](../reference/datasources.md#sqlalchemydatasourcemigrate) (or [`AsyncSQLAlchemyDatasource.migrate`](../reference/datasources.md#asyncsqlalchemydatasourcemigrate)). Pass `application_role` to grant your application's role the minimal permissions it needs to use the datasource. Then, create your datasources with `run_migrations=False`, so they only verify their tables are migrated: ```python # In your migration script, run with a privileged role: SQLAlchemyDatasource.migrate(os.environ["ADMIN_DATABASE_URL"], application_role="my_app_role") # In your application, run with my_app_role: ds = SQLAlchemyDatasource.create(os.environ["APP_DATABASE_URL"], run_migrations=False) ``` You can similarly migrate the DBOS system database separately with [`DBOS.migrate`](../reference/dbos-class.md#migrate) or the [`dbos migrate`](../reference/cli.md#dbos-migrate) command. #### Using a Datasource Inside a datasource transaction, access the current SQLAlchemy session with `ds.sql_session()` (or `ads.sql_session()` for async). ##### With the `@ds.transaction` Decorator Decorate any function with `@ds.transaction` to run it as a tracked database transaction: ```python @ds.transaction() def insert_greeting(name: str, note: str) -> None: session = ds.sql_session() # sqlalchemy.orm.Session session.execute( text("INSERT INTO greetings (name, note) VALUES (:name, :note)"), {"name": name, "note": note} ) @DBOS.workflow() def greeting_workflow(name: str, note: str) -> None: insert_greeting(name, note) ``` For async code: ```python @ads.transaction() async def insert_greeting(name: str, note: str) -> None: session = ads.sql_session() # sqlalchemy.ext.asyncio.AsyncSession await session.execute( text("INSERT INTO greetings (name, note) VALUES (:name, :note)"), {"name": name, "note": note} ) @DBOS.workflow() async def greeting_workflow(name: str, note: str) -> None: await insert_greeting(name, note) ``` Putting it together, a complete async program looks like this. The datasource is created at module scope so that `@ads.transaction()` can decorate `insert_greeting` and `insert_greeting` can call `ads.sql_session()`:
Async Datasource Example ```python import asyncio import os from dbos import DBOS, DBOSConfig, AsyncSQLAlchemyDatasource from sqlalchemy import text ads = asyncio.run(AsyncSQLAlchemyDatasource.create(os.environ["APP_DATABASE_URL"])) config: DBOSConfig = { "name": "greeting-app", "system_database_url": os.environ["DBOS_SYSTEM_DATABASE_URL"], } DBOS(config=config) @ads.transaction() async def insert_greeting(name: str, note: str) -> None: session = ads.sql_session() await session.execute( text("INSERT INTO greetings (name, note) VALUES (:name, :note)"), {"name": name, "note": note}, ) @DBOS.workflow() async def greeting_workflow(name: str, note: str) -> None: await insert_greeting(name, note) async def main() -> None: await greeting_workflow("Alice", "Hello!") if __name__ == "__main__": DBOS.launch() asyncio.run(main()) ```
The decorator accepts two optional keyword arguments: - `name` – a custom step name recorded in the workflow log (defaults to the function's qualified name) - `isolation_level` – the SQL transaction isolation level; one of `"SERIALIZABLE"` (default), `"REPEATABLE READ"`, or `"READ COMMITTED"` (SQLite supports only `"SERIALIZABLE"`) ```python @ds.transaction(isolation_level="READ COMMITTED", name="insert_greeting") def insert_greeting(name: str, note: str) -> None: session = ds.sql_session() session.execute(...) ``` ##### Inline with `run_tx_step` / `run_tx_step_async` You can also run an un-decorated function as a datasource transaction step inline: ```python def insert_greeting(name: str, note: str) -> None: session = ds.sql_session() # sqlalchemy.orm.Session session.execute( text("INSERT INTO greetings (name, note) VALUES (:name, :note)"), {"name": name, "note": note} ) @DBOS.workflow() def greeting_workflow(name: str, note: str) -> None: ds.run_tx_step({"name": "insert_greeting"}, insert_greeting, name, note) ``` For async code: ```python async def insert_greeting(name: str, note: str) -> None: session = ads.sql_session() # sqlalchemy.ext.asyncio.AsyncSession await session.execute(...) @DBOS.workflow() async def greeting_workflow(name: str, note: str) -> None: await ads.run_tx_step_async({"name": "insert_greeting"}, insert_greeting, name, note) ``` The first argument to `run_tx_step` / `run_tx_step_async` is a dict with optional keys `name` and `isolation_level`, or `None` to use the defaults. #### How Datasource Transactions Work When a datasource transaction runs inside a DBOS workflow, DBOS records the outcome atomically in the same database transaction. If the workflow is interrupted and replayed, DBOS detects the existing record and returns the stored result without re-executing the function—exactly-once semantics even for side effects on your application database. Once the workflow completes, its outcome is recorded in the system database, so DBOS deletes its records from the `datasource_outputs` table. Outside a workflow, datasource transactions execute normally as plain SQLAlchemy transactions with no recording overhead. --- ## Upgrading Workflow Code(3) One challenge you may encounter when operating long-running durable workflows in production is **how to deploy breaking changes without disrupting in-progress workflows.** A breaking change to a workflow is any change in what steps run or the order in which steps run. The issue is that if a breaking change was made to a workflow, the checkpoints created by a workflow that started on the previous version of the code may not match the steps called by the workflow in the new version of the code, which makes the workflow difficult to recover. DBOS supports two strategies for safely upgrading workflow code: **patching** and **versioning**. ### Patching When using patching, you use [`DBOS.patch()`](../reference/contexts.md#patch) to make a breaking change in a conditional. `DBOS.patch()` returns `True` for new workflows (those started after the breaking change) and `False` for old workflows (those started before the breaking change). Therefore, if `DBOS.patch()` is `True`, call the new code, else, call the old code. To use patching, you must enable it in configuration: ```python config: DBOSConfig = { "name": "dbos-app", "system_database_url": os.environ.get("DBOS_SYSTEM_DATABASE_URL"), "enable_patching": True, } DBOS(config=config) ``` For example, let's say our workflow is: ```python @DBOS.workflow() def workflow(): foo() bar() ``` We want to replace the call to `foo()` with a call to `baz()`. This is a breaking change because it changes what steps run. We can make this breaking change safely using a patch: ```python @DBOS.workflow() def workflow(): if DBOS.patch("use-baz"): baz() else: foo() bar() ``` Now, new workflows will run `baz()`, while old workflows will safely continue through `foo()`. In [coroutine workflows](./workflow-tutorial.md#coroutine-async-workflows), use `await DBOS.patch_async()` and `await DBOS.deprecate_patch_async()` instead, as `DBOS.patch()` and `DBOS.deprecate_patch()` raise an error when called from a running event loop. #### Deprecating and Removing Patches Patches don't need to stay in your code forever. Once all workflows that started before you deployed the patch are complete, you can safely remove patches from your code. You can use the [list workflows APIs](./workflow-management.md#listing-workflows) to see what workflows are still active. First, you must deprecate the patch with [`DBOS.deprecate_patch()`](../reference/contexts.md#deprecate_patch). This safely runs all workflows that contain the patch marker, but does not insert the patch marker into new workflows. For example, here's how to deprecate the patch above: ```python @DBOS.workflow() def workflow(): DBOS.deprecate_patch("use-baz") baz() bar() ``` Then, when all workflows that started before you deprecated the patch are complete, you can remove the patch entirely: ```python @DBOS.workflow() def workflow(): baz() bar() ``` If any mistakes happen during the process (a breaking change is not patched, or a patch is deprecated or removed prematurely), the workflow will throw a `DBOSUnexpectedStepError` error clearly pointing to the step where the problem occurred. #### How Patching Works Under the hood, when you call `DBOS.patch()` from a workflow, it attempts to insert a "patch marker" at its current point in your workflow history (this is a new row in the `operation_outputs` table in your database). If it successfully inserts the patch marker or if the patch marker is already present, then the workflow must be new (it started after the patch, or started before the patch but has not yet reached this point), so `DBOS.patch()` returns `True`. If there is already a record present in this point in your workflow history, then the workflow must be old (it started before the patch and has already continued past this point), so `DBOS.patch()` returns `False`. When you deprecate a patch with `DBOS.deprecate_patch()`, new workflows no longer insert patch markers into their workflow history. However, if a workflow contains the patch marker in its history, it continues past that patch marker, safely ignoring it. Once all workflows with patch markers are complete, you can safely remove the patch entirely. ### Versioning When using versioning, DBOS **versions** applications and workflows. All workflows are tagged with the application version on which they started. By default, application version is automatically computed from a hash of workflow source code. However, you can set your own version through configuration. ```python config: DBOSConfig = { "name": "dbos-app", "system_database_url": os.environ.get("DBOS_SYSTEM_DATABASE_URL"), "application_version": "1.0.0", } DBOS(config=config) ``` When DBOS tries to recover workflows, it only recovers workflows whose version matches the current application version. This prevents unsafe recovery of workflows that depend on different code. When using versioning, we recommend **blue-green** code upgrades. When deploying a new version of your code, launch new processes running your new code version, but retain some processes running your old code version. Direct new traffic to your new processes while your old processes "drain" and complete all workflows of the old code version. To direct enqueued workflows to processes running the latest version of your code, use [`DBOS.get_latest_application_version`](../reference/contexts.md#get_latest_application_version) to look up the latest version: ```python from dbos import DBOS, SetEnqueueOptions DBOS.register_queue("my_queue") latest_version = DBOS.get_latest_application_version() with SetEnqueueOptions(app_version=latest_version["version_name"]): DBOS.enqueue_workflow("my_queue", my_workflow, arg1, arg2) ``` Or using [`DBOSClient`](../reference/client.md#version-management): ```python from dbos import DBOSClient, EnqueueOptions client = DBOSClient( system_database_url=os.environ["DBOS_SYSTEM_DATABASE_URL"], application_name="my-app", ) latest_version = client.get_latest_application_version() options: EnqueueOptions = { "workflow_name": "my_workflow", "queue_name": "my_queue", "app_version": latest_version["version_name"], } handle = client.enqueue(options, arg1, arg2) ``` Note that [scheduled workflows](./scheduled-workflows.md) are automatically enqueued to their owning application's latest version. Then, once all workflows of the old version are complete, you can retire the old code version. You can use [`DBOS.list_workflows`](../reference/contexts.md#list_workflows) to check if any workflows are still active for a given version: ```python active = DBOS.list_workflows( app_version="1.0.0", status=["ENQUEUED", "DELAYED", "PENDING"], ) if not active: print("Safe to retire version 1.0.0") ``` If you need to roll back to a previous version of your code, use [`DBOS.set_latest_application_version`](../reference/contexts.md#set_latest_application_version) to promote that version to latest so new enqueued workflows are directed to it: ```python DBOS.set_latest_application_version("1.0.0") ``` --- ## Communicating with Workflows(3) DBOS provides a few different ways to communicate with your workflows. You can: - [Send messages to workflows](#workflow-messaging-and-notifications) - [Publish events from workflows for clients to read](#workflow-events) - [Stream values from workflows to clients](#workflow-streaming) ### Workflow Messaging and Notifications You can send messages to a specific workflow. This is useful for signaling a workflow or sending notifications to it while it's running. ##### Send ```python DBOS.send( destination_id: str, message: Any, topic: Optional[str] = None ) -> None ``` You can call `DBOS.send()` to send a message to a workflow. Messages can optionally be associated with a topic and are queued on the receiver per topic. You can also call [`send`](../reference/client.md#send) from outside of your DBOS application with the [DBOS Client](../reference/client.md) or with the [`dbos.send_message` PL/pgSQL function](../../explanations/system-tables.md#dbossend_message). ##### Recv ```python DBOS.recv( topic: Optional[str] = None, timeout_seconds: float = 60, ) -> Any ``` Workflows can call `DBOS.recv()` to receive messages sent to them, optionally for a particular topic. Each call to `recv()` waits for and consumes the next message to arrive in the queue for the specified topic, returning `None` if the wait times out. If the topic is not specified, this method only receives messages sent without a topic. ##### Messages Example Messages are especially useful for sending notifications to a workflow. For example, in the [widget store demo](../examples/widget-store.md), the checkout workflow, after redirecting customers to a payments page, must wait for a notification that the user has paid. To wait for this notification, the payments workflow uses `recv()`, executing failure-handling code if the notification doesn't arrive in time: ```python @DBOS.workflow() def checkout_workflow(): ... # Validate the order, then redirect customers to a payments service. payment_status = DBOS.recv(PAYMENT_STATUS) if payment_status is not None and payment_status == "paid": ... # Handle a successful payment. else: ... # Handle a failed payment or timeout. ``` An endpoint waits for the payment processor to send the notification, then uses `send()` to forward it to the workflow: ```python @app.post("/payment_webhook/{payment_id}/{payment_status}") def payment_endpoint(payment_id: str, payment_status: str) -> Response: # Send the payment status to the checkout workflow. DBOS.send(payment_id, payment_status, PAYMENT_STATUS) ``` ##### Reliability Guarantees All messages are persisted to the database, so if `send` completes successfully, the destination workflow is guaranteed to be able to `recv` it. If you're sending a message from a workflow, DBOS guarantees exactly-once delivery. If you're sending a message from normal Python code, you can pass an `idempotency_key` to [`DBOS.send`](../reference/contexts.md#send) to guarantee exactly-once delivery. ### Workflow Events Workflows can publish _events_, which are key-value pairs associated with the workflow. They are useful for publishing information about the status of a workflow or to send a result to clients while the workflow is running. ##### set_event ```python DBOS.set_event( key: str, value: Any, ) -> None ``` Any workflow or step can call [`DBOS.set_event`](../reference/contexts.md#set_event) to publish a key-value pair, or update its value if it has already been published. ##### get_event ```python DBOS.get_event( workflow_id: str, key: str, timeout_seconds: float = 60, ) -> Any ``` You can call [`DBOS.get_event`](../reference/contexts.md#get_event) to retrieve the value published by a particular workflow identity for a particular key. If the event does not yet exist, this call waits for it to be published, returning `None` if the wait times out. You can also call [`get_event`](../reference/client.md#get_event) from outside of your DBOS application with [DBOS Client](../reference/client.md). ##### get_all_events ```python DBOS.get_all_events( workflow_id: str ) -> Dict[str, Any] ``` You can use `DBOS.get_all_events` to retrieve the latest values of all events published by a workflow. ##### Events Example Events are especially useful for writing interactive workflows that communicate information to their caller. For example, in the [widget store demo](../examples/widget-store.md), the checkout workflow, after validating an order, needs to send the customer a unique payment ID. To communicate the payment ID to the customer, it uses events. The payments workflow emits the payment ID using `set_event()`: ```python @DBOS.workflow() def checkout_workflow(): ... payment_id = ... DBOS.set_event(PAYMENT_ID, payment_id) ... ``` The FastAPI handler that originally started the workflow uses `get_event()` to await this payment ID, then returns it: ```python @app.post("/checkout/{idempotency_key}") def checkout_endpoint(idempotency_key: str) -> Response: # Idempotently start the checkout workflow in the background. with SetWorkflowID(idempotency_key): handle = DBOS.start_workflow(checkout_workflow) # Wait for the checkout workflow to send a payment ID, then return it. payment_id = DBOS.get_event(handle.workflow_id, PAYMENT_ID) if payment_id is None: raise HTTPException(status_code=404, detail="Checkout failed to start") return Response(payment_id) ``` ##### Reliability Guarantees All events are persisted to the database, so the latest version of an event is always retrievable. Additionally, if `get_event` is called in a workflow, the retrieved value is persisted in the database so workflow recovery can use that value, even if the event is later updated. ### Workflow Streaming Workflows can stream data in real time to clients. This is useful for streaming results from a long-running workflow or LLM call or for monitoring or progress reporting. ##### Writing to Streams ```python DBOS.write_stream( key: str, value: Any ) -> None: ``` You can write values to a stream from a workflow or its steps using [`DBOS.write_stream`](../reference/contexts.md#write_stream). A workflow may have any number of streams, each identified by a unique key. When you are done writing to a stream, you should close it with [`DBOS.close_stream`](../reference/contexts.md#close_stream). Otherwise, streams are automatically closed when the workflow terminates. ```python DBOS.close_stream( key: str ) -> None ``` DBOS streams are immutable and append-only. Writes to a stream from a workflow happen exactly-once. Writes to a stream from a step happen at-least-once; if a step fails and is retried, it may write to the stream multiple times. Readers will see all values written to the stream from all tries of the step in the order in which they were written. **Example syntax:** ```python @DBOS.workflow() def producer_workflow(): DBOS.write_stream(example_key, {"step": 1, "data": "value1"}) DBOS.write_stream(example_key, {"step": 2, "data": "value2"}) DBOS.close_stream(example_key) # Signal completion ``` ##### Reading from Streams ```python DBOS.read_stream( workflow_id: str, key: str, *, offset: int = 0, polling_interval_sec: Optional[float] = None, timeout_seconds: Optional[float] = None, ) -> Generator[Any, Any, None] ``` You can read values from a stream from anywhere using [`DBOS.read_stream`](../reference/contexts.md#read_stream). This function reads values from a stream identified by a workflow ID and key, yielding each value in order until the stream is closed or the workflow terminates. You can also read from a stream from outside a DBOS application with a [DBOS Client](../reference/client.md#read_stream). **Example syntax:** ```python for value in DBOS.read_stream(workflow_id, example_key): print(f"Received: {value}") ``` ##### Stream Timeouts By default, a read waits indefinitely for the next value. Pass `timeout_seconds` to bound that wait; if no value arrives in time, the read raises `DBOSStreamTimeoutError`. The timeout applies to **each value**, not to the read as a whole; the clock restarts every time a value is delivered. ```python from dbos import DBOS from dbos import error as dboserror try: for value in DBOS.read_stream(workflow_id, example_key, timeout_seconds=30): print(f"Received: {value}") except dboserror.DBOSStreamTimeoutError: print("The producer stopped sending values") ``` --- ## Workflow Management(4) You can view and manage your durable workflow executions via the [DBOS Console](../../conductor/workflow-management.md), programmatically, or via command line. ### Listing Workflows You can list your application's workflows programmatically via [`DBOS.list_workflows`](../reference/contexts.md#list_workflows) or from the command line with [`dbos workflow list`](../reference/cli.md#dbos-workflow-list). You can also view a searchable and expandable list of your application's workflows from its page on the [DBOS Console](../../conductor/workflow-management.md). ### Listing Workflow Steps You can list the steps of a workflow programmatically via [`DBOS.list_workflow_steps`](../reference/contexts.md#list_workflow_steps) or from the command line with [`dbos workflow steps`](../reference/cli.md#dbos-workflow-steps). You can also visualize a workflow's execution as a trace timeline (showing the workflow, its steps, and its child workflows and their steps) from its page on the [DBOS Console](../../conductor/workflow-management.md). For example, here is the trace of a workflow that processes multiple tasks concurrently by enqueueing child workflows: ### Workflow Attributes You can attach a dictionary of custom, JSON-serializable key-value **attributes** to your workflows using [`SetWorkflowAttributes`](../reference/contexts.md#setworkflowattributes). This is useful for tagging workflows with application-specific metadata such as a customer ID, tenant, or region. ```python from dbos import DBOS, SetWorkflowAttributes with SetWorkflowAttributes({"customer": "acme", "region": "us-east-1"}): process_order(order) ``` Attributes are recorded at creation time and are not inherited by child workflows. They are stored in Postgres as GIN-indexed JSONB, so you can efficiently search for workflows by attribute by passing the `attributes` filter to [`DBOS.list_workflows`](../reference/contexts.md#list_workflows). A workflow matches if its attributes contain all the key-value pairs you provide: ```python # Retrieve all workflows tagged with this customer workflows = DBOS.list_workflows(attributes={"customer": "acme"}) ``` To change a workflow's attributes after it is created, use [`DBOS.update_workflow_attributes`](../reference/contexts.md#update_workflow_attributes), which replaces the workflow's entire attributes dictionary (pass `None` to clear them). :::note Filtering workflows by attribute is only supported when using a Postgres system database. ::: ### Cancelling Workflows You can cancel the execution of a workflow from the web UI, programmatically via [`DBOS.cancel_workflow`](../reference/contexts.md#cancel_workflow), or through the command line with [`dbos workflow cancel`](../reference/cli.md#dbos-workflow-cancel). If the workflow is currently executing, cancelling it preempts its execution (interrupting it at the beginning of its next step). If the workflow is enqueued, cancelling removes it from the queue. To cancel an executing async step immediately rather than waiting for it to complete, mark the step as [`preemptible`](../reference/decorators.md#step). ### Resuming Workflows You can resume a workflow from its last completed step from the web UI, programmatically via [`DBOS.resume_workflow`](../reference/contexts.md#resume_workflow), or through the command line with [`dbos workflow resume`](../reference/cli.md#dbos-workflow-resume). You can use this to resume workflows that are cancelled or that have exceeded their maximum recovery attempts. You can also use this to start an enqueued workflow immediately, bypassing its queue. ### Forking Workflows You can start a new execution of a workflow by **forking** it from a specific step. When you fork a workflow, DBOS generates a new workflow with a new workflow ID, copies to that workflow the original workflow's inputs and all its steps up to the selected step, then begins executing the new workflow from the selected step. Forking a workflow is useful for recovering from outages in downstream services (by forking from the step that failed after the outage is resolved) or for "patching" workflows that failed due to a bug in a previous application version (by forking from the bugged step to an application version on which the bug is fixed). You can fork a workflow programmatically using [`DBOS.fork_workflow`](../reference/contexts.md#fork_workflow). You can also fork a workflow from a step from the web UI by clicking on that step in the workflow's trace timeline: ### Rewinding Workflows You can re-execute a workflow from a specific step, keeping its workflow ID, by **rewinding** it. When you rewind a workflow, DBOS discards the workflow's recorded steps from the selected step onward, clears its output, and re-enqueues it. The workflow then re-executes from the selected step, replaying the recorded outputs of earlier steps. The difference between rewind and fork is that fork creates a copy of the workflow with a new ID, while rewind actually "rewinds" the original workflow (modifying its state) and keeps its original ID. Because the rewound workflow keeps its original ID, other workflows and clients can keep sending messages to it and reading its events and streams, and its child workflows keep the same IDs. Rewinding is useful when other code refers to a workflow by its ID, for example when the ID is an idempotency key derived from an order or request ID. You can only rewind a workflow that is in a terminal state (for example, `SUCCESS`, `ERROR`, or `CANCELLED`); cancel a running workflow before rewinding it. Like forking, you can rewind a workflow onto a new application version to "patch" a workflow that failed due to a bug. You can rewind a workflow programmatically using [`DBOS.rewind_workflow`](../reference/contexts.md#rewind_workflow): ```python # Re-execute the workflow from step 3, keeping its workflow ID handle = DBOS.rewind_workflow(workflow_id, start_step=3) result = handle.get_result() ``` --- ## Workflows(3) Workflows provide **durable execution** so you can write programs that are **resilient to any failure**. Workflows help you write fault-tolerant background tasks, data processing pipelines, AI agents, and more. You can make a function a workflow by annotating it with [`@DBOS.workflow()`](../reference/decorators.md#workflow). Workflows call [steps](./step-tutorial.md), which are Python functions annotated with [`@DBOS.step()`](../reference/decorators.md#step). If a workflow is interrupted for any reason, DBOS automatically recovers its execution from the last completed step. Here's an example of a workflow: ```python @DBOS.step() def step_one(): print("Step one completed!") @DBOS.step() def step_two(): print("Step two completed!") @DBOS.workflow() def workflow(): step_one() step_two() ``` ### Starting Workflows In The Background One common use-case for workflows is building reliable background tasks that keep running even when the program is interrupted, restarted, or crashes. You can use [`DBOS.start_workflow`](../reference/contexts.md#start_workflow) to start a workflow in the background. If you start a workflow this way, it returns a [workflow handle](../reference/workflow_handles.md), from which you can access information about the workflow or wait for it to complete and retrieve its result. Here's an example: ```python @DBOS.workflow() def background_task(input): # ... return output # Start the background task handle: WorkflowHandle = DBOS.start_workflow(background_task, input) # Wait for the background task to complete and retrieve its result. output = handle.get_result() ``` After starting a workflow in the background, you can use [`DBOS.retrieve_workflow`](../reference/contexts.md#retrieve_workflow) to retrieve a workflow's handle from its ID. You can also retrieve a workflow's handle from outside of your DBOS application with [`DBOSClient.retrieve_workflow`](../reference/client.md#retrieve_workflow). If you need to run many workflows in the background and manage their concurrency or flow control, you can also use [DBOS queues](./queue-tutorial.md). ### Workflow IDs and Idempotency Every time you execute a workflow, that execution is assigned a unique ID, by default a [UUID](https://en.wikipedia.org/wiki/Universally_unique_identifier). You can access this ID through the [`DBOS.workflow_id`](../reference/contexts.md#workflow_id) context variable. Workflow IDs are useful for communicating with workflows and developing interactive workflows. You can set the workflow ID of a workflow with [`SetWorkflowID`](../reference/contexts.md#setworkflowid). Workflow IDs must be **globally unique** for your application. An assigned workflow ID acts as an idempotency key: if a workflow is called multiple times with the same ID, it executes only once. This is useful if your operations have side effects like making a payment or sending an email. For example: ```python @DBOS.workflow() def example_workflow(): DBOS.logger.info(f"I am a workflow with ID {DBOS.workflow_id}") with SetWorkflowID("very-unique-id"): example_workflow() ``` By default, if you start a workflow with an ID that is already in use, DBOS returns a handle to (or, for a direct call, the result of) the existing workflow instead of starting a new one. To instead raise an error when a workflow ID is already in use, set `workflow_id_reuse_policy="reject"` in [`SetWorkflowID`](../reference/contexts.md#setworkflowid): ```python from dbos import error as dboserror try: with SetWorkflowID("very-unique-id", workflow_id_reuse_policy="reject"): example_workflow() except dboserror.DBOSWorkflowIDInUseError: # A workflow with this ID already exists ... ``` ### Determinism Workflows are in most respects normal Python functions. They can have loops, branches, conditionals, and so on. However, a workflow function must be **deterministic**: if called multiple times with the same inputs, it should invoke the same steps with the same inputs in the same order (given the same return values from those steps). If you need to perform a non-deterministic operation like accessing the database, calling a third-party API, generating a random number, or getting the local time, you shouldn't do it directly in a workflow function. Instead, you should do non-deterministic operations in [steps](./step-tutorial.md). For example, **don't do this**: ```python @DBOS.workflow() def example_workflow(): choice = random.randint(0, 1) if choice == 0: step_one() else: step_two() ``` Do this instead: ```python @DBOS.step() def generate_choice(): return random.randint(0, 1) @DBOS.workflow() def example_workflow(friend: str): choice = generate_choice() if choice == 0: step_one() else: step_two() ``` If DBOS detects that a single execution of a workflow recorded different results for the same step, it raises a `DBOSStepNondeterminismError`, indicating the workflow is not deterministic. ### Workflow Timeouts You can set a timeout for a workflow with [`SetWorkflowTimeout`](../reference/contexts.md#setworkflowtimeout). When the timeout expires, the workflow **and all its children** are cancelled. Cancelling a workflow sets its status to `CANCELLED` and preempts its execution at the beginning of its next step. To cancel an executing async step immediately rather than waiting for it to complete, mark the step as [`preemptible`](../reference/decorators.md#step). Timeouts are **start-to-completion**: if a workflow is enqueued, the timeout does not begin until the workflow is dequeued and starts execution. Also, timeouts are **durable**: they are stored in the database and persist across restarts, so workflows can have very long timeouts. Example syntax: ```python @DBOS.workflow() def example_workflow(): ... # If the workflow does not complete within 10 seconds, it times out and is cancelled with SetWorkflowTimeout(10): example_workflow() ``` ### Durable Sleep You can use [`DBOS.sleep()`](../reference/contexts.md#sleep) to put your workflow to sleep for any period of time. This sleep is **durable**—DBOS saves the wakeup time in the database so that even if the workflow is interrupted and restarted multiple times while sleeping, it still wakes up on schedule. Sleeping is useful for scheduling a workflow to run in the future (even days, weeks, or months from now). For example: ```python @DBOS.workflow() def schedule_task(time_to_sleep, task): # Durably sleep for some time before running the task DBOS.sleep(time_to_sleep) run_task(task) ``` ### Debouncing Workflows You can debounce workflows to delay their execution until some time has passed since the workflow has last been called. This is useful for preventing wasted work when a workflow may be triggered multiple times in quick succession. For example, if a user is editing an input field, you can debounce their changes to execute a processing workflow only after they haven't edited the field for some time: ```python @DBOS.workflow() def process_input(user_input): ... # Each time a user submits a new input, debounce the process_input workflow. # The workflow will wait until 60 seconds after the user stops submitting new inputs, # then process the last input submitted. debouncer = Debouncer.create(process_input) def on_user_input_submit(user_id, user_input): debounce_key = user_id debounce_period_sec = 60 debouncer.debounce(debounce_key, debounce_period_sec, user_input) ``` See the [debouncing reference](../reference/contexts.md#debouncing) for more details. ### Coroutine (Async) Workflows Coroutines (functions defined with `async def`, also known as async functions) can also be DBOS workflows. Coroutine workflows may invoke [coroutine steps](./step-tutorial.md#coroutine-steps) via [await expressions](https://docs.python.org/3/reference/expressions.html#await). You should start coroutine workflows using [`DBOS.start_workflow_async`](../reference/contexts.md#start_workflow_async) and enqueue them using [`DBOS.enqueue_workflow_async`](../reference/contexts.md#enqueue_workflow_async). Calling a coroutine workflow or starting it with `DBOS.start_workflow_async` always runs it in the same event loop as its caller, but a workflow enqueued with `DBOS.enqueue_workflow_async` is started by DBOS in the event loop in which `DBOS.launch()` was called (if that loop is still running) or otherwise in a separate background event loop. Additionally, coroutine workflows should use the asynchronous versions of the workflow [communication](./workflow-communication.md) context methods. ```python @DBOS.step() async def example_step(): async with aiohttp.ClientSession() as session: async with session.get("https://example.com") as response: return await response.text() @DBOS.workflow() async def example_workflow(friend: str): await DBOS.sleep_async(10) body = await example_step() return body ``` #### Running Async Steps In Parallel Initiating several concurrent steps in an `async` workflow, followed by awaiting them with `asyncio.gather(..., return_exceptions=True)`, is valid as long as the steps are started in a **deterministic order**. For example, the following is allowed: ```python # Start steps in a deterministic order (step1, step2, step3, step4), # then await them all together. # Collects exceptions instead of raising immediately results = await asyncio.gather( step1("arg1"), step2("arg2"), step3("arg3"), step4("arg4"), return_exceptions=True, ) return results ``` This is allowed because each step is started in a well-defined sequence before awaiting. By contrast, the following is not allowed: ```python async def seq_a(): await step1("arg1") await step2("arg3") async def seq_b(): await step3("arg2") await step4("arg4") results = await asyncio.gather(seq_a(), seq_b(), return_exceptions=True) return results ``` Here, `step2` and `step4` may be started in either order since their execution depends on the relative time taken by `step1` and `step3`. If you need to run sequences of operations concurrently, start child workflows and await their results, rather than interleaving step execution inside a single workflow. For proper error handling, when using `asyncio.gather()`, specify `return_exceptions=True`. Without `return_exceptions=True`, `gather` will raise any exception immediately and stop awaiting the rest of the tasks. If one of the remaining tasks later fails, its exception may go unobserved. Instead, prefer `asyncio.gather(..., return_exceptions=True)`, which safely waits for all tasks to complete and reports their outcomes. You can also use [`DBOS.asyncio_wait`](../reference/contexts.md#asyncio_wait), a durable wrapper around [`asyncio.wait`](https://docs.python.org/3/library/asyncio-task.html#asyncio.wait), to process tasks as they complete: ```python pending = [ step1("arg1"), step2("arg2"), step3("arg3"), step4("arg4"), ] # Process each result as it completes while pending: done, pending = await DBOS.asyncio_wait( pending, return_when=asyncio.FIRST_COMPLETED ) for task in done: result = task.result() DBOS.logger.info(f"Completed with result: {result}") ``` ### Workflow Guarantees Workflows provide the following guarantees. These guarantees assume that the application and database may crash and go offline at any point in time, but are always restarted and return online. 1. Workflows always run to completion. If a DBOS process is interrupted while executing a workflow and restarts, it resumes the workflow from the last completed step. 2. [Steps](./step-tutorial.md) are tried _at least once_ but are never re-executed after they complete. If a failure occurs inside a step, the step may be retried, but once a step has completed, it will never be re-executed. 3. [Transactions](./transaction-tutorial.md) commit _exactly once_. Once a workflow commits a transaction, it will never retry that transaction. If an exception is thrown from a workflow, the workflow terminates—DBOS records the exception, sets the workflow status to `ERROR`, and does not recover the workflow. This is because uncaught exceptions are assumed to be nonrecoverable. If your workflow performs operations that may transiently fail (for example, sending HTTP requests to unreliable services), those should be performed in [steps with configured retries](./step-tutorial.md#configurable-retries). DBOS provides [tooling](./workflow-management.md) to help you identify failed workflows and examine the specific uncaught exceptions. --- ## Upgrading to 3.0 DBOS Python 3.0 removes features that were deprecated in DBOS 2.x. This guide describes each removed feature and what to replace it with, as well as how to safely upgrade a running application. ### Upgrading a Running Application DBOS 3.0 changes the storage schema for workflow inputs and outputs to improve performance. Therefore, DBOS 3.0 can process workflows created by DBOS 2.x, but **DBOS 2.x cannot process workflows created by DBOS 3.0**. Therefore: - **Don't run DBOS 2.x and 3.0 processes concurrently with the same application version.** If you set `application_version` yourself, change it when you upgrade. If you use [patching](./tutorials/upgrading-workflows.md#patching), shut down all DBOS 2.x processes before launching DBOS 3.0 processes. - **Upgrade applications that use [`DBOSClient`](./reference/client.md) along with your DBOS processes.** A DBOS 2.x client cannot retrieve the inputs or results of workflows created by DBOS 3.0. - **Be careful reusing the IDs of workflows created by DBOS 2.x.** If DBOS 3.0 starts or enqueues a workflow with the same [ID](./tutorials/workflow-tutorial.md#workflow-ids-and-idempotency) as a workflow created by DBOS 2.x but with different inputs, the new inputs can replace that workflow's recorded inputs, and if it hasn't completed yet, it may run with them. ### Breaking Changes #### `@DBOS.transaction` and the Application Database The `@DBOS.transaction` decorator, `DBOS.sql_session`, and the `application_database_url` and `database_url` configuration fields have been removed. Instead, run database transactions with [datasources](./tutorials/transaction-tutorial.md#datasources). A datasource connects to your application database and runs transactions with the same exactly-once guarantees as `@DBOS.transaction`. **Before:** ```python config: DBOSConfig = { "name": "my-app", "system_database_url": os.environ["DBOS_SYSTEM_DATABASE_URL"], "application_database_url": os.environ["APP_DATABASE_URL"], } DBOS(config=config) @DBOS.transaction() def insert_greeting(name: str, note: str) -> None: sql = text("INSERT INTO greetings (name, note) VALUES (:name, :note)") DBOS.sql_session.execute(sql, {"name": name, "note": note}) ``` **After:** ```python from dbos import DBOS, DBOSConfig, SQLAlchemyDatasource config: DBOSConfig = { "name": "my-app", "system_database_url": os.environ["DBOS_SYSTEM_DATABASE_URL"], } DBOS(config=config) ds = SQLAlchemyDatasource.create(os.environ["APP_DATABASE_URL"]) @ds.transaction() def insert_greeting(name: str, note: str) -> None: sql = text("INSERT INTO greetings (name, note) VALUES (:name, :note)") ds.sql_session().execute(sql, {"name": name, "note": note}) ``` Datasources also support `async` transactions through [`AsyncSQLAlchemyDatasource`](./reference/datasources.md#asyncsqlalchemydatasource). For the full API, see the [datasource reference](./reference/datasources.md). :::warning If you previously set `database_url` or `application_database_url` but not `system_database_url`, DBOS stored its state in a separate system database. With Postgres, this database has the same name as your application database followed by `_dbos_sys`. With SQLite, it is the same database file. Set [`system_database_url`](./reference/configuration.md#database-connection-settings) to that database's connection string. ::: #### In-Memory Queues The `Queue(...)` constructor, which declared a queue in process memory, has been removed. Instead, register queues in the system database with [`DBOS.register_queue`](./reference/contexts.md#register_queue) after launching DBOS, then enqueue workflows by queue name with [`DBOS.enqueue_workflow`](./reference/contexts.md#enqueue_workflow). **Before:** ```python from dbos import DBOS, Queue queue = Queue("example_queue", worker_concurrency=5) @DBOS.workflow() def process_task(task): ... DBOS.launch() handle = queue.enqueue(process_task, task) ``` **After:** ```python from dbos import DBOS @DBOS.workflow() def process_task(task): ... DBOS.launch() DBOS.register_queue("example_queue", worker_concurrency=5) handle = DBOS.enqueue_workflow("example_queue", process_task, task) ``` When migrating your queues, note that: - In `async` code use `await DBOS.register_queue_async(...)` instead. - Register every queue your application previously declared in memory. Workflows enqueued on a queue that isn't registered stay `ENQUEUED` until the queue is registered. - [`DBOS.listen_queues`](./tutorials/queue-tutorial.md#explicit-queue-listening) now accepts only queue names, not `Queue` objects. - Queue names starting with `_dbos_` are reserved for DBOS. For more on queues, see the [queues tutorial](./tutorials/queue-tutorial.md). #### Legacy Partitioned Queues The `partition_queue` parameter of `DBOS.register_queue` and `DBOSClient.register_queue` has been removed. Instead, a queue is [partitioned](./tutorials/queue-tutorial.md#partitioning-queues) if you set any per-partition limit. With `partition_queue=True`, the `concurrency`, `worker_concurrency`, and `limiter` settings applied to each partition. Replace them with `partition_concurrency`, `partition_worker_concurrency`, and `partition_limiter`. **Before:** ```python DBOS.register_queue("partitioned_queue", partition_queue=True, concurrency=1) ``` **After:** ```python DBOS.register_queue("partitioned_queue", partition_concurrency=1) ``` Unlike legacy partitioned queues, a partitioned queue can also have queue-wide limits, as described in [Combining Queue-Wide and Per-Partition Limits](./tutorials/queue-tutorial.md#combining-queue-wide-and-per-partition-limits). The `priority_enabled` parameter and the `Queue` methods for reading and setting `priority_enabled` and `partition_queue` have also been removed. [Priority](./tutorials/queue-tutorial.md#priority) is now always enabled. #### Decorator-Based Scheduling The `@DBOS.scheduled` decorator has been removed. Instead, create schedules in the system database with [`DBOS.apply_schedules`](./reference/contexts.md#apply_schedules) or [`DBOS.create_schedule`](./reference/contexts.md#create_schedule). The second argument of a scheduled workflow is now the schedule's `context` instead of the time at which the workflow actually started. **Before:** ```python @DBOS.scheduled("*/5 * * * *") @DBOS.workflow() def my_periodic_task(scheduled_time: datetime, actual_time: datetime): ... ``` **After:** ```python @DBOS.workflow() def my_periodic_task(scheduled_time: datetime, context: Any): ... DBOS.launch() DBOS.apply_schedules([ { "schedule_name": "my-periodic-task", "workflow_fn": my_periodic_task, "schedule": "*/5 * * * *", }, ]) ``` Schedules persist in the system database, so a schedule you stop applying keeps running until you delete it with [`DBOS.delete_schedule`](./reference/contexts.md#delete_schedule). To learn more, see the [scheduling tutorial](./tutorials/scheduled-workflows.md). #### FastAPI and Flask Integrations The `fastapi` and `flask` parameters of the `DBOS` constructor have been removed. With FastAPI, launch and shut down DBOS from a [lifespan](https://fastapi.tiangolo.com/advanced/events/) function. **Before:** ```python app = FastAPI() DBOS(fastapi=app, config=config) ``` **After:** ```python from contextlib import asynccontextmanager @asynccontextmanager async def lifespan(app: FastAPI): DBOS.launch() try: yield finally: DBOS.destroy() app = FastAPI(lifespan=lifespan) DBOS(config=config) ``` With Flask, remove `flask=app` and call [`DBOS.launch()`](./reference/dbos-class.md#launch) before starting your app. The integrations created a tracing span for each HTTP request. Instead, use the OpenTelemetry instrumentation for your web framework. DBOS workflow spans automatically join your request spans, as described in the [tracing tutorial](./tutorials/logging-and-tracing.md#connecting-dbos-to-your-observability-provider). --- ## Get Started with DBOS DBOS is a library for building reliable programs. This guide shows you how to install and run it on your computer. :::tip To teach your AI coding assistant to build with DBOS, try out [skills](./python/prompting.md) and [MCP](./integrations/mcp.md). :::
##### 1. Create a Virtual Environment Create and activate a Python virtual environment in a directory. DBOS requires Python 3.10 or later.
**macOS or Linux** ```shell python3 -m venv dbos-app-starter/.venv cd dbos-app-starter source .venv/bin/activate ``` **Windows (PowerShell)** ```shell python3 -m venv dbos-app-starter/.venv cd dbos-app-starter .venv\Scripts\activate.ps1 ``` **Windows (cmd)** ```shell python3 -m venv dbos-app-starter/.venv cd dbos-app-starter .venv\Scripts\activate.bat ```
##### 2. Install and Initialize DBOS Install DBOS and FastAPI (used by the example application). Then initialize an example application.
```shell pip install dbos 'fastapi[standard]' dbos init --template dbos-app-starter ```
##### 3. Start Your App
Now, start your app!
```bash python3 main.py ```
To see that your app is working, visit this URL in your browser: http://localhost:8000/ This app lets you test the reliability of DBOS for yourself. Launch a durable workflow and watch it execute its three steps. At any point, crash the app. Then, restart it with `python3 main.py` and watch it seamlessly recover from where it left off. Congratulations, you've run your first durable workflow with DBOS!
##### 4. Connect to DBOS Conductor
[Conductor](./conductor/overview.md) is the control plane for your durable workflows, providing distributed workflow recovery, observability, and management. To connect your app to Conductor, first sign up for an account on the [DBOS Console](https://console.dbos.dev/login-redirect).
Then, install [`dbosctl`](./conductor/reference/dbosctl.md), the Conductor command-line client. On Windows, [download a release binary](https://github.com/dbos-inc/dbos-ctl/releases) instead.
```bash curl -sSfL https://raw.githubusercontent.com/dbos-inc/dbos-ctl/main/install.sh | sh ```
Next, configure a `dbosctl` profile and log in. `dbosctl login` prints a URL and a code for you to approve in your browser.
```bash dbosctl config set dbos --managed dbosctl login ```
Then, register your application with Conductor and create an API key. The name you register must match your app's name in its DBOS configuration. The key's secret is printed once and cannot be retrieved afterwards, so copy it now.
```bash dbosctl app register dbos-app-starter dbosctl api-key create dbos-app-starter-key ```
Finally, provide your API key to your app through the `DBOS_CONDUCTOR_KEY` environment variable, then restart it to connect it to Conductor.
```bash export DBOS_CONDUCTOR_KEY= python3 main.py ```
Your app is now connected to Conductor! You can view and manage its workflows from the [DBOS Console](https://console.dbos.dev).
Next: - Check out the [**DBOS programming guide**](./python/programming-guide.md) to learn how to build reliable applications with DBOS. - Learn how to [**add DBOS to your application**](./python/integrating-dbos.md) to make it reliable with just a few lines of code.
:::tip To teach your AI coding assistant to build with DBOS, try out [skills](./typescript/prompting.md) and [MCP](./integrations/mcp.md). :::
##### 1. Initialize an Application Initialize a starter application and enter its directory. DBOS requires Node v20 or later.
```shell npx @dbos-inc/create@latest --template dbos-node-starter cd dbos-node-starter ```
##### 2. Build Your Application Install dependencies, then build your application.
```shell npm install npm run build ```
##### 3. Start Your App
DBOS requires a Postgres database. If you already have Postgres, you can set the `DBOS_SYSTEM_DATABASE_URL` environment variable to your connection string. Otherwise, you can start Postgres in a Docker container with this command:
```bash npx dbos postgres start ```
Now, start your app!
```bash npm run start ```
To see that your app is working, visit this URL in your browser: http://localhost:3000/ This app lets you test the reliability of DBOS for yourself. Launch a durable workflow and watch it execute its three steps. At any point, crash the app. Then, restart it with `npm run start` and watch it seamlessly recover from where it left off. Congratulations, you've run your first durable workflow with DBOS!
##### 4. Connect to DBOS Conductor
[Conductor](./conductor/overview.md) is the control plane for your durable workflows, providing distributed workflow recovery, observability, and management. To connect your app to Conductor, first sign up for an account on the [DBOS Console](https://console.dbos.dev/login-redirect).
Then, install [`dbosctl`](./conductor/reference/dbosctl.md), the Conductor command-line client. On Windows, [download a release binary](https://github.com/dbos-inc/dbos-ctl/releases) instead.
```bash curl -sSfL https://raw.githubusercontent.com/dbos-inc/dbos-ctl/main/install.sh | sh ```
Next, configure a `dbosctl` profile and log in. `dbosctl login` prints a URL and a code for you to approve in your browser.
```bash dbosctl config set dbos --managed dbosctl login ```
Then, register your application with Conductor and create an API key. The name you register must match your app's name in its DBOS configuration. The key's secret is printed once and cannot be retrieved afterwards, so copy it now.
```bash dbosctl app register dbos-node-starter dbosctl api-key create dbos-node-starter-key ```
Finally, provide your API key to your app through the `DBOS_CONDUCTOR_KEY` environment variable, then restart it to connect it to Conductor.
```bash export DBOS_CONDUCTOR_KEY= npm run start ```
Your app is now connected to Conductor! You can view and manage its workflows from the [DBOS Console](https://console.dbos.dev).
Next: - Check out the [**DBOS programming guide**](./typescript/programming-guide.md) to learn how to build reliable applications with DBOS. - Learn how to [**add DBOS to your application**](./typescript/integrating-dbos.md) to make it reliable with just a few lines of code.
:::tip To teach your AI coding assistant to build with DBOS, try out [skills](./golang/prompting.md) and [MCP](./integrations/mcp.md). ::: ##### 1. Initialize an Application
Install the DBOS Go CLI, then initialize a starter application and enter its directory. DBOS requires Go 1.25.0 or higher.
```shell go install github.com/dbos-inc/dbos-transact-golang/cmd/dbos@latest dbos init cd dbos-go-starter ```
##### 2. Launch Postgres
DBOS requires a Postgres database. If you already have Postgres, you can set the `DBOS_SYSTEM_DATABASE_URL` environment variable to your connection string. Otherwise, you can start Postgres in a Docker container with this command:
```shell dbos postgres start export DBOS_SYSTEM_DATABASE_URL=postgres://postgres:dbos@localhost:5432/dbos_go_starter ```
##### 3. Start Your App
Now, download dependencies and start your app!
```bash go mod tidy go run main.go ```
To see that your app is working, visit this URL in your browser: http://localhost:8080/ This app lets you test the reliability of DBOS for yourself. Launch a durable workflow and watch it execute its three steps. At any point, crash the app. Then, restart it with `go run main.go` and watch it seamlessly recover from where it left off. Congratulations, you've run your first durable workflow with DBOS!
##### 4. Connect to DBOS Conductor
[Conductor](./conductor/overview.md) is the control plane for your durable workflows, providing distributed workflow recovery, observability, and management. To connect your app to Conductor, first sign up for an account on the [DBOS Console](https://console.dbos.dev/login-redirect).
Then, install [`dbosctl`](./conductor/reference/dbosctl.md), the Conductor command-line client. On Windows, [download a release binary](https://github.com/dbos-inc/dbos-ctl/releases) instead.
```bash curl -sSfL https://raw.githubusercontent.com/dbos-inc/dbos-ctl/main/install.sh | sh ```
Next, configure a `dbosctl` profile and log in. `dbosctl login` prints a URL and a code for you to approve in your browser.
```bash dbosctl config set dbos --managed dbosctl login ```
Then, register your application with Conductor and create an API key. The name you register must match your app's name in its DBOS configuration. The key's secret is printed once and cannot be retrieved afterwards, so copy it now.
```bash dbosctl app register dbos-go-starter dbosctl api-key create dbos-go-starter-key ```
Finally, provide your API key to your app through the `DBOS_CONDUCTOR_KEY` environment variable, then restart it to connect it to Conductor.
```bash export DBOS_CONDUCTOR_KEY= go run main.go ```
Your app is now connected to Conductor! You can view and manage its workflows from the [DBOS Console](https://console.dbos.dev).
Next: - Check out the [**DBOS programming guide**](./golang/programming-guide.md) to learn how to build reliable applications with DBOS. - Learn how to [**add DBOS to your application**](./golang/integrating-dbos.md) to make it reliable with just a few lines of code.
:::tip To teach your AI coding assistant to build with DBOS, try out [skills](./java/prompting.md) and [MCP](./integrations/mcp.md). ::: ##### 1. Download an Application
Download an example application and enter its directory. DBOS requires Java 17 or higher.
```shell git clone https://github.com/dbos-inc/dbos-demo-apps.git cd dbos-demo-apps/java/dbos-starter ```
##### 2. Launch Postgres
DBOS requires a Postgres database. If you already have Postgres, you can set the `DBOS_SYSTEM_JDBC_URL` environment variable to your connection string. Otherwise, you can start Postgres in a Docker container and set environment variables with these commands:
```shell docker run -d --name dbos-postgres -e POSTGRES_PASSWORD=dbos -p 5432:5432 postgres:17 export PGUSER=postgres export PGPASSWORD=dbos export DBOS_SYSTEM_JDBC_URL=jdbc:postgresql://localhost:5432/dbos_java_starter ```
##### 3. Start Your App
Now, start your app!
```bash ./gradlew run ```
To see that your app is working, visit this URL in your browser: http://localhost:7070/ This app lets you test the reliability of DBOS for yourself. Launch a durable workflow and watch it execute its three steps. At any point, crash the app. Then, restart it with `./gradlew run` and watch it seamlessly recover from where it left off. Congratulations, you've run your first durable workflow with DBOS!
##### 4. Connect to DBOS Conductor
[Conductor](./conductor/overview.md) is the control plane for your durable workflows, providing distributed workflow recovery, observability, and management. To connect your app to Conductor, first sign up for an account on the [DBOS Console](https://console.dbos.dev/login-redirect).
Then, install [`dbosctl`](./conductor/reference/dbosctl.md), the Conductor command-line client. On Windows, [download a release binary](https://github.com/dbos-inc/dbos-ctl/releases) instead.
```bash curl -sSfL https://raw.githubusercontent.com/dbos-inc/dbos-ctl/main/install.sh | sh ```
Next, configure a `dbosctl` profile and log in. `dbosctl login` prints a URL and a code for you to approve in your browser.
```bash dbosctl config set dbos --managed dbosctl login ```
Then, register your application with Conductor and create an API key. The name you register must match your app's name in its DBOS configuration. The key's secret is printed once and cannot be retrieved afterwards, so copy it now.
```bash dbosctl app register dbos-starter-java dbosctl api-key create dbos-starter-java-key ```
Finally, provide your API key to your app through the `DBOS_CONDUCTOR_KEY` environment variable, then restart it to connect it to Conductor.
```bash export DBOS_CONDUCTOR_KEY= ./gradlew run ```
Your app is now connected to Conductor! You can view and manage its workflows from the [DBOS Console](https://console.dbos.dev).
Next: - Check out the [**DBOS programming guide**](./java/programming-guide.md) to learn how to build reliable applications with DBOS. - Learn how to [**add DBOS to your application**](./java/integrating-dbos.md) to make it reliable with just a few lines of code.
--- ## Fault-Tolerant Checkout(4) :::info This example is also available in [Python](../../python/examples/widget-store), [Java](../../java/examples/widget-store), and [Go](../../golang/examples/widget-store.md). ::: In this example, we use DBOS and Fastify to deploy an online storefront that's resilient to any failure. You can see the application live [here](https://demo-widget-store.cloud.dbos.dev/). Try playing with it and pressing the crash button as often as you want. Within a few seconds, the app will recover and resume as if nothing happened. All source code is [available on GitHub](https://github.com/dbos-inc/dbos-demo-apps/tree/main/typescript/widget-store). ![Widget store UI](../../python/examples/assets/widget_store_ui.png) ### Building the Checkout Workflow The heart of this application is the checkout workflow, which orchestrates the entire purchase process. This workflow is triggered whenever a customer buys a widget and handles the complete order lifecycle: 1. Creates a new order in the system 2. Reserves inventory to ensure the item is available 3. Processes payment 4. Marks the order as paid and initiates fulfillment 5. Handles failures gracefully by releasing reserved inventory and canceling orders when necessary DBOS **durably executes** this workflow. It checkpoints each step in the database so that if the app fails or is interrupted during checkout, it will automatically recover from the last completed step. This means that customers never lose their order progress, no matter what breaks. You can try this yourself! On the [live application](https://demo-widget-store.cloud.dbos.dev/), start an order and press the crash button at any time. Within seconds, your app will recover to exactly the state it was in before the crash and continue as if nothing happened. ```javascript export const checkoutWorkflow = DBOS.registerWorkflow( async () => { // Attempt to reserve inventory, failing if no inventory remains try { await subtractInventory(); } catch (error) { console.error(`Failed to update inventory: ${(error as Error).message}`); await DBOS.setEvent(PAYMENT_ID_EVENT, null); return; } // Create a new order const orderID = await createOrder(); // Send a unique payment ID to the checkout endpoint so it can // redirect the customer to the payments page await DBOS.setEvent(PAYMENT_ID_EVENT, DBOS.workflowID); const notification = await DBOS.recv(PAYMENT_TOPIC, 120); // If payment succeeded, mark the order as paid and start the order dispatch workflow. // Otherwise, return reserved inventory and cancel the order. if (notification && notification === 'paid') { console.info(`Payment successful!`); await markOrderPaid(orderID); await DBOS.startWorkflow(dispatchOrder)(orderID); } else { console.warn(`Payment failed...`); await errorOrder(orderID); await undoSubtractInventory(); } // Finally, send the order ID to the payment endpoint so it can redirect // the customer to the order status page. await DBOS.setEvent(ORDER_ID_EVENT, orderID); }, { name: 'checkoutWorkflow' }, ); ``` ### The Checkout and Payment Endpoints Now let's implement the HTTP endpoints that handle customer interactions with the checkout system. The checkout endpoint is triggered when a customer clicks the "Buy Now" button. It starts the checkout workflow in the background, then waits for the workflow to generate and send it a unique payment ID. It then returns the payment ID so the browser can redirect the user to the payments page. The endpoint accepts an [idempotency key](../tutorials/workflow-tutorial.md#workflow-ids-and-idempotency) so that even if the customer presses "buy now" multiple times, only one checkout workflow is started. ```javascript const fastify = Fastify({logger: true}); fastify.post<{ Params: { key: string }; }>('/checkout/:key', async (req, reply) => { const key = req.params.key; // Idempotently start the checkout workflow in the background. const handle = await DBOS.startWorkflow(checkoutWorkflow, { workflowID: key })(); // Wait for the checkout workflow to send a payment ID, then return it. const paymentID = await DBOS.getEvent(handle.workflowID, PAYMENT_ID_EVENT, 300); if (paymentID === null) { DBOS.logger.error('checkout failed'); return reply.code(500).send('Error starting checkout'); } return paymentID; }); ``` The payment endpoint handles the communication between the payment system and the checkout workflow. It uses the payment ID to signal the checkout workflow whether the payment succeeded or failed. It then retrieves the order ID from the checkout workflow so the browser can redirect the customer to the order status page. ```javascript fastify.post<{ Params: { key: string; status: string }; }>('/payment_webhook/:key/:status', async (req, reply) => { const { key, status } = req.params; // Send the payment status to the checkout workflow. await DBOS.send(key, status, PAYMENT_TOPIC); // Wait for the checkout workflow to send an order ID, then return it. const orderID = await DBOS.getEvent(key, ORDER_ID_EVENT, 300); if (orderID === null) { DBOS.logger.error('retrieving order ID failed'); return reply.code(500).send('Error retrieving order ID'); } return orderID; }); ``` ### Database Operations Now, let's implement the checkout workflow's steps. Each step performs a database operation, like updating inventory or order status. Because these steps access the database, they are implemented using [datasource transactions](../tutorials/transaction-tutorial.md).
Database Operations ```javascript export const knexds = new KnexDataSource('app-db', config); export async function subtractInventory(): Promise { return knexds.runTransaction( async () => { const numAffected = await knexds.client('products') .where('product_id', PRODUCT_ID) .andWhere('inventory', '>=', 1) .update({ inventory: knexds.client.raw('inventory - ?', 1), }); if (numAffected <= 0) { throw new Error('Insufficient Inventory'); } }, { name: 'subtractInventory' }, ); } export async function undoSubtractInventory(): Promise { return knexds.runTransaction( async () => { await knexds.client('products') .where({ product_id: PRODUCT_ID }) .update({ inventory: knexds.client.raw('inventory + ?', 1) }); }, { name: 'undoSubtractInventory' }, ); } export async function setInventory(inventory: number): Promise { return knexds.runTransaction( async () => { await knexds.client('products').where({ product_id: PRODUCT_ID }).update({ inventory }); }, { name: 'setInventory' }, ); } export async function retrieveProduct(): Promise { return knexds.runTransaction( async () => { const item = await knexds.client('products').select('*').where({ product_id: PRODUCT_ID }); if (!item.length) { throw new Error(`Product ${PRODUCT_ID} not found`); } return item[0]; }, { name: 'retrieveProduct' }, ); } export async function createOrder(): Promise { return knexds.runTransaction( async () => { const orders = await knexds.client('orders') .insert({ order_status: OrderStatus.PENDING, product_id: PRODUCT_ID, last_update_time: knexds.client.fn.now(), progress_remaining: 10, }) .returning('order_id'); const orderID = orders[0].order_id; return orderID; }, { name: 'createOrder' }, ); } export async function markOrderPaid(order_id: number): Promise { return knexds.runTransaction( async () => { await knexds.client('orders').where({ order_id: order_id }).update({ order_status: OrderStatus.PAID, last_update_time: knexds.client.fn.now(), }); }, { name: 'markOrderPaid' }, ); } export async function errorOrder(order_id: number): Promise { return knexds.runTransaction( async () => { await knexds.client('orders').where({ order_id: order_id }).update({ order_status: OrderStatus.CANCELLED, last_update_time: knexds.client.fn.now(), }); }, { name: 'errorOrder' }, ); } export async function retrieveOrder(order_id: number): Promise { return knexds.runTransaction( async () => { const item = await knexds.client('orders').select('*').where({ order_id: order_id }); if (!item.length) { throw new Error(`Order ${order_id} not found`); } return item[0]; }, { name: 'retrieveOrder' }, ); } export async function retrieveOrders() { return knexds.runTransaction( async () => { return knexds.client('orders').select('*'); }, { name: 'retrieveOrders' }, ); } export const dispatchOrder = DBOS.registerWorkflow( async (order_id: number) => { for (let i = 0; i < 10; i++) { await DBOS.sleep(1000); await updateOrderProgress(order_id); } }, { name: 'dispatchOrder' }, ); export async function updateOrderProgress(order_id: number): Promise { return knexds.runTransaction( async () => { const orders = await knexds.client('orders').where({ order_id: order_id, order_status: OrderStatus.PAID, }); if (!orders.length) { throw new Error(`No PAID order with ID ${order_id} found`); } const order = orders[0]; if (order.progress_remaining > 1) { await knexds.client('orders') .where({ order_id: order_id }) .update({ progress_remaining: order.progress_remaining - 1 }); } else { await knexds.client('orders').where({ order_id: order_id }).update({ order_status: OrderStatus.DISPATCHED, progress_remaining: 0, }); } }, { name: 'updateOrderProgress' }, ); } ```
### Finishing Up Let's add the final touches to the app. This Fastify endpoint serves its frontend: ```javascript fastify.get('/', async (req, reply) => { async function render(file: string, ctx?: object): Promise { const engine = new Liquid({ root: path.resolve(__dirname, '..', 'public'), }); return (await engine.renderFile(file, ctx)) as string; } const html = await render('app.html', {}); return reply.type('text/html').send(html); }); ``` Here is the crash endpoint. It crashes your app. Trigger it as many times as you want—DBOS always comes back, resuming from exactly where it left off! ```javascript fastify.post('/crash_application', () => { process.exit(1); }); ``` Finally, let's start DBOS and the Fastify server: ```javascript async function main() { const PORT = parseInt(process.env.NODE_PORT || '3000'); DBOS.setConfig({ "name": 'widget-store-node', "applicationVersion": '0.1.0', "systemDatabaseUrl": process.env.DBOS_SYSTEM_DATABASE_URL, }); DBOS.logRegisteredEndpoints(); await DBOS.launch(); await fastify.listen({ port: PORT, host: '0.0.0.0' }); console.log(`🚀 Server is running on http://localhost:${PORT}`); } if (require.main === module) { main().catch(console.log); } ``` ### Try it Yourself! First, clone and enter the [dbos-demo-apps](https://github.com/dbos-inc/dbos-demo-apps) repository: ```shell git clone https://github.com/dbos-inc/dbos-demo-apps.git cd dbos-demo-apps/typescript/widget-store ``` Then install dependencies and build the application: ```shell npm install npm run build ``` Then, start Postgres in a local Docker container. If you already use Postgres, you can set the `DBOS_DATABASE_URL` (for application data) and `DBOS_SYSTEM_DATABASE_URL` (for DBOS system data) environment variables to your database connection string. ```shell npx dbos postgres start ``` Create database tables: ```shell npm run db:setup ``` Then start your app: ```shell npm run start ``` Visit [http://localhost:3000](http://localhost:3000) to see your app! --- ## Hacker News Research Agent(Examples) :::info This example is also available in [Python](../../python/examples/hacker-news-agent). ::: In this example, we use DBOS to build an AI deep research agent that autonomously searches Hacker News for information on any topic. This example demonstrates how to build **reliable, durable AI agents** with DBOS. The agent starts with a research topic, autonomously searches for related information, makes decisions about when to continue research, and synthesizes findings into a comprehensive report. Because the agent is implemented as a DBOS durable workflow, it can automatically recover from any failure and continue research from where it left off, ensuring no work is lost. This example also demonstrates how easy it is to add DBOS to an existing agentic application. Adding DBOS to this agent to make it reliable and observable required changing **<20 lines of code**. All you have to do is annotate workflows and steps. All source code is [available on GitHub](https://github.com/dbos-inc/dbos-demo-apps/tree/main/typescript/hacker-news-agent). ### Main Research Workflow The core of the agent is the main research workflow. It starts with a topic and autonomously explores related queries until it has enough information, then synthesizes a final report. ```typescript async function agenticResearchWorkflowFunction( topic: string, maxIterations: number, ): Promise { console.log(`Starting agentic research for: ${topic}`); const allFindings: Finding[] = []; const researchHistory: IterationResult[] = []; let currentIteration = 0; let currentQuery = topic; // Main agentic research loop while (currentIteration < maxIterations) { currentIteration++; console.log(`🔄 Starting iteration ${currentIteration}/${maxIterations}`); // Research the next query const iterationResult = await researchQueryWorkflow( topic, currentQuery, currentIteration, ); researchHistory.push(iterationResult); allFindings.push(iterationResult.evaluation); // Handle cases where no results are found const storiesFound = iterationResult.stories_found; if (storiesFound === 0) { console.log( `⚠️ No stories found for '${currentQuery}', trying alternative approach...`, ); // Generate alternative queries when hitting dead ends const alternativeQuery = await DBOS.runStep( () => generateFollowUps(topic, allFindings, currentIteration), { name: "generateFollowUps" }, ); if (alternativeQuery) { currentQuery = alternativeQuery; console.log(`🔄 Retrying with: '${currentQuery}'`); continue; } else { console.log("❌ No alternative queries available, continuing..."); } } // Evaluate whether to continue research console.log("🤔 Agent evaluating whether to continue research..."); const shouldContinueDecision = await DBOS.runStep( () => shouldContinue(topic, allFindings, currentIteration, maxIterations), { name: "shouldContinue" }, ); if (!shouldContinueDecision) { console.log("✅ Agent decided to conclude research"); break; } // Generate next research question based on findings if (currentIteration < maxIterations) { console.log("💭 Agent generating next research question..."); const followUpQuery = await DBOS.runStep( () => generateFollowUps(topic, allFindings, currentIteration), { name: "generateFollowUps" }, ); if (followUpQuery) { currentQuery = followUpQuery; console.log(`➡️ Next research focus: '${currentQuery}'`); } else { console.log("💡 No new research directions found, concluding..."); break; } } } // Final step: Synthesize all findings into comprehensive report console.log("📋 Agent synthesizing final research report..."); const finalReport = await DBOS.runStep( () => synthesizeFindings(topic, allFindings), { name: "synthesizeFindings" }, ); // Return complete research results return { topic, total_iterations: currentIteration, max_iterations: maxIterations, research_history: researchHistory, final_report: finalReport, summary: { total_stories: researchHistory.reduce( (sum, r) => sum + r.stories_found, 0, ), total_comments: researchHistory.reduce( (sum, r) => sum + r.comments_analyzed, 0, ), queries_executed: researchHistory.map((r) => r.query), avg_relevance: allFindings.length > 0 ? allFindings.reduce((sum, f) => sum + (f.relevance_score || 0), 0) / allFindings.length : 0, }, }; } export const agenticResearchWorkflow = DBOS.registerWorkflow( agenticResearchWorkflowFunction, { name: "agenticResearchWorkflow" }, ); ``` ### Research Query Workflow Each iteration of the main research workflow calls a child workflow that searches Hacker News for information about a query, then evaluates and returns its findings. ```typescript async function researchQueryWorkflowFunction( topic: string, query: string, iteration: number, ): Promise { console.log(`🔍 Searching for stories: '${query}'`); // Step 1: Search Hacker News for stories about the topic const stories = await DBOS.runStep(() => searchHackerNews(query, 30), { name: "searchHackerNews", }); if (stories.length > 0) { console.log(`📚 Found ${stories.length} stories, analyzing all stories...`); stories.forEach((story, i) => { const title = (story.title || "No title").slice(0, 80); const points = story.points || 0; const numComments = story.num_comments || 0; console.log( ` 📖 Story ${i + 1}: ${title}... (${points} points, ${numComments} comments)`, ); }); } else { console.log("❌ No stories found for this query"); } // Step 2: Gather comments from all stories found const comments: any[] = []; if (stories.length > 0) { console.log(`💬 Reading comments from ALL ${stories.length} stories...`); for (let i = 0; i < stories.length; i++) { const story = stories[i]; const storyId = story.objectID; const title = (story.title || "Unknown").slice(0, 50); const numComments = story.num_comments || 0; if (storyId && numComments > 0) { console.log( ` 💭 Reading comments from: ${title}... (${numComments} comments)`, ); const storyComments = await DBOS.runStep( () => getComments(storyId, 10), { name: "getComments" }, ); comments.push(...storyComments); console.log(` ✓ Read ${storyComments.length} comments`); } else if (storyId) { console.log(` 📖 Story has no comments: ${title}`); } else { console.log(` ❌ No story ID available for: ${title}`); } } } // Step 3: Evaluate gathered data and return findings console.log( `🤔 Analyzing findings from ${stories.length} stories and ${comments.length} comments...`, ); const evaluation = await DBOS.runStep( () => evaluateResults(topic, query, stories, comments), { name: "evaluateResults" }, ); return { iteration, query, stories_found: stories.length, comments_analyzed: comments.length, evaluation, stories, comments, }; } export const researchQueryWorkflow = DBOS.registerWorkflow( researchQueryWorkflowFunction, { name: "researchQueryWorkflow" }, ); ``` ### Agent Decision-Making Steps The agent's intelligence comes from three key step functions that handle decision-making:
Agent Evaluation Step ```typescript export async function evaluateResults( topic: string, query: string, stories: any[], comments?: any[], ): Promise { let storiesText = ""; const topStories: Story[] = []; // Evaluate only the top 10 most relevant stories stories.slice(0, 10).forEach((story, i) => { const title = story.title || "No title"; const url = story.url || "No URL"; const hnUrl = `https://news.ycombinator.com/item?id=${story.objectID || ""}`; const points = story.points || 0; const numComments = story.num_comments || 0; const author = story.author || "Unknown"; storiesText += `Story ${i + 1}:\n`; storiesText += ` Title: ${title}\n`; storiesText += ` Points: ${points}, Comments: ${numComments}\n`; storiesText += ` URL: ${url}\n`; storiesText += ` HN Discussion: ${hnUrl}\n`; storiesText += ` Author: ${author}\n\n`; topStories.push({ title, url, hn_url: hnUrl, points, num_comments: numComments, author, objectID: story.objectID || "", }); }); let commentsText = ""; if (comments) { comments.slice(0, 20).forEach((comment, i) => { const commentText = comment.comment_text || ""; if (commentText) { const author = comment.author || "Unknown"; const excerpt = commentText.length > 400 ? commentText.slice(0, 400) + "..." : commentText; commentsText += `Comment ${i + 1}:\n`; commentsText += ` Author: ${author}\n`; commentsText += ` Text: ${excerpt}\n\n`; } }); } const prompt = ` You are a research agent evaluating search results for: ${topic} Query used: ${query} Stories found: ${storiesText} Comments analyzed: ${commentsText} Provide a DETAILED analysis with specific insights, not generalizations. Focus on: - Specific technical details, metrics, or benchmarks mentioned - Concrete tools, libraries, frameworks, or techniques discussed - Interesting problems, solutions, or approaches described - Performance data, comparison results, or quantitative insights - Notable opinions, debates, or community perspectives - Specific use cases, implementation details, or real-world examples Return JSON with: - "insights": Array of specific, technical insights with context - "relevance_score": Number 1-10 - "summary": Brief summary of findings - "key_points": Array of most important points discovered `; const messages = [ { role: "system" as const, content: "You are a research evaluation agent. Analyze search results and provide structured insights in JSON format.", }, { role: "user" as const, content: prompt }, ]; try { const response = await callLLM(messages, "gpt-4o-mini", 0.1, 2000); const cleanedResponse = cleanJsonResponse(response); const evaluation = JSON.parse(cleanedResponse); evaluation.query = query; evaluation.top_stories = topStories; return evaluation; } catch (error) { return { insights: [`Found ${stories.length} stories about ${topic}`], relevance_score: 7, summary: `Basic search results for ${query}`, key_points: [], query, }; } } ```
Follow-up Query Generation Step ```typescript export async function generateFollowUps( topic: string, currentFindings: Finding[], iteration: number, ): Promise { let findingsSummary = ""; currentFindings.forEach((finding) => { findingsSummary += `Query: ${finding.query || "Unknown"}\n`; findingsSummary += `Summary: ${finding.summary || "No summary"}\n`; findingsSummary += `Key insights: ${JSON.stringify(finding.insights || [])}\n`; findingsSummary += `Unanswered questions: ${JSON.stringify(finding.unanswered_questions || [])}\n\n`; }); const prompt = ` You are a research agent investigating: ${topic} This is iteration ${iteration} of your research. Current findings: ${findingsSummary} Generate 2-4 SHORT KEYWORD-BASED search queries for Hacker News that explore DIVERSE aspects of ${topic}. CRITICAL RULES: 1. Use SHORT keywords (2-4 words max) - NOT long sentences 2. Focus on DIFFERENT aspects of ${topic}, not just one narrow area 3. Use terms that appear in actual Hacker News story titles 4. Avoid repeating previous focus areas 5. Think about what tech people actually discuss about ${topic} For ${topic}, consider diverse areas like: - Performance/optimization - Tools/extensions - Comparisons with other technologies - Use cases/applications - Configuration/deployment - Recent developments GOOD examples: ["postgres performance", "database tools", "sql optimization"] BAD examples: ["What are the best practices for PostgreSQL optimization?"] Return only a JSON array of SHORT keyword queries: ["query1", "query2", "query3"] `; const messages = [ { role: "system" as const, content: "You are a research agent. Generate focused follow-up queries based on current findings. Return only JSON array.", }, { role: "user" as const, content: prompt }, ]; try { const response = await callLLM(messages); const cleanedResponse = cleanJsonResponse(response); const queries = JSON.parse(cleanedResponse); return Array.isArray(queries) && queries.length > 0 ? queries[0] : null; } catch (error) { return null; } } ```
Continuation Decision Step ```typescript export async function shouldContinue( topic: string, allFindings: Finding[], currentIteration: number, maxIterations: number, ): Promise { if (currentIteration >= maxIterations) { return false; } let findingsSummary = ""; let totalRelevance = 0; allFindings.forEach((finding) => { findingsSummary += `Query: ${finding.query || "Unknown"}\n`; findingsSummary += `Summary: ${finding.summary || "No summary"}\n`; findingsSummary += `Relevance: ${finding.relevance_score || 5}/10\n`; totalRelevance += finding.relevance_score || 5; }); const avgRelevance = allFindings.length > 0 ? totalRelevance / allFindings.length : 0; const prompt = ` You are a research agent investigating: ${topic} Current iteration: ${currentIteration}/${maxIterations} Findings so far: ${findingsSummary} Average relevance score: ${avgRelevance.toFixed(1)}/10 Decide whether to continue research or conclude. PRIORITIZE THOROUGH EXPLORATION - continue if: 1. Current iteration is less than 75% of max_iterations 2. Average relevance is above 6.0 and there are likely unexplored aspects 3. Recent queries found significant new information 4. The research seems to be discovering diverse perspectives on the topic Only stop early if: - Average relevance is below 5.0 for multiple iterations - No new meaningful information in the last 2 iterations - Research appears to be hitting diminishing returns Return JSON with: - "should_continue": boolean `; const messages = [ { role: "system" as const, content: "You are a research decision agent. Evaluate research completeness and decide whether to continue. Return JSON.", }, { role: "user" as const, content: prompt }, ]; try { const response = await callLLM(messages); const cleanedResponse = cleanJsonResponse(response); const decision = JSON.parse(cleanedResponse); return decision.should_continue || true; } catch (error) { return true; } } ```
### Search API Steps After deciding what terms to search for, the agent calls these steps to retrieve stories and comments from Hacker News.
Hacker News API Steps ```typescript export const searchHackerNews = async ( query: string, maxResults = 20, ): Promise => { try { const response = await fetch( `${HN_SEARCH_URL}?${new URLSearchParams({ query, hitsPerPage: maxResults.toString(), tags: "story", })}`, { signal: AbortSignal.timeout(30000) }, ); if (!response.ok) throw new Error(`HTTP error! status: ${response.status}`); const data = (await response.json()) as { hits: HackerNewsStory[] }; return data.hits ?? []; } catch (error) { console.error("Error searching Hacker News:", error); return []; } }; export const getComments = async ( storyId: string, maxComments = 50, ): Promise => { try { const response = await fetch( `${HN_SEARCH_URL}?${new URLSearchParams({ tags: `comment,story_${storyId}`, hitsPerPage: maxComments.toString(), })}`, { signal: AbortSignal.timeout(30000) }, ); if (!response.ok) throw new Error(`HTTP error! status: ${response.status}`); const data = (await response.json()) as { hits: HackerNewsComment[] }; return data.hits ?? []; } catch (error) { console.error("Error getting comments:", error); return []; } }; ```
### Synthesize Findings Step Finally, after concluding its research, the agentic workflow calls this step to synthesize its findings into a report.
Synthesize Findings Step ```typescript export async function synthesizeFindings( topic: string, allFindings: Finding[], ): Promise { let findingsText = ""; const storyLinks: Story[] = []; allFindings.forEach((finding, i) => { findingsText += `\n=== Finding ${i + 1} ===\n`; findingsText += `Query: ${finding.query || "Unknown"}\n`; findingsText += `Summary: ${finding.summary || "No summary"}\n`; findingsText += `Key Points: ${JSON.stringify(finding.key_points || [])}\n`; findingsText += `Insights: ${JSON.stringify(finding.insights || [])}\n`; if (finding.top_stories) { finding.top_stories.forEach((story) => { storyLinks.push({ title: story.title || "Unknown", url: story.url || "", hn_url: `https://news.ycombinator.com/item?id=${story.objectID || ""}`, points: story.points || 0, num_comments: story.num_comments || 0, }); }); } }); const storyCitations: Record = {}; let citationId = 1; allFindings.forEach((finding) => { if (finding.top_stories) { finding.top_stories.forEach((story) => { const storyId = story.objectID || ""; if (storyId && !storyCitations[storyId]) { storyCitations[storyId] = { id: citationId, title: story.title || "Unknown", url: story.url || "", hn_url: story.hn_url || "", points: story.points || 0, comments: story.num_comments || 0, }; citationId++; } }); } }); const citationsText = Object.values(storyCitations) .map( (cite) => `[${cite.id}] ${cite.title} (${cite.points} points, ${cite.comments} comments) - ${cite.hn_url}` + (cite.url ? ` - ${cite.url}` : ""), ) .join("\n"); const prompt = ` You are a research analyst. Synthesize the following research findings into a comprehensive, detailed report about: ${topic} Research Findings: ${findingsText} Available Citations: ${citationsText} IMPORTANT: You must return ONLY a valid JSON object with no additional text, explanations, or formatting. Create a comprehensive research report that flows naturally as a single narrative. Include: - Specific technical details and concrete examples - Actionable insights practitioners can use - Interesting discoveries and surprising findings - Specific tools, libraries, or techniques mentioned - Performance metrics, benchmarks, or quantitative data when available - Notable opinions or debates in the community - INLINE LINKS: When making claims, include clickable links directly in the text using this format: [link text](HN_URL) - Use MANY inline links throughout the report. Aim for at least 4-5 links per paragraph. CRITICAL CITATION RULES - FOLLOW EXACTLY: 1. NEVER replace words with bare URLs like "(https://news.ycombinator.com/item?id=123)" 2. ALWAYS write complete sentences with all words present 3. Add citations using descriptive link text in brackets: [descriptive text](URL) 4. Every sentence must be grammatically complete and readable without the links 5. Links should ALWAYS be to the Hacker News discussion, NEVER directly to the article. CORRECT examples: "PostgreSQL's performance improvements have been significant in recent versions, as discussed in [community forums](https://news.ycombinator.com/item?id=123456), with developers highlighting [specific optimizations](https://news.ycombinator.com/item?id=789012) in query processing." "Redis performance issues can stem from common configuration mistakes, which are well-documented in [troubleshooting guides](https://news.ycombinator.com/item?id=345678) and [community discussions](https://news.ycombinator.com/item?id=901234)." "React's licensing changes have sparked significant community debate, as seen in [detailed discussions](https://news.ycombinator.com/item?id=15316175) about the implications for open-source projects." WRONG examples (NEVER DO THIS): "Community discussions reveal a strong interest in the (https://news.ycombinator.com/item?id=18717168) and the common pitfalls" "One significant topic is the (https://news.ycombinator.com/item?id=15316175), which raises important legal considerations" Always link to relevant discussions for: - Every specific tool, library, or technology mentioned - Performance claims and benchmarks - Community opinions and debates - Technical implementation details - Companies or projects referenced - Version releases or updates - Problem reports or solutions Return a JSON object with this exact structure: { "report": "A comprehensive research report written as flowing narrative text with inline clickable links [like this](https://news.ycombinator.com/item?id=123). Include specific technical details, tools, performance metrics, community opinions, and actionable insights. Make it detailed and informative, not just a summary." } `; const messages: Message[] = [ { role: "system", content: "You are a research analyst. Provide comprehensive synthesis in JSON format.", }, { role: "user", content: prompt }, ]; try { const response = await callLLM( messages, DEFAULT_MODEL, DEFAULT_TEMPERATURE, 3000, ); const cleanedResponse = cleanJsonResponse(response); const result = JSON.parse(cleanedResponse); return result; } catch (error) { return { report: "JSON parsing error, report could not be generated.", error: `JSON parsing failed, created basic synthesis. Error: ${error}`, }; } } ```
### Try it Yourself! #### Setting Up OpenAI To run this agent, you need an OpenAI developer account. Obtain an API key [here](https://platform.openai.com/api-keys) and set up a payment method for your account [here](https://platform.openai.com/account/billing/overview). This agent uses `gpt-4o-mini` for decision-making. Set your API key as an environment variable: ```shell export OPENAI_API_KEY= ``` #### Running Locally First, clone this repository: ```shell git clone https://github.com/dbos-inc/dbos-demo-apps.git cd dbos-demo-apps/typescript/hacker-news-agent ``` Install dependencies and build the project: ```bash npm install npm run build ``` Start Postgres (if you already use Postgres, instead set the `DBOS_SYSTEM_DATABASE_URL` environment variable to your database connection string): ```bash npx dbos postgres start ``` Run the agent with any research topic: ```bash npx agent "artificial intelligence" ``` Or try other topics: ```shell npx agent "rust" npx agent "postgres" npx agent "kubernetes" ``` The agent will autonomously research your topic, make decisions about what to investigate next, and produce a research report with insights from Hacker News. If the agent fails at any point during its research, you can restart it using its workflow ID to recover it from where it left off: ```shell npx agent "artificial intelligence" --workflow-id ``` --- ## Queue Worker(Examples) :::info This example is also available in [Python](../../python/examples/queue-worker.md). ::: This example demonstrates how to run DBOS workflows in their own "queue worker" service while enqueueing and managing them from other services. This design pattern lets you separate concerns and separately scale the workers that execute your durable workflows from your other services. Architecturally, this example contains two services: a web server and a worker service. The web server uses the [DBOS Client](../reference/client.md) to enqueue workflows and monitor their status. The worker service dequeues and executes workflows. All source code is [available on GitHub](https://github.com/dbos-inc/dbos-demo-apps/tree/main/typescript/queue-worker). ### Worker Service The worker service implements your durable workflows and their steps. Notably, this workflow periodically reports its progress using [`DBOS.setEvent`](../tutorials/workflow-communication.md). This lets the web server query the event to monitor workflow progress. ```ts DBOS.registerWorkflow( async function (numSteps: number): Promise { const progress = { steps_completed: 0, num_steps: numSteps, }; // The server can query this event to obtain // the current progress of the workflow await DBOS.setEvent(WF_PROGRESS_KEY, progress); for (let i = 0; i < numSteps; i++) { await DBOS.runStep(() => stepFunction(i)); // Update workflow progress each time a step completes progress.steps_completed = i + 1; await DBOS.setEvent(WF_PROGRESS_KEY, progress); } }, { name: 'workflow' }, ); async function stepFunction(i: number): Promise { console.log(`Step ${i} completed!`); // Sleep one second await new Promise((resolve) => setTimeout(resolve, 1000)); } ``` In its main function, the worker service configures and launches DBOS, registers the queue on which the web server can submit workflows for execution, then waits indefinitely, dequeuing and executing workflows: ```ts async function main(): Promise { const systemDatabaseUrl = process.env.DBOS_SYSTEM_DATABASE_URL || 'postgresql://postgres:dbos@localhost:5432/dbos_queue_worker'; DBOS.setConfig({ name: 'dbos-queue-worker', applicationVersion: '0.1.0', systemDatabaseUrl: systemDatabaseUrl, }); await DBOS.launch(); // Define a queue on which the web server // can submit workflows for execution. await DBOS.registerQueue('workflow-queue'); // After launching DBOS, the worker waits indefinitely, // dequeuing and executing workflows. console.log('Worker started, waiting for workflows...'); await new Promise(() => {}); } main().catch(console.log); ``` ### Web Server The web server first creates a DBOS Client: ```ts const systemDatabaseUrl = process.env.DBOS_SYSTEM_DATABASE_URL || 'postgresql://postgres:dbos@localhost:5432/dbos_queue_worker'; const client = await DBOSClient.create({ systemDatabaseUrl }); ``` It then enqueues workflows using the client: ```ts app.post('/api/workflows', async (_req: Request, res: Response) => { const numSteps = 10; await client.enqueue( { queueName: 'workflow-queue', workflowName: 'workflow', }, numSteps, ); res.json({ status: 'enqueued' }); }); ``` The web server can also report workflow status. This function first lists all workflows, then uses [`getEvent`](../tutorials/workflow-communication.md) to query the progress of each workflow. This is a useful pattern for showing workflow progress or status to end users of your application. ```ts app.get('/api/workflows', async (_req: Request, res: Response) => { // Use the DBOS client to list all workflows const workflows = await client.listWorkflows({ workflowName: 'workflow', sortDesc: true, }); const statuses: WorkflowStatus[] = []; for (const workflow of workflows) { // Query each workflow's progress event. This may not be available // if the workflow has not yet started executing. const progress = await client.getEvent(workflow.workflowID, WF_PROGRESS_KEY, 0); const status: WorkflowStatus = { workflow_id: workflow.workflowID, workflow_status: workflow.status, steps_completed: progress ? progress.steps_completed : null, num_steps: progress ? progress.num_steps : null, }; statuses.push(status); } res.json(statuses); }); ``` ### Try it Yourself! Clone and enter the [dbos-demo-apps](https://github.com/dbos-inc/dbos-demo-apps) repository: ```shell git clone https://github.com/dbos-inc/dbos-demo-apps.git cd dbos-demo-apps/typescript/queue-worker ``` Then follow the instructions in the [README](https://github.com/dbos-inc/dbos-demo-apps/tree/main/typescript/queue-worker) to run the app. --- ## Add DBOS To Your App(Typescript) This guide shows you how to add the open-source [DBOS Transact](https://github.com/dbos-inc/dbos-transact-ts) library to your existing application to **durably execute** it and make it resilient to any failure. :::info Also check out the integration guides for popular TypeScript frameworks: - [Nest.js + DBOS](../integrations/nestjs.md) ::: :::warning Due to its internal workflow registry, the DBOS library and DBOS workflows cannot be bundled with JavaScript or TypeScript bundlers (Webpack, Vite, Rollup, esbuild, Parcel, etc.) and must be treated as an external library by these tools. ::: #### Using DBOS Transact ##### 1. Install DBOS `npm install` DBOS into your application. Note that DBOS requires Node.js 20 or later. ```shell npm install @dbos-inc/dbos-sdk@latest ``` **Optionally**, if you want to use TypeScript decorators, enable them in your `tsconfig.json` file: ```json title="tsconfig.json" "compilerOptions": { "experimentalDecorators": true, } ``` DBOS requires a Postgres database. If you already have Postgres, you can set the `DBOS_SYSTEM_DATABASE_URL` environment variable to your connection string (later we'll pass that value into DBOS). Otherwise, you can start Postgres in a Docker container with this command: ```shell npx dbos postgres start ``` ##### 2. Launch DBOS in Your App In your app's main entrypoint, add code to configure and launch DBOS: ```javascript import { DBOS } from "@dbos-inc/dbos-sdk"; DBOS.setConfig({ "name": "my-app", "applicationVersion": "0.1.0", "systemDatabaseUrl": process.env.DBOS_SYSTEM_DATABASE_URL, }); await DBOS.launch(); ``` Both `DBOS.setConfig` and `DBOS.launch` should be called after your app is created, but before it begins processing requests. ##### 3. Start Your Application Try starting your application. If everything is set up correctly, your app should run normally, but log `DBOS launched!` on startup. Congratulations! You've integrated DBOS into your application. ##### 4. Start Building With DBOS At this point, you can apply DBOS durability to your functions. For example, you can register one of your functions as a [workflow](./tutorials/workflow-tutorial.md) and call other functions as [steps](./tutorials/step-tutorial.md). DBOS durably executes the workflow so if it is ever interrupted, upon restart it automatically resumes from the last completed step. ```typescript async function stepOne() { DBOS.logger.info("Step one completed!"); } async function stepTwo() { DBOS.logger.info("Step two completed!"); } async function workflowFunction() { await DBOS.runStep(() => stepOne(), {name: "stepOne"}); await DBOS.runStep(() => stepTwo(), {name: "stepTwo"}); } const workflow = DBOS.registerWorkflow(workflowFunction) await workflow(); ``` **You must register all workflows before calling `DBOS.launch()`** As workflow recovery will commence after `DBOS.launch()`, it is essential that all workflows be registered before this point. You can add DBOS to your application incrementally—it won't interfere with code that's already there. It's totally okay for your application to have one DBOS workflow alongside thousands of lines of non-DBOS code. To learn more about programming with DBOS, check out [the programming guide](./programming-guide.md). --- ## Learn DBOS TypeScript This guide shows you how to use DBOS to build TypeScript apps that are **resilient to any failure**. :::tip To teach your AI coding assistant to build with DBOS, try out [skills](./prompting.md) and [MCP](../integrations/mcp.md). ::: ### 1. Setting Up Your App To get started, initialize a DBOS template and install dependencies: ```shell npx @dbos-inc/create@latest -t dbos-node-starter cd dbos-node-starter ``` DBOS requires a Postgres database. If you already have Postgres, you can set the `DBOS_SYSTEM_DATABASE_URL` environment variable to your connection string (later we'll pass that value into DBOS). Otherwise, you can start Postgres in a Docker container with this command: ```shell npx dbos postgres start ``` ### 2. Workflows and Steps DBOS helps you add reliability to your TypeScript programs. The key feature of DBOS is **workflow functions** comprised of **steps**. DBOS checkpoints the state of your workflows and steps to its system database. If your program crashes or is interrupted, DBOS uses this checkpointed state to recover each of your workflows from its last completed step. Thus, DBOS makes your application **resilient to any failure**. Let's create a simple DBOS program that runs a workflow of two steps. Replace all the code in `src/main.ts` with the following: ```javascript showLineNumbers title="src/main.ts" import { DBOS } from "@dbos-inc/dbos-sdk"; async function stepOne() { DBOS.logger.info("Step one completed!"); } async function stepTwo() { DBOS.logger.info("Step two completed!"); } async function exampleFunction() { await DBOS.runStep(() => stepOne(), {name: "stepOne"}); await DBOS.runStep(() => stepTwo(), {name: "stepTwo"}); } const exampleWorkflow = DBOS.registerWorkflow(exampleFunction); async function main() { DBOS.setConfig({ "name": "dbos-node-starter", "applicationVersion": "0.1.0", "systemDatabaseUrl": process.env.DBOS_SYSTEM_DATABASE_URL, }); await DBOS.launch(); await exampleWorkflow(); await DBOS.shutdown(); } main().catch(console.log); ``` Now, build and run this code with: ```shell npm run build npm run start ``` Your program should print output like: ``` DBOS launched! Step one completed! Step two completed! ``` To see durable execution in action, let's modify the app to serve a DBOS workflow from an HTTP endpoint using Express.js. Replace the contents of `src/main.ts` with: ```javascript showLineNumbers title="src/main.ts" import { DBOS } from "@dbos-inc/dbos-sdk"; import express from "express"; export const app = express(); app.use(express.json()); async function stepOne() { DBOS.logger.info("Step one completed!"); } async function stepTwo() { DBOS.logger.info("Step two completed!"); } async function exampleFunction() { await DBOS.runStep(() => stepOne(), {name: "stepOne"}); for (let i = 0; i < 5; i++) { console.log("Press Control + C to stop the app..."); await DBOS.sleep(1000); } await DBOS.runStep(() => stepTwo(), {name: "stepTwo"}); } const exampleWorkflow = DBOS.registerWorkflow(exampleFunction); app.get("/", async (req, res) => { await exampleWorkflow(); res.send(); }); async function main() { DBOS.setConfig({ "name": "dbos-node-starter", "applicationVersion": "0.1.0", "systemDatabaseUrl": process.env.DBOS_SYSTEM_DATABASE_URL, }); await DBOS.launch(); const PORT = 3000; app.listen(PORT, () => { console.log(`🚀 Server is running on http://localhost:${PORT}`); }); } main().catch(console.log); ``` Now, install Express.js and its types, then rebuild and restart your app with: ```shell npm install express @types/express npm run build npm run start ``` Then, visit this URL: http://localhost:3000. In your terminal, you should see an output like: ``` 🚀 Server is running on http://localhost:3000 Step one completed! Press Control + C to stop the app... Press Control + C to stop the app... Press Control + C to stop the app... Press Control + C to stop the app... ``` Now, press CTRL+C to stop your app. Then, run `npm run start` to restart it. You should see an output like: ``` 🚀 Server is running on http://localhost:3000 Press Control + C to stop the app... Press Control + C to stop the app... Press Control + C to stop the app... Press Control + C to stop the app... Press Control + C to stop the app... Step two completed! ``` You can see how DBOS **recovers your workflow from the last completed step**, executing step two without re-executing step one. Learn more about workflows, steps, and their guarantees [here](./tutorials/workflow-tutorial.md). ### 3. Queues and Parallelism If you need to run many functions concurrently, use DBOS _queues_. To try them out, copy this code into `src/main.ts`: ```javascript showLineNumbers title="src/main.ts" import { DBOS } from "@dbos-inc/dbos-sdk"; import express from "express"; export const app = express(); app.use(express.json()); async function taskFunction(n: number) { await DBOS.sleep(5000); DBOS.logger.info(`Task ${n} completed!`) } const taskWorkflow = DBOS.registerWorkflow(taskFunction); async function queueFunction() { DBOS.logger.info("Enqueueing tasks!") const handles = [] for (let i = 0; i < 10; i++) { handles.push(await DBOS.startWorkflow(taskWorkflow, { queueName: "example_queue" })(i)) } const results = [] for (const h of handles) { results.push(await h.getResult()) } DBOS.logger.info(`Successfully completed ${results.length} tasks`) } const queueWorkflow = DBOS.registerWorkflow(queueFunction) app.get("/", async (req, res) => { await queueWorkflow(); res.send(); }); async function main() { DBOS.setConfig({ "name": "dbos-node-starter", "applicationVersion": "0.1.0", "systemDatabaseUrl": process.env.DBOS_SYSTEM_DATABASE_URL, }); await DBOS.launch(); await DBOS.registerQueue("example_queue"); const PORT = 3000; app.listen(PORT, () => { console.log(`🚀 Server is running on http://localhost:${PORT}`); }); } main().catch(console.log); ``` When you enqueue a function with `DBOS.startWorkflow`, DBOS executes it _asynchronously_, running it in the background without waiting for it to finish. `DBOS.startWorkflow` returns a handle representing the state of the enqueued function. This example enqueues ten functions, then waits for them all to finish using `getResult()` to wait for each of their handles. Now, rebuild and restart your app with: ```shell npm run build npm run start ``` Then, visit this URL: http://localhost:3000. Wait five seconds and you should see an output like: ``` 🚀 Server is running on http://localhost:3000 Enqueueing tasks! Task 0 completed! Task 1 completed! Task 2 completed! Task 3 completed! Task 4 completed! Task 5 completed! Task 6 completed! Task 7 completed! Task 8 completed! Task 9 completed! Successfully completed 10 tasks ``` You can see how all ten steps run concurrently—even though each takes five seconds, they all finish at the same time. Learn more about DBOS queues [here](./tutorials/queue-tutorial.md). ### 4. Connecting to DBOS Conductor [Conductor](../conductor/overview.md) is the control plane for your durable workflows, providing distributed workflow recovery, observability, and management. Once you connect your app to Conductor, you can view and manage all its workflows and queued tasks from the [DBOS Console](https://console.dbos.dev). To connect your app to Conductor, first sign up for an account on the [DBOS Console](https://console.dbos.dev/login-redirect). Then, install [`dbosctl`](../conductor/reference/dbosctl.md), the Conductor command-line client. On Windows, [download a release binary](https://github.com/dbos-inc/dbos-ctl/releases) instead. ```shell curl -sSfL https://raw.githubusercontent.com/dbos-inc/dbos-ctl/main/install.sh | sh ``` Next, configure a `dbosctl` profile and log in. `dbosctl login` prints a URL and a code for you to approve in your browser. ```shell dbosctl config set dbos --managed dbosctl login ``` Then, register your application with Conductor and create an API key. The name you register must match the `name` in your DBOS configuration. The key's secret is printed once and cannot be retrieved afterwards, so copy it now. ```shell dbosctl app register dbos-node-starter dbosctl api-key create dbos-node-starter-key ``` Next, supply your API key to your app through the `conductorKey` launch option. Update the call to `DBOS.launch` in `main` to read the key from an environment variable: ```javascript await DBOS.launch({ conductorKey: process.env.DBOS_CONDUCTOR_KEY }); ``` Finally, set the `DBOS_CONDUCTOR_KEY` environment variable to the key you created, then rebuild and restart your app: ```shell export DBOS_CONDUCTOR_KEY= npm run build npm run start ``` Your app is now connected to Conductor! Launch a workflow by visiting http://localhost:3000, then watch it execute in real time from the [DBOS Console](https://console.dbos.dev). Learn more about Conductor [here](../conductor/overview.md). Congratulations! You've finished the DBOS TypeScript guide. Next, you should: - Learn how to [**add DBOS to your own application**](./integrating-dbos.md). - Check out some [**example applications**](../examples/index.md). --- ## AI-Assisted Development(Typescript) If you're using an AI coding agent to build a DBOS application, make sure it has the latest information on DBOS by either: 1. [Installing DBOS skills.](#dbos-agent-skills) 2. [Providing your agent with a DBOS prompt.](#dbos-prompt) You may also want to use the [DBOS MCP server](../integrations/mcp.md) so your model can directly access your application's workflows and steps. ### DBOS Agent Skills [Agent Skills](https://agentskills.io/home) help developers use AI agents to add DBOS durable workflows to their applications. DBOS provides open-source skills you can check out [here](https://github.com/dbos-inc/agent-skills). To install them into your coding agent, run: ``` npx skills add dbos-inc/agent-skills ``` The [Skills CLI](https://skills.sh/) is compatible with most coding agents, including Claude Code, Codex, Antigravity, and Cursor. ### DBOS Prompt You can use this prompt to add rich information about DBOS to your AI coding agent's context. You can copy and paste it directly into your context, or follow these directions to add it to your AI-powered IDE or coding agent of choice: - Claude Code: Add the prompt, or a link to it, to your CLAUDE.md file. - Cursor: Add the prompt to [your project rules](https://docs.cursor.com/context/rules-for-ai). - GitHub Copilot: Create a [`.github/copilot-instructions.md`](https://docs.github.com/en/copilot/customizing-copilot/adding-repository-custom-instructions-for-github-copilot) file in your repository and add the prompt to it.
DBOS TypeScript Prompt ````markdown # Build Reliable Applications With DBOS ## Guidelines - Respond in a friendly and concise manner - Ask clarifying questions when requirements are ambiguous - Generate code in TypeScript using the DBOS library. Make sure to fully type everything. - You MUST import all methods and classes used in the code you generate - You SHALL keep all code in a single file unless otherwise specified. - You MUST await all promises. - DBOS does NOT stand for anything. ## Workflow Guidelines Workflows provide durable execution so you can write programs that are resilient to any failure. Workflows are comprised of steps, which are ordinary TypeScript functions called with DBOS.runStep(). When using DBOS workflows, you should call any function that performs complex operations or accesses external APIs or services as a step using DBOS.runStep. If a workflow is interrupted for any reason (e.g., an executor restarts or crashes), when your program restarts the workflow automatically resumes execution from the last completed step. - If asked to add DBOS to existing code, you MUST ask which function to make a workflow. Do NOT recommend any changes until they have told you what function to make a workflow. Do NOT make a function a workflow unless SPECIFICALLY requested. - When making a function a workflow, you should make all functions it calls steps. Do NOT change the functions in any way. - Do NOT make functions steps unless they are DIRECTLY called by a workflow. - If the workflow function performs a non-deterministic action, you MUST move that action to its own function and make that function a step. Examples of non-deterministic actions include accessing an external API or service, accessing files on disk, generating a random number, or getting the current time. - Do NOT use Promise.all() due to the risks posed by multiple rejections. Using Promise.allSettled() for parallelism is allowed for single-step promises only. For any complex parallel execution, you should instead use DBOS.startWorkflow and DBOS queues to achieve the parallelism. - DBOS workflows and steps should NOT have side effects in memory outside of their own scope. They can access global variables, but they should NOT create or update global variables or variables outside their scope. - Do NOT call any DBOS context method (DBOS.send, DBOS.recv, DBOS.startWorkflow, DBOS.sleep, DBOS.setEvent, DBOS.getEvent) from a step. - Do NOT start workflows from inside a step. - Do NOT call DBOS.setEvent and DBOS.recv from outside a workflow function. - Do NOT use DBOS.getApi, DBOS.postApi, or other DBOS HTTP annotations. These are no longer supported. Instead, use Express for HTTP serving by default, unless another web framework is specified. ## DBOS Lifecycle Guidelines DBOS should be installed and imported from the `@dbos-inc/dbos-sdk` package. Due to its internal workflow registry, the DBOS library and DBOS workflows cannot be bundled with JavaScript or TypeScript bundlers (Webpack, Vite, Rollup, esbuild, Parcel, etc.) and must be treated as an external library by these tools. Configuration for bundlers should be suggested if these tools are in use and cannot be avoided. DBOS programs MUST have a starting file (typically 'main.ts' or 'server.ts') that creates all objects and workflow functions during startup, before calling DBOS.launch(). Because DBOS runs workflows as long-running background jobs, prefer a long-lived process. DBOS can run on a "serverless" platform, but only where that startup file is evaluated on every invocation and the invocation lives long enough to dequeue and execute workflows; for that pattern, enqueue work from your application with a DBOS client and run a separate worker invocation that launches DBOS and drains the queue. Any DBOS program MUST call DBOS.setConfig and DBOS.launch in its main function, like so. You MUST use this default configuration (changing the name as appropriate) unless otherwise specified. ```javascript DBOS.setConfig({ "name": "dbos-node-starter", "applicationVersion": "0.1.0", "systemDatabaseUrl": process.env.DBOS_SYSTEM_DATABASE_URL, }); await DBOS.launch(); ``` Here is an example main function using Express: ```javascript import { DBOS } from "@dbos-inc/dbos-sdk"; async function main() { DBOS.setConfig({ "name": "dbos-node-starter", "applicationVersion": "0.1.0", "systemDatabaseUrl": process.env.DBOS_SYSTEM_DATABASE_URL, }); await DBOS.launch(); const PORT = 3000; app.listen(PORT, () => { console.log(`🚀 Server is running on http://localhost:${PORT}`); }); } main().catch(console.log); ``` ### Workflow and Steps Examples Simple example: ```javascript import { DBOS } from "@dbos-inc/dbos-sdk"; async function stepOne() { DBOS.logger.info("Step one completed!"); } async function stepTwo() { DBOS.logger.info("Step two completed!"); } async function exampleFunction() { await DBOS.runStep(() => stepOne()); await DBOS.runStep(() => stepTwo()); } const exampleWorkflow = DBOS.registerWorkflow(exampleFunction); async function main() { DBOS.setConfig({ "name": "dbos-node-starter", "applicationVersion": "0.1.0", "systemDatabaseUrl": process.env.DBOS_SYSTEM_DATABASE_URL, }); await DBOS.launch(); await exampleWorkflow(); await DBOS.shutdown(); } main().catch(console.log); ``` Example with Express: ```javascript import { DBOS } from "@dbos-inc/dbos-sdk"; import express from "express"; export const app = express(); app.use(express.json()); async function stepOne() { DBOS.logger.info("Step one completed!"); } async function stepTwo() { DBOS.logger.info("Step two completed!"); } async function exampleFunction() { await DBOS.runStep(() => stepOne()); await DBOS.runStep(() => stepTwo()); } const exampleWorkflow = DBOS.registerWorkflow(exampleFunction); app.get("/", async (req, res) => { await exampleWorkflow(); res.send(); }); async function main() { DBOS.setConfig({ "name": "dbos-node-starter", "applicationVersion": "0.1.0", "systemDatabaseUrl": process.env.DBOS_SYSTEM_DATABASE_URL, }); await DBOS.launch(); const PORT = 3000; app.listen(PORT, () => { console.log(`🚀 Server is running on http://localhost:${PORT}`); }); } main().catch(console.log); ``` Example with queues: ```javascript import { DBOS } from "@dbos-inc/dbos-sdk"; import express from "express"; export const app = express(); app.use(express.json()); async function taskFunction(n: number) { await DBOS.sleep(5000); DBOS.logger.info(`Task ${n} completed!`) } const taskWorkflow = DBOS.registerWorkflow(taskFunction); async function queueFunction() { DBOS.logger.info("Enqueueing tasks!") const handles = [] for (let i = 0; i < 10; i++) { handles.push(await DBOS.startWorkflow(taskWorkflow, { queueName: "example_queue" })(i)) } const results = [] for (const h of handles) { results.push(await h.getResult()) } DBOS.logger.info(`Successfully completed ${results.length} tasks`) } const queueWorkflow = DBOS.registerWorkflow(queueFunction) app.get("/", async (req, res) => { await queueWorkflow(); res.send(); }); async function main() { DBOS.setConfig({ "name": "dbos-node-starter", "applicationVersion": "0.1.0", "systemDatabaseUrl": process.env.DBOS_SYSTEM_DATABASE_URL, }); await DBOS.launch(); await DBOS.registerQueue("example_queue"); const PORT = 3000; app.listen(PORT, () => { console.log(`🚀 Server is running on http://localhost:${PORT}`); }); } main().catch(console.log); ``` #### Scheduled Workflow You can schedule DBOS workflows to run on a cron schedule. Schedules are stored in the database and can be created, paused, resumed, and deleted at runtime. A scheduled workflow MUST take two arguments: a `Date` (the scheduled execution time) and a context object: ```typescript import { DBOS } from "@dbos-inc/dbos-sdk"; async function myPeriodicTask(scheduledTime: Date, context: unknown) { DBOS.logger.info(`Running task scheduled for ${scheduledTime.toISOString()}`); } const myPeriodicTaskWorkflow = DBOS.registerWorkflow(myPeriodicTask); async function main() { DBOS.setConfig({ "name": "dbos-node-starter", "applicationVersion": "0.1.0", "systemDatabaseUrl": process.env.DBOS_SYSTEM_DATABASE_URL, }); await DBOS.launch(); await DBOS.applySchedules([ { scheduleName: "my-task-schedule", workflowFn: myPeriodicTaskWorkflow, schedule: "*/5 * * * *", // Every 5 minutes }, ]); } main().catch(console.log); ``` - You MUST create schedules after `DBOS.launch()`; schedule methods throw if DBOS has not been launched. - Use `DBOS.createSchedule({ scheduleName, workflowFn, schedule, context?, options? })` to create a schedule with a crontab expression. It throws if a schedule with that name already exists. - Use `DBOS.applySchedules` to atomically create or update multiple schedules at once. To define static schedules on program start, use `DBOS.applySchedules`, which updates schedules that already exist. - Optional settings (`automaticBackfill`, `cronTimezone`, `queueName`) go in the `options` object of `DBOS.createSchedule`, but are top-level fields of each `DBOS.applySchedules` entry. - Use `DBOS.pauseSchedule` and `DBOS.resumeSchedule` to pause and resume schedules. - Use `DBOS.updateSchedule` to change only some fields of an existing schedule. - Use `DBOS.deleteSchedule` to delete a schedule. - Use `DBOS.listSchedules` and `DBOS.getSchedule` to inspect schedules. - Use `DBOS.backfillSchedule` to enqueue missed executions for a time range. - Use `DBOS.triggerSchedule` to immediately trigger a schedule. - Scheduled workflows MUST be free functions or static class methods, not methods on `ConfiguredInstance` objects. - Each workflow enqueued by a schedule is tagged with its schedule's name (recorded in the workflow's status). Retrieve all runs of a schedule with `DBOS.listWorkflows({ scheduleName: "my-task-schedule" })`. ### Workflow Documentation: Workflows provide **durable execution** so you can write programs that are **resilient to any failure**. Workflows are comprised of steps, which wrap ordinary TypeScript (or JavaScript) functions. If a workflow is interrupted for any reason (e.g., an executor restarts or crashes), when your program restarts the workflow automatically resumes execution from the last completed step. To write a workflow, register a TypeScript function with `DBOS.registerWorkflow`. The function's inputs and outputs must be serializable to JSON. For example: ```typescript async function stepOne() { DBOS.logger.info("Step one completed!"); } async function stepTwo() { DBOS.logger.info("Step two completed!"); } async function workflowFunction() { await DBOS.runStep(() => stepOne(), {name: "stepOne"}); await DBOS.runStep(() => stepTwo(), {name: "stepTwo"}); } const workflow = DBOS.registerWorkflow(workflowFunction) await workflow(); ``` Alternatively, you can register workflows and steps with decorators: ```typescript export class Example { @DBOS.step() static async stepOne() { DBOS.logger.info("Step one completed!"); } @DBOS.step() static async stepTwo() { DBOS.logger.info("Step two completed!"); } // Call steps from workflows @DBOS.workflow() static async exampleWorkflow() { await Example.stepOne(); await Example.stepTwo(); } } await Example.exampleWorkflow(); ``` ### Starting Workflows In The Background One common use-case for workflows is building reliable background tasks that keep running even when your program is interrupted, restarted, or crashes. You can use `DBOS.startWorkflow` to start a workflow in the background. If you start a workflow this way, it returns a workflow handle, from which you can access information about the workflow or wait for it to complete and retrieve its result. Here's an example: ```javascript class Example { @DBOS.workflow() static async exampleWorkflow(var1: string, var2: string) { return var1 + var2; } } async function main() { // Start exampleWorkflow in the background const handle = await DBOS.startWorkflow(Example).exampleWorkflow("one", "two"); // Wait for the workflow to complete and return its results const result = await handle.getResult(); } ``` After starting a workflow in the background, you can use `DBOS.retrieveWorkflow` to retrieve a workflow's handle from its ID. You can also retrieve a workflow's handle from outside of your DBOS application with `DBOSClient.retrieveWorkflow`. If you need to run many workflows in the background and manage their concurrency or flow control, you can also use DBOS queues. ### Workflow IDs and Idempotency Every time you execute a workflow, that execution is assigned a unique ID, by default a UUID. You can access this ID through the `DBOS.workflowID` context variable. Workflow IDs are useful for communicating with workflows and developing interactive workflows. You can set the workflow ID of a workflow as an argument to `DBOS.startWorkflow()`. Workflow IDs must be **globally unique** for your application. An assigned workflow ID acts as an idempotency key: if a workflow is called multiple times with the same ID, it executes only once. This is useful if your operations have side effects like making a payment or sending an email. Workflow IDs are also useful for communicating with workflows and developing interactive workflows - see Communicating with Workflows for more details. For example: ```javascript class Example { @DBOS.workflow() static async exampleWorkflow(var1: string, var2: string) { // ... } } async function main() { const myID: string = ... const handle = await DBOS.startWorkflow(Example, {workflowID: myID}).exampleWorkflow("one", "two"); const result = await handle.getResult(); } ``` By default, starting a workflow with an ID that is already in use returns a handle to the existing workflow instead of starting a new one. To instead throw `DBOSWorkflowIDInUseError`, pass `workflowIDReusePolicy: 'reject'` to `DBOS.startWorkflow` (or in `DBOSClient.enqueue` options), and match the error with `isWorkflowIDInUseError` from the SDK's `Error` namespace (`import { Error as DBOSErrors } from "@dbos-inc/dbos-sdk"`) rather than `instanceof`. ### Determinism Workflows are in most respects normal TypeScript functions. They can have loops, branches, conditionals, and so on. However, a workflow function must be **deterministic**: if called multiple times with the same inputs, it should invoke the same steps with the same inputs in the same order (given the same return values from those steps). If you need to perform a non-deterministic operation like accessing the database, calling a third-party API, generating a random number, or getting the local time, you shouldn't do it directly in a workflow function. Instead, you should do all database operations in transactions and all other non-deterministic operations in steps. For example, **don't do this**: ```javascript async function exampleWorkflowFunction() { const choice = Math.random() > 0.5 ? 1 : 0; if (choice === 0) { await stepOne(); } else { await stepTwo(); } } const exampleWorkflow = DBOS.registerWorkflow(exampleWorkflowFunction); ``` Do this instead: ```javascript async function exampleWorkflowFunction() { const choice = await DBOS.runStep( () => Promise.resolve(Math.random() > 0.5 ? 1 : 0), { name: "generateChoice" } ); if (choice === 0) { await stepOne(); } else { await stepTwo(); } } const exampleWorkflow = DBOS.registerWorkflow(exampleWorkflowFunction); ``` #### Running Steps In Parallel Initiating several concurrent steps in a workflow, followed by awaiting them with `Promise.allSettled`, is valid as long as the steps are started in a deterministic order. For example the following is allowed: ```typescript const results = await Promise.allSettled([ step1("arg1"), step2("arg2"), step3("arg3"), step4("arg4"), ]) ``` This is allowed because each step is started in a well-defined sequence before awaiting. By contrast, the following is not allowed: ```typescript const results = await Promise.allSettled([ (async () => { await step1("arg1"); await step2("arg3"); })(), (async () => { await step3("arg2"); await step4("arg4"); })(), ]); ``` Here, `step2` and `step4` may be started in either order since their execution depends on the relative time taken by `step1` and `step3`. If you need to run sequences of operations concurrently, start child workflows with `startWorkflow` and await the results from their `WorkflowHandle`s. Avoid using `Promise.all` because of how it handles errors and rejections. When any promise rejects, `Promise.all` immediately fails, leaving the other promises unresolved. If one of those later throws an unhandled exception, it can crash your Node.js process. Instead, prefer `Promise.allSettled`, which safely waits for all promises to complete and reports their outcomes. If DBOS detects that a single execution of a workflow recorded different results for the same step, it throws a `DBOSStepNondeterminismError`, indicating the workflow is not deterministic. ### Workflow Timeouts You can set a timeout for a workflow by passing a `timeoutMS` argument to `DBOS.startWorkflow`. When the timeout expires, the workflow **and all its children** are cancelled. Cancelling a workflow sets its status to `CANCELLED` and preempts its execution at the beginning of its next step. Timeouts are **start-to-completion**: a workflow's timeout does not begin until the workflow starts execution. Also, timeouts are **durable**: they are stored in the database and persist across restarts, so workflows can have very long timeouts. Example syntax: ```javascript async function taskFunction(task) { // ... } const taskWorkflow = DBOS.registerWorkflow(taskFunction); async function main() { const task = ... const timeout = ... // Timeout in milliseconds const handle = await DBOS.startWorkflow(taskWorkflow, {timeoutMS: timeout})(task); } ``` ### Workflow Attributes You can attach a record of custom, JSON-serializable key-value attributes to a workflow by passing `workflowAttributes` to `DBOS.startWorkflow`. This is useful for tagging workflows with application-specific metadata (such as a customer ID, tenant, or region) so you can find them later. Attributes must be a key-value object (not a scalar or array), are recorded at creation time, and are **not** inherited by child workflows. ```javascript const handle = await DBOS.startWorkflow(taskWorkflow, { workflowAttributes: { customer: "acme", region: "us-east-1" }, })(task); ``` Attributes are stored in Postgres as GIN-indexed JSONB, so you can efficiently search for workflows by attribute by passing the `attributes` filter to `DBOS.listWorkflows`. A workflow matches if its attributes contain all the key-value pairs you provide: ```javascript // Retrieve all workflows tagged with this customer const workflows = await DBOS.listWorkflows({ attributes: { customer: "acme" } }); ``` ### Durable Sleep You can use `DBOS.sleep()` to put your workflow to sleep for any period of time. This sleep is **durable**—DBOS saves the wakeup time in the database so that even if the workflow is interrupted and restarted multiple times while sleeping, it still wakes up on schedule. Sleeping is useful for scheduling a workflow to run in the future (even days, weeks, or months from now). For example: ```javascript @DBOS.workflow() static async exampleWorkflow(timeToSleep, task) { await DBOS.sleep(timeToSleep); await runTask(task); } ``` ### Debouncing Workflows You can create a `Debouncer` to debounce your workflows. Debouncing delays workflow execution until some time has passed since the workflow has last been called. This is useful for preventing wasted work when a workflow may be triggered multiple times in quick succession. For example, if a user is editing an input field, you can debounce their changes to execute a processing workflow only after they haven't edited the field for some time: #### Debouncer ```typescript new Debouncer( params: DebouncerConfig ) ``` ```typescript interface DebouncerConfig { workflow: (...args: Args) => Promise; startWorkflowParams?: StartWorkflowParams; debounceTimeoutMs?: number; } ``` **Parameters:** - **workflow**: The workflow to debounce. Note that workflows from configured instances cannot be debounced. - **startWorkflowParams**: Optional workflow parameters, as in `startWorkflow`. Applied to all workflows started from this debouncer. - **debounceTimeoutMs**: After this time elapses since the first time a workflow is submitted from this debouncer, the workflow is started regardless of the debounce period. #### debouncer.debounce ```typescript debouncer.debounce( debounceKey: string, debouncePeriodMs: number, ...args: Args ): Promise> ``` Submit a workflow for execution but delay it by `debouncePeriodMs`. Returns a handle to the workflow. The workflow may be debounced again, which further delays its execution (up to `debounceTimeoutMs`). When the workflow eventually executes, it uses the **last** set of inputs passed into `debounce`. Once the debounce period expires and the workflow is released for execution, the next call to `debounce` starts the debouncing process again for a new workflow execution. **Parameters:** - **debounceKey**: A key used to group workflow executions that will be debounced together. For example, if the debounce key is set to customer ID, each customer's workflows would be debounced separately. - **debouncePeriodMs**: Delay this workflow's execution by this period in milliseconds. - **...args**: Variadic workflow arguments. **Example Syntax**: ```typescript async function processInput(userInput: string) { ... } const processInputWorkflow = DBOS.registerWorkflow(processInput); // Each time a user submits a new input, debounce the processInput workflow. // The workflow will wait until 60 seconds after the user stops submitting new inputs, // then process the last input submitted. const debouncer = new Debouncer({ workflow: processInputWorkflow, }); async function onUserInputSubmit(userId: string, userInput: string) { const debounceKey = userId; const debouncePeriodMs = 60000; // 60 seconds await debouncer.debounce(debounceKey, debouncePeriodMs, userInput); } ``` ### Workflow Communication DBOS provides a few different ways to communicate with your workflows. You can: - Send messages to workflows - Publish events from workflows for clients to read - Stream values from workflows to clients ### Workflow Messaging and Notifications You can send messages to a specific workflow. This is useful for signaling a workflow or sending notifications to it while it's running. #### Send ```typescript DBOS.send(destinationID: string, message: T, topic?: string, idempotencyKey?: string): Promise; ``` You can call `DBOS.send()` to send a message to a workflow. Messages can optionally be associated with a topic and are queued on the receiver per topic. You can also call `send` from outside of your DBOS application with the DBOS Client. #### Recv ```typescript DBOS.recv(topic?: string, options?: RecvOptions): Promise interface RecvOptions { timeoutSeconds?: number; deadlineEpochMS?: number; pollingIntervalMs?: number; } ``` Workflows can call `DBOS.recv()` to receive messages sent to them, optionally for a particular topic. Each call to `recv()` waits for and consumes the next message to arrive in the queue for the specified topic, returning `null` if the wait times out. The default timeout is 60 seconds. If the topic is not specified, this method only receives messages sent without a topic. #### Messages Example Messages are especially useful for sending notifications to a workflow. For example, in the e-commerce demo, the checkout workflow, after redirecting customers to a secure payments service, must wait for a notification from that service that the payment has finished processing. To wait for this notification, the payments workflow uses `recv()`, executing failure-handling code if the notification doesn't arrive in time: ```javascript @DBOS.workflow() static async checkoutWorkflow(...): Promise { ... const notification = await DBOS.recv(PAYMENT_STATUS, timeout); if (notification) { ... // Handle the notification. } else { ... // Handle a timeout. } } ``` A webhook waits for the payment processor to send the notification, then uses `send()` to forward it to the workflow: ```javascript static async paymentWebhook(): Promise { const notificationMessage = ... // Parse the notification. const workflowID = ... // Retrieve the workflow ID from notification metadata. await DBOS.send(workflowID, notificationMessage, PAYMENT_STATUS); } ``` #### Reliability Guarantees All messages are persisted to the database, so if `send` completes successfully, the destination workflow is guaranteed to be able to `recv` it. If you're sending a message from a workflow, DBOS guarantees exactly-once delivery. If you're sending a message from normal TypeScript code, you can specify an idempotency key for `send` to guarantee exactly-once delivery. ### Workflow Events Workflows can publish _events_, which are key-value pairs associated with the workflow. They are useful for publishing information about the status of a workflow or to send a result to clients while the workflow is running. #### setEvent ```typescript DBOS.setEvent(key: string, value: T): Promise ``` Any workflow can call `DBOS.setEvent` to publish a key-value pair, or update its value if it has already been published. #### getEvent ```typescript DBOS.getEvent(workflowID: string, key: string, options?: GetEventOptions): Promise interface GetEventOptions { timeoutSeconds?: number; deadlineEpochMS?: number; pollingIntervalMs?: number; } ``` You can call `DBOS.getEvent` to retrieve the value published by a particular workflow ID for a particular key. If the event does not yet exist, this call waits for it to be published, returning `null` if the wait times out. Default timeout is 60 seconds. You can also call `getEvent` from outside of your DBOS application with DBOS Client. #### Events Example Events are especially useful for writing interactive workflows that communicate information to their caller. For example, in the e-commerce demo, the checkout workflow, after validating an order, directs the customer to a secure payments service to handle credit card processing. To communicate the payments URL to the customer, it uses events. The checkout workflow emits the payments URL using `setEvent()`: ```javascript @DBOS.workflow() static async checkoutWorkflow(...): Promise { ... const paymentsURL = ... await DBOS.setEvent(PAYMENT_URL, paymentsURL); ... } ``` The HTTP handler that originally started the workflow uses `getEvent()` to await this URL, then redirects the customer to it: ```javascript app.post("/checkout", async (req, res) => { const handle = await DBOS.startWorkflow(Shop).checkoutWorkflow(...); const url = await DBOS.getEvent(handle.workflowID, PAYMENT_URL); if (url === null) { res.redirect(`${origin}/checkout/cancel`); } else { res.redirect(url); } }); ``` #### Reliability Guarantees All events are persisted to the database, so the latest version of an event is always retrievable. Additionally, if `getEvent` is called in a workflow, the retrieved value is persisted in the database so workflow recovery can use that value, even if the event is later updated. ### Workflow Streaming Workflows can stream data in real time to clients. This is useful for streaming results from a long-running workflow or LLM call or for monitoring or progress reporting. #### Writing to Streams ```typescript DBOS.writeStream(key: string, value: T): Promise ``` You can write values to a stream from a workflow or its steps using `DBOS.writeStream`. A workflow may have any number of streams, each identified by a unique key. When you are done writing to a stream, you should close it with `DBOS.closeStream`, which you can call from a workflow or its steps. Otherwise, streams are automatically closed when the workflow terminates. ```typescript DBOS.closeStream(key: string): Promise ``` DBOS streams are immutable and append-only. Writes to a stream from a workflow happen exactly-once. Writes to a stream from a step happen at-least-once; if a step fails and is retried, it may write to the stream multiple times. Readers will see all values written to the stream from all tries of the step in the order in which they were written. **Example syntax:** ```typescript async function producerWorkflowFunction() { await DBOS.writeStream("example_key", { step: 1, data: "value1" }); await DBOS.writeStream("example_key", { step: 2, data: "value2" }); await DBOS.closeStream("example_key"); // Signal completion } const producerWorkflow = DBOS.registerWorkflow(producerWorkflowFunction); ``` #### Reading from Streams ```typescript DBOS.readStream( workflowID: string, key: string, options?: ReadStreamOptions ): AsyncGenerator interface ReadStreamOptions { offset?: number; pollingIntervalMs?: number; timeoutSeconds?: number; } ``` You can read values from a stream from anywhere using `DBOS.readStream`. This function reads values from a stream identified by a workflow ID and key, yielding each value in order until the stream is closed or the workflow terminates. You can also read from a stream from outside a DBOS application with a DBOS Client. **Parameters:** - `offset`: The offset to start reading from. Defaults to `0`, the start of the stream. - `pollingIntervalMs`: The interval, in milliseconds, between system database polls while waiting for new values. Must be at least `1`. - `timeoutSeconds`: How long to wait for **each** value before throwing `DBOSStreamTimeoutError`. The clock restarts every time a value is delivered, so this bounds the gap between values, not the total duration of the read. Defaults to waiting indefinitely. **Example syntax:** ```typescript for await (const value of DBOS.readStream(workflowID, "example_key")) { console.log(`Received: ${JSON.stringify(value)}`); } ``` ```typescript import { DBOS, Error as DBOSErrors } from "@dbos-inc/dbos-sdk"; try { for await (const value of DBOS.readStream(workflowID, "example_key", { timeoutSeconds: 30 })) { console.log(`Received: ${JSON.stringify(value)}`); } } catch (e) { if (DBOSErrors.isStreamTimeoutError(e)) { console.log("The producer stopped sending values"); } else { throw e; } } ``` Match a stream timeout with `isStreamTimeoutError` rather than `instanceof DBOSStreamTimeoutError`: when a workflow replays a checkpointed timeout, the error is revived as a plain `Error`, so `instanceof` does not hold. ### Steps When using DBOS workflows, you should call any function that performs complex operations or accesses external APIs or services as a _step_. If a workflow is interrupted, upon restart it automatically resumes execution from the **last completed step**. You can use `DBOS.runStep` to call a function as a step. For a function to be used as a step, it should have a return value that can be serialized as JSON, and should not have non-durable side effects. ALWAYS call steps this way unless otherwise specify. ```javascript async function generateRandomNumber() { return Math.random(); } async function workflowFunction() { const randomNumber = await DBOS.runStep(() => generateRandomNumber(), {name: "generateRandomNumber"}); } const workflow = DBOS.registerWorkflow(workflowFunction) ``` Alternatively, you can register a function as a step using `DBOS.registerStep`: NEVER do this unless specifically asked, ALWAYS use DBOS.runStep instead. ```javascript async function generateRandomNumber() { return Math.random(); } const randomStep = DBOS.registerStep(generateRandomNumber); async function workflowFunction() { const randomNumber = await randomStep(); } const workflow = DBOS.registerWorkflow(workflowFunction) ``` Or use the `@DBOS.step()` decorator: NEVER do this unless specifically asked, ALWAYS use DBOS.runStep instead. ```typescript export class Example { @DBOS.step() static async generateRandomNumber() { return Math.random(); } @DBOS.workflow() static async exampleWorkflow() { await Example.generateRandomNumber(); } } ``` You **cannot** call, start, or enqueue workflows from within steps. These operations should be performed from workflow functions. You can call one step from another step, but the called step becomes part of the calling step's execution rather than functioning as a separate step. Steps are only checkpointed when called from a workflow: a step called outside a workflow runs as an ordinary function call, with no checkpoint, retries, or timeout. Only workflows can be started or enqueued; calling `DBOS.startWorkflow` on a step throws an error. #### Configurable Retries You can optionally configure a step to automatically retry any exception a set number of times with exponential backoff. This is useful for automatically handling transient failures, like making requests to unreliable APIs. Retries are configurable through arguments to the step decorator: ```typescript export interface StepConfig { retriesAllowed?: boolean; // Should failures be retried? (default false) intervalSeconds?: number; // Seconds to wait before the first retry attempt (default 1). maxAttempts?: number; // Maximum number of attempts, including the first (default 3). If every attempt fails, throw an exception. backoffRate?: number; // Multiplier by which the retry interval increases after a retry attempt (default 2). timeoutMS?: number; // Maximum duration in milliseconds of a single attempt of the step. } ``` For example, let's configure this step to retry exceptions (such as if `example.com` is temporarily down) up to 10 times: ```javascript async function fetchFunction() { return await fetch("https://example.com").then(r => r.text()); } async function workflowFunction() { const randomNumber = await DBOS.runStep(() => fetchFunction(), { name: "fetchFunction", retriesAllowed: true, maxAttempts: 10 }); } ``` Or if registering the step: ```javascript async function fetchFunction() { return await fetch("https://example.com").then(r => r.text()); } const fetchStep = DBOS.registerStep(fetchFunction, { retriesAllowed: true, maxAttempts: 10 }); ``` Or if using decorators: ```javascript @DBOS.step({retriesAllowed: true, maxAttempts: 10}) static async exampleStep() { return await fetch("https://example.com").then(r => r.text()); } ``` If a step fails on all `maxAttempts` attempts, it throws a `DBOSMaxStepRetriesError` to the calling workflow. If that exception is not caught, the workflow terminates. #### Step Timeouts You can set a timeout on a step by passing `timeoutMS` to its `StepConfig`. If a single attempt of the step runs longer than the timeout, it fails with a `DBOSStepTimeoutError` (retried like any other failure if `retriesAllowed` is `true`). Step timeouts are **cooperative**: DBOS does not forcibly terminate a running step. When the timeout expires, DBOS aborts the `AbortSignal` exposed at `DBOS.stepStatus.timeoutSignal`; pass it to APIs like `fetch` so they stop work promptly. A step that ignores the signal keeps running in the background, but its result is discarded. ```javascript async function fetchFunction() { return await fetch("https://example.com", { signal: DBOS.stepStatus?.timeoutSignal }).then(r => r.text()); } async function workflowFunction() { return await DBOS.runStep(() => fetchFunction(), { name: "fetchFunction", timeoutMS: 5000 }); } ``` ### Queues You can use queues to run many workflows at once with managed concurrency. Queues provide _flow control_, letting you manage how many workflows run at once or how often workflows are started. Queue configuration is persisted to the system database, so any DBOS process or DBOSClient connected to the same system database can register, retrieve, and reconfigure queues. To create a queue, register it with `DBOS.registerQueue`. This must be called **after** `DBOS.launch()`: ```javascript import { DBOS } from "@dbos-inc/dbos-sdk"; await DBOS.registerQueue("example_queue"); ``` For brevity, some snippets below omit `DBOS.setConfig`, which you MUST still call before `DBOS.launch()`. Full signature: ```typescript DBOS.registerQueue( name: string, options?: RegisterQueueOptions, ): Promise interface RegisterQueueOptions { // Applied to the queue as a whole globalConcurrency?: number; workerConcurrency?: number; rateLimit?: { limitPerPeriod: number; periodSec: number }; // Applied to each partition separately partitionConcurrency?: number; partitionWorkerConcurrency?: number; partitionRateLimit?: { limitPerPeriod: number; periodSec: number }; minPollingIntervalMs?: number; onConflict?: 'update_if_latest_version' | 'always_update' | 'never_update'; } ``` Queue names must be unique within the system database, and names starting with `_dbos_` are reserved. Setting any partition limit makes the queue partitioned: every enqueue must supply a `queuePartitionKey`, and deduplication IDs are unique across the whole queue, including all its partitions. The queue-wide limits still apply across all partitions. If a queue with this name already exists in the database, `onConflict` controls whether the configuration is overwritten: - `'update_if_latest_version'` (default): only overwrite when the running application is the latest registered version. Safe for rolling deploys. - `'always_update'`: always overwrite. - `'never_update'`: leave the existing configuration unchanged. You can then enqueue any workflow by passing the queue name as an argument to `DBOS.startWorkflow`. Enqueuing a function submits it for execution and returns a handle to it. Queued tasks are started in first-in, first-out (FIFO) order. ```javascript class Tasks { @DBOS.workflow() static async processTask(task) { // ... } } async function main() { await DBOS.launch(); await DBOS.registerQueue("example_queue"); const task = ... const handle = await DBOS.startWorkflow(Tasks, { queueName: "example_queue" }).processTask(task) } ``` #### Reconfiguring Queues at Runtime Because queue configuration lives in the system database, you can change a queue's configuration at runtime without redeploying or restarting your workers. Use `DBOS.retrieveQueue` to fetch a queue, then call its `setX` methods. Workers pick up the new configuration on their next polling iteration. ```javascript const queue = await DBOS.retrieveQueue("example_queue"); if (queue !== null) { await queue.setGlobalConcurrency(20); await queue.setRateLimit({ limitPerPeriod: 25, periodSec: 30 }); } ``` Available methods on the returned `WorkflowQueue`: - `setGlobalConcurrency(value)`, `setWorkerConcurrency(value)`, `setRateLimit(value)`, `setPartitionConcurrency(value)`, `setPartitionWorkerConcurrency(value)`, `setPartitionRateLimit(value)`, `setMinPollingIntervalMs(value)` — write through to the database. - `getGlobalConcurrency()`, `getWorkerConcurrency()`, `getRateLimit()`, `getPartitionConcurrency()`, `getPartitionWorkerConcurrency()`, `getPartitionRateLimit()`, `getMinPollingIntervalMs()` — re-read from the database. To delete a queue, use `DBOS.deleteQueue("name")`. **Warning:** workflows already enqueued on a deleted queue can no longer be dequeued, executed, or recovered — unless a queue with the same name is later registered, in which case it will dequeue the leftover workflows. Do not rely on this behavior: cancel or drain pending workflows before deleting. #### Queue Example Here's an example of a workflow using a queue to process tasks in parallel: ```javascript import { DBOS } from "@dbos-inc/dbos-sdk"; async function taskFunction(task) { // ... } const taskWorkflow = DBOS.registerWorkflow(taskFunction, {"name": "taskWorkflow"}); async function queueFunction(tasks) { const handles = [] // Enqueue each task so all tasks are processed concurrently. for (const task of tasks) { handles.push(await DBOS.startWorkflow(taskWorkflow, { queueName: "example_queue" })(task)) } // Wait for each task to complete and retrieve its result. // Return the results of all tasks. const results = [] for (const h of handles) { results.push(await h.getResult()) } return results } const queueWorkflow = DBOS.registerWorkflow(queueFunction, {"name": "queueWorkflow"}) async function main() { await DBOS.launch(); await DBOS.registerQueue("example_queue"); // ... } ``` #### Enqueueing from Another Application Often, you want to enqueue a workflow from outside your DBOS application. For example, let's say you have an API server and a data processing service. You're using DBOS to build a durable data pipeline in the data processing service. When the API server receives a request, it should enqueue the data pipeline for execution on the data processing service. You can use the DBOS Client to register queues and enqueue workflows from outside your DBOS application by connecting directly to your DBOS application's system database. Since the DBOS Client is designed to be used from outside your DBOS application, workflow and queue metadata must be specified explicitly. For example, this code registers `pipelineQueue` and enqueues the `dataPipeline` workflow on it with `task` as an argument. ```ts import { DBOSClient } from "@dbos-inc/dbos-sdk"; const client = await DBOSClient.create({ systemDatabaseUrl: process.env.DBOS_SYSTEM_DATABASE_URL!, // The name of the application that runs the data pipeline applicationName: "data-processing-service", }); await client.registerQueue("pipelineQueue"); type ProcessTask = typeof Tasks.processTask; await client.enqueue( { workflowName: 'dataPipeline', queueName: 'pipelineQueue', }, task); ``` Note: `client.registerQueue` defaults `onConflict` to `'always_update'` because clients are not associated with an application version. `'update_if_latest_version'` is not supported on the client. #### Managing Concurrency You can control how many workflows from a queue run simultaneously by configuring concurrency limits. This helps prevent resource exhaustion when workflows consume significant memory or processing power. ##### Worker Concurrency Worker concurrency sets the maximum number of workflows from a queue that can run concurrently on a single DBOS process. This is particularly useful for resource-intensive workflows to avoid exhausting the resources of any process. For example, this queue has a worker concurrency of 5, so each process will run at most 5 workflows from this queue simultaneously: ```javascript import { DBOS } from "@dbos-inc/dbos-sdk"; await DBOS.registerQueue("example_queue", { workerConcurrency: 5 }); ``` ##### Global Concurrency Global concurrency limits the total number of workflows from a queue that can run concurrently across all DBOS processes in your application. For example, this queue will have a maximum of 10 workflows running simultaneously across your entire application. :::warning Worker concurrency limits are recommended for most use cases. Take care when using a global concurrency limit as any `PENDING` workflow on the queue counts toward the limit, including workflows from previous application versions ::: ```javascript import { DBOS } from "@dbos-inc/dbos-sdk"; await DBOS.registerQueue("example_queue", { globalConcurrency: 10 }); ``` ##### In-Order Processing You can use a queue with `globalConcurrency: 1` to guarantee sequential, in-order processing of events. Only a single event will be processed at a time. For example, this app processes events sequentially in the order of their arrival: ```javascript import { DBOS } from "@dbos-inc/dbos-sdk"; import express from "express"; const app = express(); class Tasks { @DBOS.workflow() static async processTask(task){ // ... process task } } app.get("/events/:event", async (req, res) => { await DBOS.startWorkflow(Tasks, { queueName: "in_order_queue" }).processTask(req.params); await res.send("Workflow Started!"); }); // Launch DBOS, register the queue, and start the server async function main() { DBOS.setConfig({ "name": "dbos-node-starter", "applicationVersion": "0.1.0", "systemDatabaseUrl": process.env.DBOS_SYSTEM_DATABASE_URL, }); await DBOS.launch(); await DBOS.registerQueue("in_order_queue", { globalConcurrency: 1 }); app.listen(3000, () => {}); } main().catch(console.log); ``` #### Rate Limiting You can set _rate limits_ for a queue, limiting the number of functions that it can start in a given period. Rate limits are global across all DBOS processes using this queue. For example, this queue has a limit of 50 with a period of 30 seconds, so it may not start more than 50 functions in 30 seconds: ```javascript await DBOS.registerQueue("example_queue", { rateLimit: { limitPerPeriod: 50, periodSec: 30 } }); ``` Rate limits are especially useful when working with a rate-limited API, such as many LLM APIs. #### Setting Timeouts You can set a timeout for an enqueued workflow by passing a `timeoutMS` argument to `DBOS.startWorkflow`. When the timeout expires, the workflow **and all its children** are cancelled. Cancelling a workflow sets its status to `CANCELLED` and preempts its execution at the beginning of its next step. Timeouts are **start-to-completion**: a workflow's timeout does not begin until the workflow is dequeued and starts execution. Also, timeouts are **durable**: they are stored in the database and persist across restarts, so workflows can have very long timeouts. Example syntax: ```javascript async function taskFunction(task) { // ... } const taskWorkflow = DBOS.registerWorkflow(taskFunction, {"name": "taskWorkflow"}); async function main() { await DBOS.launch(); await DBOS.registerQueue("example_queue"); const task = ... const timeout = ... // Timeout in milliseconds const handle = await DBOS.startWorkflow(taskWorkflow, { queueName: "example_queue", timeoutMS: timeout })(task); } ``` #### Partitioning Queues You can **partition** queues to distribute work across dynamically created queue partitions. A queue is partitioned if you register it with any per-partition flow control limit: - `partitionConcurrency`: Maximum workflows from any one partition running at once across all processes. - `partitionWorkerConcurrency`: Maximum workflows from any one partition running at once on a single process. - `partitionRateLimit`: Maximum workflows that may be started from any one partition in a given period. When you enqueue a workflow on a partitioned queue, you must supply a queue partition key. Essentially, you can think of each partition as a "subqueue" you dynamically create by enqueueing a workflow with a partition key. For example, suppose you want your users to each be able to run at most one task at a time. You can do this with a queue whose `partitionConcurrency` is 1, where the partition key is user ID. **Example Syntax** ```ts await DBOS.registerQueue("example_queue", { partitionConcurrency: 1 }); async function onUserTaskSubmission(userID: string, task: Task) { // Partition the task queue by user ID. As the queue has a // per-partition concurrency of 1, this means that at most one // task can run at once per user (but tasks from different // users can run concurrently). await DBOS.startWorkflow(taskWorkflow, { queueName: "example_queue", enqueueOptions: { queuePartitionKey: userID } })(task); } ``` A partitioned queue enforces its per-partition limits **and** its queue-wide limits (`globalConcurrency`, `workerConcurrency`, and `rateLimit`) at the same time. This lets you protect your workers from overload while still fairly distributing work between partitions. For example, this "fair queue" runs at most one task per user, but no more than 10 tasks on any single process: ```ts await DBOS.registerQueue("fair_queue", { partitionConcurrency: 1, workerConcurrency: 10 }); ``` Each queue-wide limit has a per-partition counterpart, so you can mix and match them freely: ```ts // At most 100 tasks running globally and 25 running per tenant, // at most 10 tasks running per process and 2 per tenant per process, // and at most 1000 tasks started per minute globally and 50 per tenant. await DBOS.registerQueue("tenant_queue", { globalConcurrency: 100, workerConcurrency: 10, rateLimit: { limitPerPeriod: 1000, periodSec: 60 }, partitionConcurrency: 25, partitionWorkerConcurrency: 2, partitionRateLimit: { limitPerPeriod: 50, periodSec: 60 }, }); ``` Each per-partition concurrency limit must be less than or equal to its queue-wide counterpart, and `partitionWorkerConcurrency` must be less than or equal to `partitionConcurrency`. On a partitioned queue, deduplication IDs are unique across the whole queue, not per partition. To deduplicate within each partition separately, include the partition key in the deduplication ID. #### Deduplication You can set a deduplication ID for an enqueued workflow as an argument to `DBOS.startWorkflow`. At any given time, only one workflow with a specific deduplication ID can be enqueued in the specified queue. If a workflow with a deduplication ID is currently enqueued, delayed, or actively executing (status `ENQUEUED`, `DELAYED`, or `PENDING`), subsequent workflow enqueue attempt with the same deduplication ID in the same queue will raise a `DBOSQueueDuplicatedError` exception. For example, this is useful if you only want to have one workflow active at a time per user—set the deduplication ID to the user's ID. Example syntax: ```javascript async function taskFunction(task) { // ... } const taskWorkflow = DBOS.registerWorkflow(taskFunction, {"name": "taskWorkflow"}); async function main() { await DBOS.launch(); await DBOS.registerQueue("example_queue"); const task = ... const dedup: string = ... try { const handle = await DBOS.startWorkflow(taskWorkflow, { queueName: "example_queue", enqueueOptions: { deduplicationID: dedup } })(task); } catch (e) { // Handle DBOSQueueDuplicatedError } } ``` #### Priority You can set a priority for an enqueued workflow as an argument to `DBOS.startWorkflow`. Workflows with the same priority are dequeued in **FIFO (first in, first out)** order. Priority values can range from `0` to `2,147,483,647`, where **a low number indicates a higher priority**. Priority is enabled on every queue; no extra configuration is needed. :::tip Workflows without assigned priorities have priority `0`, the highest priority. ::: Example syntax: ```javascript async function taskFunction(task) { // ... } const taskWorkflow = DBOS.registerWorkflow(taskFunction, {"name": "taskWorkflow"}); async function main() { await DBOS.launch(); await DBOS.registerQueue("example_queue"); const task = ... const priority: number = ... const handle = await DBOS.startWorkflow(taskWorkflow, { queueName: "example_queue", enqueueOptions: { priority: priority } })(task); } ``` #### Explicit Queue Listening By default, a process running DBOS listens to (dequeues workflows from) all queues owned by its application in its system database. However, sometimes you only want a process to listen to a specific list of queues. You can configure `listenQueues` in your DBOS configuration to explicitly tell a process running DBOS to only listen to a specific set of queues. Each entry is a queue name; names that don't match any queue at launch are deferred until a queue is registered with that name. This is particularly useful when managing heterogeneous workers, where specific tasks should execute on specific physical servers. For example, say you have a mix of CPU workers and GPU workers and you want CPU tasks to only execute on CPU workers and GPU tasks to only execute on GPU workers. You can create separate queues for CPU and GPU tasks and configure each type of worker to only listen to the appropriate queue: ```javascript import { DBOS } from "@dbos-inc/dbos-sdk"; async function main() { const workerType = process.env.WORKER_TYPE; // "cpu" or "gpu" const config = // ... if (workerType === "gpu") { // GPU workers will only dequeue and execute workflows from the GPU queue config.listenQueues = ["gpu_queue"]; } else if (workerType === "cpu") { // CPU workers will only dequeue and execute workflows from the CPU queue config.listenQueues = ["cpu_queue"]; } DBOS.setConfig(config); await DBOS.launch(); await DBOS.registerQueue("cpu_queue"); await DBOS.registerQueue("gpu_queue"); } ``` Note that `listenQueues` only controls what workflows are dequeued, not what workflows can be enqueued, so you can freely enqueue tasks onto the GPU queue from a CPU worker for execution on a GPU worker, and vice versa. ### Classes You can use class instance methods as workflows and steps. Any class instance method can be freely used as a step using DBOS.runStep; there are no special requirements. To use the DBOS.step decorator on a class instance method, the class must inherit from `ConfiguredInstance`. To use a class instance method as a workflow, the class must inherit from `ConfiguredInstance`, and you must either use the DBOS.workflow decorator or register the method with DBOS.registerWorkflow, passing the class's prototype (`MyClass.prototype`) as `ctorOrProto`. For example: ```typescript class MyClass extends ConfiguredInstance { cfg: MyConfig; constructor(name: string, config: MyConfig) { super(name); this.cfg = config; } override async initialize() : Promise { // ... Validate this.cfg; will be called at DBOS.launch() } @DBOS.workflow() async testWorkflow(p: string): Promise { // ... Operations that use this.cfg } } const myClassInstance = new MyClass('instanceA', myConfig); ``` When you create a new instance of such a class, the constructor for the base `ConfiguredInstance` must be called with a `name`. This `name` should be unique among instances of the same class. Additionally, all `ConfiguredInstance` classes must be instantiated before DBOS.launch() is called. The reason for these requirements is to enable workflow recovery. When you create a new instance of a `ConfiguredInstance` class, DBOS stores it in a global registry indexed by `name`. When DBOS needs to recover a workflow belonging to that class, it looks up the `name` so it can run the workflow using the right class instance. While names are used by DBOS Transact internally to find the correct object instance across system restarts, they are also potentially useful for monitoring, tracing, and debugging. You should AVOID using ConfiguredInstance if at all possible and instead use registerWorkflow on regular non-class functions. #### Testing DBOS Functions Unless otherwise specified, recommend testing DBOS applications with jest. There are two approaches to testing workflows: 1. **Unit testing.** Test workflow logic in isolation by mocking the DBOS interface. Does not require Postgres. 2. **Integration testing.** Test workflows with real DBOS infrastructure. Requires a Postgres test database. ##### Unit Testing You can unit test workflows by mocking the DBOS interface with `jest.mock`. It's important to mock `DBOS.registerWorkflow` to directly return the workflow function instead of wrapping it with durable workflow code: ```ts // Mock DBOS jest.mock('@dbos-inc/dbos-sdk', () => ({ DBOS: { // IMPORTANT: Mock DBOS.registerWorkflow to return the workflow function registerWorkflow: jest.fn((fn) => fn), setEvent: jest.fn(), recv: jest.fn(), startWorkflow: jest.fn(), workflowID: 'test-workflow-id-123', }, })); ``` Then call the workflow function directly in your tests and assert on the mocked calls. ##### Integration Testing Integration tests require a Postgres database. You should reset DBOS and its system database between tests: ```ts describe('example integration tests', () => { beforeEach(async () => { const databaseUrl = process.env.DBOS_TEST_DATABASE_URL; // Shut down DBOS (in case a previous test launched it). await DBOS.shutdown(); // Reset the system database here (for example, drop and recreate it). // Configure and launch DBOS const dbosTestConfig: DBOSConfig = { name: "my-integration-test", applicationVersion: "0.1.0", systemDatabaseUrl: databaseUrl, }; DBOS.setConfig(dbosTestConfig); await DBOS.launch(); }, 10000); afterEach(async () => { await DBOS.shutdown(); }); it('my integration test', async () => { // test goes here }); }); ``` #### Logging ALWAYS log errors like this: ```typescript DBOS.logger.error(`Error: ${(error as Error).message}`); ``` ### Workflow Handles A workflow handle represents the state of a particular active or completed workflow execution. You obtain a workflow handle when using `DBOS.startWorkflow` to start a workflow in the background. If you know a workflow's identity, you can also retrieve its handle using `DBOS.retrieveWorkflow`. Workflow handles have the following methods: #### handle.workflowID ```typescript handle.workflowID: string; ``` Retrieve the ID of the workflow. #### handle.getResult ```typescript handle.getResult( options?: { pollingIntervalMs?: number } ): Promise; ``` Wait for the workflow to complete, then return its result. The optional `pollingIntervalMs` sets the interval between system database polls while waiting. It only applies to handles that wait by polling the database (such as handles from `DBOS.retrieveWorkflow`, from `DBOS.startWorkflow` with a `queueName`, or from the DBOS Client), not to a handle for a workflow that `DBOS.startWorkflow` runs directly in the same process. #### handle.getStatus ```typescript handle.getStatus(): Promise; ``` Retrieve the WorkflowStatus of the workflow, or `null` if not found: #### Workflow Status Some workflow introspection and management methods return a `WorkflowStatus`. This object has the following definition: ```typescript export interface WorkflowStatus { // The workflow ID readonly workflowID: string; // The status of the workflow. One of PENDING, SUCCESS, ERROR, ENQUEUED, DELAYED, CANCELLED, or MAX_RECOVERY_ATTEMPTS_EXCEEDED. readonly status: string; // The name of the workflow function. readonly workflowName: string; // The name of the workflow's class, if any. readonly workflowClassName: string; // The name with which the workflow's class instance was configured, if any. readonly workflowConfigName?: string; // If the workflow was enqueued, the name of the queue. readonly queueName?: string; // The deserialized workflow inputs. readonly input?: unknown[]; // The workflow's deserialized output, if any. readonly output?: unknown; // The error thrown by the workflow, if any. readonly error?: unknown; // The ID of the executor (process) that most recently executed this workflow. readonly executorId?: string; // The application version on which this workflow started. readonly applicationVersion?: string; // Workflow start time, as a UNIX epoch timestamp in milliseconds readonly createdAt: number; // Last time the workflow status was updated, as a UNIX epoch timestamp in milliseconds. For a completed workflow, this is the workflow completion timestamp. readonly updatedAt?: number; // The timeout specified for this workflow, if any. Timeouts are start-to-close. readonly timeoutMS?: number; // The deadline at which this workflow times out, if any. Not set until the workflow begins execution. readonly deadlineEpochMS?: number; // Unique queue deduplication ID, if any. Deduplication IDs are unset when the workflow completes. readonly deduplicationID?: string; // Priority of the workflow on a queue, 0 ~ 2,147,483,647. Default 0 (highest priority). readonly priority: number; // If this workflow is enqueued on a partitioned queue, its partition key readonly queuePartitionKey?: string; // If this workflow was forked from another, that workflow's ID. readonly forkedFrom?: string; // Whether this workflow has ever been forked from by another workflow. readonly wasForkedFrom?: boolean; // Custom key-value attributes attached to the workflow at creation, if any. readonly attributes?: Record; // If this workflow was enqueued by a named schedule, that schedule's name. readonly scheduleName?: string; } ``` ### DBOS Variables #### DBOS.workflowID ```typescript DBOS.workflowID: string | undefined; ``` Return the ID of the current workflow, if in a workflow. #### DBOS.stepID ```typescript DBOS.stepID: number | undefined; ``` Return the unique ID of the current step within a workflow. #### DBOS.stepStatus ```typescript DBOS.stepStatus: StepStatus | undefined; ``` Return the status of the currently executing step. This object has the following properties: ```typescript interface StepStatus { // The unique ID of this step in its workflow. stepID: number; // For steps with automatic retries, which attempt number (starting from 1) is currently executing. currentAttempt?: number; // For steps with automatic retries, the maximum number of attempts that will be made before the step fails. maxAttempts?: number; // For steps with a timeout, an AbortSignal that fires when the current attempt's timeout expires. timeoutSignal?: AbortSignal; // An AbortSignal that fires (within about a second) when the step's workflow is cancelled. cancelSignal: AbortSignal; } ``` Pass `cancelSignal` to APIs like `fetch` so a step stops promptly when its workflow is cancelled; otherwise, a running step completes before cancellation takes effect at the next step. #### DBOS.applicationVersion ```typescript DBOS.applicationVersion: string ``` Return the current application version. #### DBOS.executorID ```typescript DBOS.executorID: string ``` Retrieve the current executor ID, a unique process ID used to identify the application instance in distributed environments. ### Workflow Management Methods #### DBOS.listWorkflows ```typescript DBOS.listWorkflows( input: GetWorkflowsInput ): Promise ``` ```typescript interface GetWorkflowsInput { workflowIDs?: string[]; // Retrieve workflows with these IDs. workflowName?: string; // Retrieve workflows with this name. status?: WorkflowStatusString | WorkflowStatusString[]; // Retrieve workflows with this status or any of these statuses (each must be `ENQUEUED`, `DELAYED`, `PENDING`, `SUCCESS`, `ERROR`, `CANCELLED`, or `MAX_RECOVERY_ATTEMPTS_EXCEEDED`) startTime?: string; // Retrieve workflows started after this (RFC 3339-compliant) timestamp. endTime?: string; // Retrieve workflows started before this (RFC 3339-compliant) timestamp. authenticatedUser?: string; // Retrieve workflows run by this authenticated user. applicationVersion?: string; // Retrieve workflows started on this application version. executorId?: string; // Retrieve workflows run by this executor ID. workflow_id_prefix?: string; // Retrieve workflows whose ID have this prefix queueName?: string; // If this workflow is enqueued, on which queue queuesOnly?: boolean; // Return only workflows that are actively enqueued forkedFrom?: string; // Get workflows forked from this workflow ID. hasParent?: boolean; // If true, only return workflows that have a parent. If false, only return workflows without a parent. attributes?: Record; // Retrieve workflows whose custom attributes contain all of these key-value pairs. scheduleName?: string; // Retrieve workflows enqueued by this named schedule. limit?: number; // Return up to this many workflows IDs. IDs are ordered by workflow creation time. offset?: number; // Skip this many workflows IDs. IDs are ordered by workflow creation time. sortDesc?: boolean; // Sort the workflows in descending order by creation time (default ascending order). loadInput?: boolean; // Load the input of the workflow (default true). loadOutput?: boolean; // Load the output of the workflow (default true). } ``` Retrieve a list of WorkflowStatus of all workflows matching specified criteria. #### DBOS.listQueuedWorkflows ```typescript DBOS.listQueuedWorkflows( input: GetWorkflowsInput ): Promise ``` Retrieve a list of WorkflowStatus of all **currently enqueued** (status `PENDING`, `ENQUEUED`, or `DELAYED`) workflows matching specified criteria. The input type is the same as `DBOS.listWorkflows`; this method is equivalent to calling `DBOS.listWorkflows` with `queuesOnly` set and `loadOutput` set to `false`. #### DBOS.listWorkflowSteps ```typescript DBOS.listWorkflowSteps( workflowID: string, options?: ListWorkflowStepsOptions ): Promise interface ListWorkflowStepsOptions { limit?: number; offset?: number; } ``` Retrieve the steps of a workflow. Returns `undefined` if the workflow is not found. Steps are ordered by `functionID`. Use `limit` and `offset` to paginate results. This is a list of `StepInfo` objects, with the following structure: ```typescript interface StepInfo { // The unique ID of the step in the workflow. Zero-indexed. readonly functionID: number; // The name of the step readonly name: string; // The step's output, if any readonly output: unknown; // The error the step threw, if any readonly error: Error | null; // If the step starts or retrieves the result of a workflow, its ID readonly childWorkflowID: string | null; // The Unix epoch timestamp at which this step started readonly startedAtEpochMs?: number; // The Unix epoch timestamp at which this step completed readonly completedAtEpochMs?: number; } ``` #### DBOS.setWorkflowPriority ```typescript DBOS.setWorkflowPriority( workflowID: string, priority: number ): Promise ``` Set the priority of a queued workflow. Only affects workflows with `ENQUEUED` or `DELAYED` status. Priority value must be between `0` and `2,147,483,647`. Lower values are dequeued first. Throws `DBOSInvalidQueuePriorityError` if the priority is out of range. #### DBOS.setWorkflowDelay ```typescript DBOS.setWorkflowDelay( workflowID: string, options: SetWorkflowDelayOptions ): Promise interface SetWorkflowDelayOptions { delaySeconds?: number; delayUntilEpochMS?: number; } ``` Set or update the delay on a workflow. Only affects workflows with `DELAYED` status. Accepts a `SetWorkflowDelayOptions` object with `delaySeconds` (relative) or `delayUntilEpochMS` (absolute). #### DBOS.cancelWorkflow ```typescript cancelWorkflow( workflowID: string, options?: { cancelChildren?: boolean } ): Promise ``` Cancel a workflow. This sets its status to `CANCELLED`, removes it from its queue (if it is enqueued) and preempts its execution (interrupting it at the beginning of its next step) If `cancelChildren` is true, also cancel all child workflows recursively. #### DBOS.resumeWorkflow ```typescript DBOS.resumeWorkflow( workflowID: string, options?: { queueName?: string } ): Promise>> ``` Resume a workflow. This immediately starts it from its last completed step. You can use this to resume workflows that are cancelled or have exceeded their maximum recovery attempts. You can also use this to start an enqueued workflow immediately, bypassing its queue. If `queueName` is provided, the resumed workflow is enqueued on the specified queue instead of starting immediately. Throws `DBOSNonExistentWorkflowError` if the workflow does not exist. #### DBOS.forkWorkflow ```typescript static async forkWorkflow( workflowID: string, startStep: number, options?: { newWorkflowID?: string; applicationVersion?: string; timeoutMS?: number; queueName?: string; queuePartitionKey?: string; }, ): Promise>> ``` Start a new execution of a workflow from a specific step. The input step ID (`startStep`) must match the `functionID` of the step returned by `listWorkflowSteps`. The specified `startStep` is the step from which the new workflow will start, so any steps whose ID is less than `startStep` will not be re-executed. **Parameters:** - **workflowID**: The ID of the workflow to fork. - **startStep**: The ID of the step from which to start the forked workflow. Must match the `functionID` of the step in the original workflow execution. - **newWorkflowID**: The ID of the new workflow created by the fork. If not specified, a random UUID is used. - **applicationVersion**: The application version on which the forked workflow will run. Useful for "patching" workflows that failed due to a bug in the previous application version. - **timeoutMS**: A timeout for the forked workflow in milliseconds. - **queueName**: If provided, the forked workflow is enqueued on the specified queue instead of starting immediately. - **queuePartitionKey**: If the queue is partitioned, the partition key for the forked workflow. #### DBOS.rewindWorkflow ```typescript static async rewindWorkflow( workflowID: string, options?: { startStep?: number; applicationVersion?: string; queueName?: string; queuePartitionKey?: string; }, ): Promise>> ``` Rewind a workflow to a specific step and re-execute it from that step, keeping its workflow ID (unlike `forkWorkflow`, which creates a new workflow). Steps with IDs greater than or equal to `startStep` are discarded and re-executed; `startStep` defaults to `0`, re-executing the whole workflow. Only a workflow in a terminal state (`SUCCESS`, `ERROR`, `CANCELLED`, or `MAX_RECOVERY_ATTEMPTS_EXCEEDED`) can be rewound; cancel a running workflow first. Rewinding clears the workflow's output, rolls back events it set at or after `startStep`, deletes messages it received at or after `startStep`, and deletes the checkpoints of data source transactions at or after `startStep`. `applicationVersion`, `queueName`, and `queuePartitionKey` behave as in `forkWorkflow`. ### Upgrading Workflow Code A challenge encountered when operating long-running durable workflows in production is **how to deploy breaking changes without disrupting in-progress workflows.** A breaking change to a workflow is one that changes which steps are run, or the order in which the steps are run. If a breaking change was made to a workflow and that workflow is replayed by the recovery system, the checkpoints created by the previous version of the code may not match the steps called by the workflow in the new version of the code, causing recovery to fail. DBOS supports two strategies for safely upgrading workflow code: **patching** and **versioning**. #### Patching In patching, the result of a call to `DBOS.patch()` is used to conditionally execute the new code. `DBOS.patch()` returns `true` for new calls (those executing after the breaking change) and `false` for old calls (those that executed before the breaking change). Therefore, if `DBOS.patch()` returns `true`, the workflow should follow the new code path, otherwise it must follow the prior codepath. To use patching, you MUST enable it in the configuration; otherwise `DBOS.patch()` and `DBOS.deprecatePatch()` throw an error: ```typescript const config: DBOSConfig = { // ... enablePatching: true, } ``` ```typescript DBOS.patch( patchName: string ): Promise ``` **Parameters:** - `patchName`: The name to give the patch marker that will be inserted into workflow history. For example, let's say our original workflow is: ```typescript @DBOS.workflow() static async workflow() { await foo(); await bar(); } ``` We want to replace the call to `foo()` with a call to `baz()`. This is a breaking change because it changes what steps run. We can make this breaking change safely using a patch: ```typescript @DBOS.workflow() static async workflow() { if (await DBOS.patch('use-baz')) { await baz(); } else { await foo(); } await bar(); } ``` Now, new workflows will run `baz()`, while old workflows will reexecute `foo()`. #### Deprecating and Removing Patches Patches add complexity and runtime overhead; fortunately they don't need to stay in your code forever. Once all workflows that started before you deployed the patch are complete, you can safely remove patches from your code. First, you must deprecate the patch with `DBOS.deprecatePatch()`. `DBOS.deprecatePatch` must be used for a transition period prior to fully removing the patch, as it allows coexistence with any ongoing workflows that used `DBOS.patch()`. ```typescript DBOS.deprecatePatch( patchName: string ): Promise ``` Always returns `true`. Safely bypasses a patch marker at the current point in workflow history if present. **Parameters:** - `patchName`: The name of the patch marker to be bypassed. For example, here's how to deprecate the patch above: ```typescript @DBOS.workflow() static async workflow(){ if (await DBOS.deprecatePatch("use-baz")) { // always true await baz(); } await bar(); } ``` Then, when all workflows that started before you deprecated the patch are complete, you can remove the patch entirely: ```typescript @DBOS.workflow() static async workflow() { await baz() await bar() } ``` If any mistakes happen during the process (a breaking change is not patched, or a patch is deprecated or removed prematurely), the workflow will throw a `DBOSUnexpectedStepError` pointing to the step where the problem occurred. #### Versioning When using versioning, DBOS **versions** applications and workflows, and only continues workflow execution with the same application version that started the workflow. All workflows are tagged with the application version on which they started. By default, application version is automatically computed from a hash of workflow source code (or is fixed to `PATCHING_ENABLED` if patching is enabled). However, you can set your own version through configuration. ```typescript const config: DBOSConfig = { // ... applicationVersion: '1.0.0', } ``` When DBOS tries to recover workflows, it only recovers workflows whose version matches the current application version. This prevents recovery of workflows that depend on different code. You cannot change the version of a workflow, but you can use `DBOS.forkWorkflow` to restart a workflow from a specific step on a specific code version. ### Configuring DBOS To configure DBOS, pass in a configuration with `DBOS.setConfig` before you call `DBOS.launch`. For example: ```javascript DBOS.setConfig({ name: 'my-app', applicationVersion: '0.1.0', systemDatabaseUrl: process.env.DBOS_SYSTEM_DATABASE_URL, }); await DBOS.launch(); ``` A configuration object has the following fields. All fields except `name` are optional. ```javascript export interface DBOSConfig { name: string; applicationVersion?: string; executorID?: string; systemDatabaseUrl?: string; systemDatabasePoolSize?: number; systemDatabasePollingConcurrency?: number; systemDatabaseSchemaName?: string; systemDatabasePool?: Pool; observabilityQueryTimeoutMs?: number; systemDatabaseIdleTransactionTimeoutMs?: number; enableOTLP?: boolean; logLevel?: string; logger?: DLogger; otlpLogsEndpoints?: string[]; otlpTracesEndpoints?: string[]; listenQueues?: string[]; maxConcurrentQueueDispatches?: number; enablePatching?: boolean; serializer?: DBOSSerializer; } ``` - **name**: Your application's name. - **applicationVersion**: The code version for this application and its workflows. - **executorID**: A unique process ID used to identify the application instance in distributed environments. - **systemDatabaseUrl**: A connection string to a Postgres database in which DBOS can store internal state. The supported format is: ``` postgresql://[username]:[password]@[hostname]:[port]/[database name] ``` The default is: ``` postgresql://postgres:dbos@localhost:5432/[application name]_dbos_sys ``` If the Postgres database referenced by this connection string does not exist, DBOS will attempt to create it. - **systemDatabasePoolSize**: The size of the connection pool used for the DBOS system database. Defaults to 10. - **systemDatabasePollingConcurrency**: The maximum number of database-backed polling reads from wait operations (such as `getResult`, `waitAll`, `waitFirst`, `recv`, and `getEvent`) that may run concurrently against the system database pool. This prevents high-fan-out polling from starving control-plane operations such as enqueue/dequeue, status writes, recovery, and cancellation. Defaults to half the `systemDatabasePoolSize` (minimum 1). Set to a non-positive value to disable the limit. - **systemDatabaseSchemaName**: Postgres schema name for DBOS system tables. Defaults to `dbos`. - **systemDatabasePool**: A custom `node-postgres` connection pool to use to connect to your system database. If provided, DBOS will not create a connection pool but use this instead. - **observabilityQueryTimeoutMs**: The statement timeout, in milliseconds, applied to observability queries against the system database (such as `DBOS.listWorkflows`, `DBOS.listQueuedWorkflows`, and `DBOS.listWorkflowSteps`), so a slow query on a large database does not hold resources indefinitely. A query that exceeds the timeout throws `DBOSQueryTimeoutError`. Defaults to 30000 (30 seconds). Set to zero or a negative value to disable the timeout. - **systemDatabaseIdleTransactionTimeoutMs**: The Postgres `idle_in_transaction_session_timeout`, in milliseconds, set on the system database connections DBOS creates. Defaults to 60000 (60 seconds). - **enableOTLP**: Enable DBOS OpenTelemetry tracing and export. Defaults to False (True in DBOS Cloud). - **logLevel**: Configure the DBOS logger severity. Defaults to `info`. - **logger**: A custom logger implementing the `DLogger` interface, to which DBOS directs all its internal logging, replacing the built-in console and OTLP log sinks. When set, `logLevel` does not filter calls to it (level routing is the logger's job), logs are not exported over OTLP even if `enableOTLP` is on (traces are unaffected), and DBOS never flushes or closes it (the caller owns its lifecycle). `error()` receives the message of an `Error`, with its stack trace in `metadata.stack` and the original `Error` object in `metadata.error`. - **otlpTracesEndpoints**: A list of OTLP-compatible receivers to which to send traces. Only used when `enableOTLP` is enabled. - **otlpLogsEndpoints**: A list of OTLP-compatible receivers to which to send logs. Only used when `enableOTLP` is enabled. - **listenQueues**: This process should only listen to (dequeue and execute workflows from) these queues. Each entry is a queue name. Names that do not match any queue at launch are deferred — a queue registered later under that name will be picked up automatically. - **maxConcurrentQueueDispatches**: The maximum number of queues this process may dequeue from concurrently. Defaults to 3. Must be a positive integer; set to 1 to dequeue from one queue at a time. This prevents dequeuing from a large queue (especially a partitioned queue with many active partitions) from delaying work on smaller queues. A single queue is never dequeued from twice concurrently in the same process. Does not affect workflow concurrency, rate limits, or `systemDatabasePollingConcurrency`. - **enablePatching**: Enable workflow patching with `DBOS.patch()` and `DBOS.deprecatePatch()`. Defaults to false. - **serializer**: A custom serializer for the system database. Must match the `DBOSSerializer` interface with `name`, `stringify`, and `parse` methods. ````
--- ## DBOS CLI(3) ### Workflow Management Commands These commands all require the URL of your DBOS system database. You can supply this URL through the `--sys-db-url` argument or through a [`dbos-config.yaml` configuration file](./configuration.md#dbos-configuration-file). #### npx dbos workflow list **Description:** List workflows run by your application in JSON format ordered by recency (most recently started workflows last). **Arguments:** - `-s, --sys-db-url `: Your DBOS system database URL - `-n, --name ` Retrieve functions with this name - `-l, --limit ` Limit the results returned (default: "10") - `-u, --user ` Retrieve workflows run by this user - `-t, --start-time ` Retrieve workflows starting after this timestamp (ISO 8601 format) - `-e, --end-time ` Retrieve workflows starting before this timestamp (ISO 8601 format) - `-S, --status ` Retrieve workflows with this status (`PENDING`, `SUCCESS`, `ERROR`, `MAX_RECOVERY_ATTEMPTS_EXCEEDED`, `ENQUEUED`, `DELAYED`, or `CANCELLED`) - `-v, --application-version ` Retrieve workflows with this application version - `-a, --application-name ` Retrieve workflows owned by this application (workflows owned by no application are always included) **Output:** A JSON-formatted list of [workflow statuses](./methods.md#workflow-status). The `input`, `output`, and `error` fields are rendered as human-readable strings rather than as JSON values. #### npx dbos workflow get **Description:** Retrieve information on a workflow run by your application. **Arguments:** - `-s, --sys-db-url `: Your DBOS system database URL - ``: The ID of the workflow to retrieve. **Output:** A JSON-formatted [workflow status](./methods.md#workflow-status). The `input`, `output`, and `error` fields are rendered as human-readable strings rather than as JSON values. #### npx dbos workflow steps **Arguments:** - `-s, --sys-db-url `: Your DBOS system database URL - ``: The ID of the workflow to retrieve **Output:** A JSON-formatted list of [workflow steps](./methods.md#dboslistworkflowsteps). The `output` and `error` fields are rendered as human-readable strings rather than as JSON values. #### npx dbos workflow cancel **Description:** Cancel a workflow so it is no longer automatically retried or restarted. If the workflow is executing, it is interrupted at the beginning of its next step. **Arguments:** - `-s, --sys-db-url `: Your DBOS system database URL - ``: The ID of the workflow to cancel. #### npx dbos workflow resume **Description:** Resume a workflow from its last completed step. You can use this to resume workflows that are cancelled or that have exceeded their maximum recovery attempts. You can also use this to start an `ENQUEUED` workflow, bypassing its queue. **Arguments:** - `-s, --sys-db-url `: Your DBOS system database URL - ``: The ID of the workflow to resume. #### npx dbos workflow fork **Description:** Fork a new execution of a workflow, starting at a given step. This new workflow has a new workflow ID but the same code version, unless you specify a different one with `--application-version`. Forking from step N copies the results of all previous steps to the new workflow, which then starts running from step N. **Arguments:** * ``: The ID of the workflow to fork. - `-s, --sys-db-url URL`: Your DBOS system database URL. * `-f, --forked-workflow-id`: Custom ID for the forked workflow * `-v, --application-version`: Custom application version for the forked workflow * `-S, --step INTEGER`: Restart from this step (required) #### npx dbos workflow queue list **Description:** Lists all currently enqueued workflows in JSON format ordered by recency (most recently enqueued workflows last). **Arguments:** - `-s, --sys-db-url `: Your DBOS system database URL - `-n, --name ` Retrieve functions with this name - `-t, --start-time ` Retrieve functions starting after this timestamp (ISO 8601 format) - `-e, --end-time ` Retrieve functions starting before this timestamp (ISO 8601 format) - `-S, --status ` Retrieve functions with this status (`ENQUEUED`, `PENDING`, or `DELAYED`) - `-l, --limit ` Limit the results returned - `-q, --queue ` Retrieve functions run on this queue - `-a, --application-name ` Retrieve functions owned by this application (functions owned by no application are always included) **Output:** A JSON-formatted list of [workflow statuses](./methods.md#workflow-status). The `input`, `output`, and `error` fields are rendered as human-readable strings rather than as JSON values. ### Application Management Commands #### npx dbos schema **Description:** Create the DBOS system database and internal tables. By default, a DBOS application automatically creates these on startup. However, in production environments, a DBOS application may not run with sufficient privilege to create databases or tables. In that case, this command can be run with a privileged user to create all DBOS database tables. After creating the DBOS database tables with this command, a DBOS application can run with minimum permissions, requiring only access to the DBOS schema in the system database. Use the `-r` flag to grant a role access to that schema. Such an application should also be configured with [`runMigrations: false`](./configuration.md#database-connection-settings), so it never attempts to alter the schema and instead verifies at launch that this command has brought the system database up to date. You can also run these migrations from code with [`DBOS.migrate`](./dbos-class.md#dbosmigrate). This command does not migrate the tables used by [datasources](./datasource.md); use your datasource's [`initializeDBOSSchema`](./datasource.md#installing-the-dbos-schema) method for those. **Arguments:** - `systemDatabaseUrl`: A connection string for your DBOS [system database](../../explanations/system-tables.md), in which DBOS stores its internal state. This command will create that database if it does not exist and create or update the DBOS system tables within it. - `-r, --app-role `: The role with which you will run your DBOS app. This role is granted the minimum permissions needed to access the DBOS schema in your system database. - `-s, --schema `: The schema name for the DBOS system tables. Defaults to `dbos`. - `--print-migrations `: Instead of running the migrations, print their SQL to standard output, either all of them (for a fresh database) or starting from a migration number (to upgrade an existing database). Postgres only. - `--print-user-role`: Instead of executing them, print the SQL statements granting `--app-role` access to the DBOS system tables. Use these last two flags to emit SQL you can apply yourself, for example if your database is managed by a DBA. They never connect to a database, so the `systemDatabaseUrl` argument is optional when using them (if supplied, it is only used to annotate the output). The output is only SQL and comments, but it contains `CREATE INDEX CONCURRENTLY`, so it must run outside a transaction block. ```shell npx dbos schema --print-migrations all ${DBOS_SYSTEM_DATABASE_URL} > migrations.sql npx dbos schema --print-user-role -r my_app_role ${DBOS_SYSTEM_DATABASE_URL} > grants.sql ``` #### npx dbos reset Reset your DBOS [system database](../../explanations//system-tables.md), deleting metadata about past workflows and steps. **Use only in a development environment.** **Arguments:** - `--sys-db-url, -s `: Your DBOS system database URL - `--yes, -y`: Skip confirmation prompt. #### npx dbos rename-application After renaming an application, transfer ownership of everything in the system database (workflows, steps, queues, schedules, and application versions) from its old name to its new name. Equivalent to [`DBOSClient.renameApplication`](./client.md#renameapplication); see there for details. Prints the number of rows transferred, by table. **Stop the application being renamed before running this.** **Arguments:** - `-s, --sys-db-url `: Your DBOS system database URL - `-f, --from `: The application's previous name. Omit to only adopt rows owned by no application (requires `--adopt-unclaimed-rows`). - `-t, --to `: The application that ends up owning the rows. Required. - `--adopt-unclaimed-rows`: Also transfer rows owned by no application. - `--batch-size `: The number of completed workflows and steps transferred per transaction (default: 10000) - `--schema `: The schema name for the DBOS system tables. Defaults to `dbos`. - `-y, --yes`: Skip confirmation prompt. #### npx @dbos-inc/create **Description:** This command initializes a new DBOS application from a template into a target directory. **Arguments:** - `-n, --appName `: The name and directory to which to instantiate the application. Application names should be between 3 and 256 characters and must contain only lowercase letters and numbers, dashes (`-`), and underscores (`_`). - `-t, --template