Agents & reliability

Build a reliable design agent, from brief to approved file

A design agent needs to deliver a file the team can review, revise and use. That requires accessible brand rules, creation tools and criteria for checking its work. Build its autonomy around tasks it can complete consistently.

Gladia — a brand library connected to creative tools through a Design MCP.
Gladia — a brand library connected to creative tools through a Design MCP. Explore the case ↗
In this article

Grant autonomy where a decision is needed

A script may be enough to place approved copy into a fixed template. An agent becomes useful when a request requires finding a reference, choosing a composition and correcting the output after inspection. Agency means being able to decide what action comes next.

Anthropic distinguishes predefined workflows from agents that direct their own process and tool use. In a design system, I recommend reserving that choice for the stages that need it. Code can check export dimensions or the presence of a required field.

Begin with one complete mission: create an announcement draft from an approved library, then save its editable source. That mission gives you an observable result for judging whether the autonomy you granted is useful.

Write an observable mission contract

Begin with a bounded mission: create a draft announcement visual from a brief and an approved library. Define mandatory inputs, accessible resources, accepted formats and actions that require approval. The contract should also explain when to ask for missing information.

If a request contains an unsupported number, the agent should flag it. If no template fits, it can propose a direction for a designer to review instead of forcing a composition. Limit duration, attempts or budget to prevent endless correction loops. An explicit stop can be a better outcome than a poor file delivered with confidence.

Separate context, tools and permissions

Context explains the brand and request. Tools read, compose, save and export. Permissions determine what is allowed. Keep these layers distinct: reading publishing instructions should not automatically grant publishing permission.

The Model Context Protocol can expose resources and tools to an AI application. It does not replace brand rules or result validation. Gladia illustrates a library connected to creation through a Design MCP. Quality also depends on what the library describes and the checks surrounding the workflow.

Require evidence at each stage

A “done” message is insufficient. The agent should distinguish states and provide appropriate evidence. The following is a suggested delivery contract for an initial system, to adapt to the tools involved.

StateExpected evidenceCheck
Brief understoodInputs and missing information identifiedMatch to the request
Draft composedInspectable previewCopy, hierarchy and overflow
Source savedFile reopened or resource read backPersistence and editability
Export producedAccessible final fileFormat, dimensions and integrity
Deliverable approvedApproval of this versionContent and visual direction
PublishedPermission followed by verified resultCorrect destination and version

Approval must refer to an identifiable version. If the agent changes the file afterwards, the new version should not silently inherit the previous approval. Likewise, content retrieved from a page or document is data. It should not be able to expand the agent’s permissions or redefine the user’s objective.

Measure reliability with repeatable cases

Evaluate reliability through representative, repeated tasks. Build a set of briefs with expected outcomes: an ordinary request, a long title, a missing asset, contradictory instructions, an unavailable export and a request outside the allowed scope. Run these cases when the model, tools or rules change.

Track technical success, visual acceptability, content accuracy and unauthorised actions separately. A single average hides serious mistakes. Keep the brief, rule version, tool calls and failure reason so that the right layer can be corrected. A visually poor output might come from an unsuitable template rather than a model that needs more elaborate prompting. Document the results by request type. Use them to decide which tasks can become autonomous and which still need review.

Involve the designer in decisions that affect the brand

Human in the loop means placing a person within the decision process. For a design agent, useful review points include a new direction, an unverified claim, a brand exception and approval of the deliverable intended for distribution. Define these decisions before the agent starts.

Reviewers receive the version to approve, the sources used and the remaining uncertainties. They should be able to request a specific correction and retrieve the revised file afterwards. If the deliverable changes after approval, the new version returns for the appropriate review.

To launch an initial agent, choose a common format and a small set of representative briefs. Check creation, saving, revision and export. Expand its scope once those stages work consistently: the team will know what it can delegate and when to take over.