Agents can now see images and act on files

The agent runtime has learnt to see. Agents can now analyse images as part of a task, intake and process several files in a single pass, and delete workspace files when a process calls for it — all scoped to an account and its files, with the memory graph as context. This closes the gap between reasoning about text and acting on the mixed media that real work produces.

Because every agent action maps to the same service functions exposed to the UI and the REST API, nothing here is a private back door. Consequential writes still route through human approval, and the run history remains auditable, so wider capability comes without loosening the controls around it.

Published