Cara-4
Anam's expressive realtime digital-person model for live conversations with image-based avatars, emotion direction, and bring-your-own AI components.
Anam's expressive realtime digital-person model for live conversations with image-based avatars, emotion direction, and bring-your-own AI components.
Plain-English overview
Cara-4 generates a live visual performance rather than waiting to render a complete video. Teams can start from realistic, 3D, or illustrated identities, direct expression with natural-language notes, and use Anam's stack or bring audio from their own language and speech systems.
Anam publishes portrait and landscape output sizes plus server-side avatar latency, which helps with interface planning. A real deployment still needs end-to-end testing for turn-taking, connection time, device performance, and the latency of every connected AI service.
Category comparison
These are provider-published specifications, not Cody benchmark scores. Follow the linked sources for current limits and endpoint-specific exceptions.
API and provider access
Availability
Available in Anam's API and Lab, with integrations for web applications and major video-meeting products.
Anam documents service and enterprise controls but does not give one exhaustive model-specific country list on the launch page.
Check live availabilityData and training
Anam says customer session content is not used for training unless separately agreed. Default recordings may be retained for up to 30 days, while enterprise zero-data-retention is available; verify the exact workspace settings.
This is a concise reading of the cited provider material, not legal advice. A third-party gateway can have different storage, routing, training, and residency terms from the model maker's direct API.
Read the provider policyFrequently asked questions
Anam's expressive realtime digital-person model for live conversations with image-based avatars, emotion direction, and bring-your-own AI components. Cara-4 generates a live visual performance rather than waiting to render a complete video. Teams can start from realistic, 3D, or illustrated identities, direct expression with natural-language notes, and use Anam's stack or bring audio from their own language and speech systems.
Cara-4 was released on July 14, 2026 according to the cited provider materials.
Available in Anam's API and Lab, with integrations for web applications and major video-meeting products. The access routes listed in this guide are Anam.
30 free minutes; paid overages roughly $0.11–$0.16/min. The effective rate depends on Anam's plan and included minutes. Enterprise capacity and zero-retention controls are separately negotiated.
Available in Anam's API and Lab, with integrations for web applications and major video-meeting products. Anam documents service and enterprise controls but does not give one exhaustive model-specific country list on the launch page.
Anam says customer session content is not used for training unless separately agreed. Default recordings may be retained for up to 30 days, while enterprise zero-data-retention is available; verify the exact workspace settings. The policy belongs to the provider route and account terms, so verify it again before production use.
Related comparisons
A realtime avatar model that can animate a person, illustration, or nonhuman character from one image and connect to an existing voice-agent stack.
Open full comparisonHeyGen's rendered avatar model for producing polished presenter videos from a short identity recording, script, or supplied audio.
Open full comparisonHedra's audio-driven character model for creating talking videos from one image, one to four speakers, and optional performance direction.
Open full comparisonSynthesia's business-video avatar model for script-aware speech, gestures, body motion, and multilingual presenter content inside a complete editor.
Open full comparison