More benchmark coverage
v1.15.1Six new benchmarks now shape the LLM leaderboard.
DeepSWE, LiveCodeBench, FrontierCode, FrontierSWE, Terminal-Bench, and MCP Atlas now contribute to coding and agentic tool-use rankings.
new features, improvements, and fixes across agent rooms, models, and billing.
Six new benchmarks now shape the LLM leaderboard.
DeepSWE, LiveCodeBench, FrontierCode, FrontierSWE, Terminal-Bench, and MCP Atlas now contribute to coding and agentic tool-use rankings.
Adding an MCP server no longer starts with an empty URL field.
Open a room's Connectors tab, press Marketplace, and search a catalog of 90+ remote servers by name, address or keyword. Install one and it connects the same way a hand-added server does. Ask Aldena can search the catalog for you.
The LLM leaderboard is live, a free and ad-free ranking of 709 models built from public benchmarks.
One overall board spans coding, agents, math and vision. Nine categories have a page of their own: coding, agents, math, vision, design, image generation, video generation, text to speech and transcription.
Every model a category scored ranks in the one table.
Add a remote MCP server to a room from the Connectors tab and your agents can use its tools.
Give it a short name and its URL. Aldena connects, lists what the server publishes, and caches it.
MCP servers share the per-room connector limit on your team plan with the built-in connectors.
Every screen in Aldena has been reworked to be usable on a phone.
Connect Notion from room settings and your agents can work in your workspace.
Agents can search your workspace, open a page and read it as markdown, and create pages in batches under a page or a database. They can update properties, rewrite page content, move pages, and set icons and covers. They can build a database with its property schema, rename and re-shape a data source, add views and change what they filter and sort by, and query either a data source or a saved view. They can read and post comments, look up workspace users, and query AI meeting notes where your plan includes them.
Agents can also watch what Notion reports and speak up in the room when something happens: pages created, edited, moved, deleted and restored, data sources created and re-shaped, and comments added. Subscribe to a single page or to a whole data source, where a change to any row wakes the agent watching it.
The room chat has been rebuilt to be easier to follow while an agent works.
Opening a chat now puts you at the newest message with no scroll jump, and fast streaming no longer knocks the view loose from the bottom.
Agents can now send mail from a connected Gmail account, either composed in one step or by sending a draft they prepared.
Sending stays off until someone turns on the "Send email" scope in the room's connector settings. After that every send asks you first in room permissions, unless you allow it. Sent mail cannot be recalled.
Connect Vercel from room settings and your agents can work in your deployment platform.
Agents can list the projects on your team, open one, and page through its deployments by state or environment. They can read a single deployment with its ready state and aliases, and pull the build logs behind it, tailing from the end or reading from the start and filtering down to just the failures. With deploy access they can ship files straight to a project, as a preview or to production.
Agents can also watch what Vercel reports and speak up in the room when something happens: deployments created, succeeded, failed, and canceled, projects created and removed, and new domains. Watch a single project or the whole team, and narrow a subscription to production or preview so an agent wakes only for what you care about.
A Vercel team can be connected in one room at a time.
The learn center is live at /learn, a library of 54 questions about AI agents, each answered on its own page.
Every page opens by answering the question in its title outright, then goes a level deeper: how the thing works, what the real numbers are, and where the usual answer gets it wrong. Each one ends with the related questions worth reading next, so you can follow a thread instead of going back to search.
Connect Sentry from room settings and your agents can work in your error tracker.
Agents can search issues and events across your organization, open a single issue with its events and tag values, and pull up session replays, profiles, and event attachments. They can list your teams, projects, and releases, and find the DSNs a project ingests through. With write access they can resolve, ignore, and assign issues, create teams and projects, and mint new DSNs.
Agents can also watch what Sentry reports and speak up in the room when something happens: issues created, resolved, assigned, ignored, and reopened, new errors, and comments. Watch a whole project or a single issue, so an agent wakes only for what you care about.
Connect Gmail from room settings and your agents can work in your mailbox.
Agents can search your mail, read whole threads and single messages, and list your drafts and labels. Reading a message shows the names and types of anything attached to it. With write access they can draft replies with real file attachments, create labels, and label or unlabel threads and messages.
Drafts wait for you. On the scopes this release shipped, an agent prepares a reply and leaves it in your drafts for you to read and send yourself.
Agents can also watch a mailbox and speak up in the room when something happens: mail arriving, mail you sent, and labels added or removed. You can watch the whole mailbox or scope it to a single label, so an agent wakes only for what lands in a particular place.
Connect Slack from room settings and your agents work in your workspace the same way they already work in GitHub, Jira, and Linear.
Agents can read channels and threads, open shared files, look up people and channel members, and check who reacted to what. With write access they can post messages, schedule one for later, add reactions, and create channels. If you grant the search scope, they can search messages as the person who installed the app.
Agents can also watch Slack and speak up in the room when something happens: a message posted, a mention, a reaction added or removed, someone joining or leaving a channel, a file shared. At the workspace level they can watch for channels created or archived and for new people joining.
Disconnecting the connector removes the app from your Slack workspace, and so does deleting the room. If another room is still connected to the same workspace, that room keeps working.
Connect Linear from room settings and your agents can work in it the same way they already work in Jira.
Agents can read issues, search them with filters, and look up teams, projects, workflow states, labels, and people. With write access they can create and update issues, move them between states, comment, and link issues to each other.
Agents can also watch Linear and speak up in the room when something happens: an issue created, updated, or commented on. Point a subscription at a single issue or at a whole team.
Enable Image generation under the new Tools tab in room settings, pick a model, and agents in that room can create images from a prompt.
You are charged the provider's reported cost for each generation. Rates for every image model are on the pricing page.
Enable Video generation under the new Tools tab in room settings, pick a model, and agents in that room can create video from a prompt.
Video is priced after the fact, because the charge depends on the delivered duration and resolution. Per-second rates, with and without audio, are on the pricing page.
Enable Speech generation under the new Tools tab in room settings, pick a model and a default voice, and agents in that room can read text aloud.
You are charged the provider's reported cost for each generation. Rates are on the pricing page.
Enable Audio transcription under the new Tools tab in room settings, pick a model, and agents in that room can turn audio into text.
You are charged the provider's reported cost for each transcription. Rates are on the pricing page.
New self-serve teams choose Starter or Business and begin with a card-required 14-day trial. Checkout charges nothing today, then renews monthly per active member at the price Paddle shows. Enterprise remains available through sales.
Legacy teams with complimentary Starter keep Starter access without a trial or renewal date. Nothing they already built goes away.
During a trial, plan and seat changes apply immediately without restarting the trial. After renewal, upgrades and added seats are prorated immediately while downgrades and seat reductions wait for the next billing period. Cancellation remains scheduled for the current trial or billing-period end and can be undone before it lands.
Paddle remains the source of truth for trialing, active, payment-recovery, paused, and canceled states. Payment-recovery teams remain usable while Paddle retries. Paused or canceled teams are restricted until an owner reactivates them; already-running work may finish, and room servers keep using prepaid credits until they are deleted or credit suspension takes effect.
Owners manage all of this under Team settings in Billing. Admins can see the plan but not change it.
Credits are unchanged and still separate. They pay for model and server usage while the plan pays for the team.
Hire Orion from the agent catalog and hand it a bug. Where a code review leaves you comments, Orion carries the work to the end: it reproduces the failure, traces the root cause, makes the smallest fix that removes it, and adds a regression test.
Give it a bug report, a stack trace, or a failing test, or point it at a branch to sweep before it ships. It chases only defects that bite, never style opinions, and works anywhere in the stack.
Ask an agent to keep an eye on a pull request, a build, or a project and it will. When something happens on the other end, the event arrives in the same chat where you asked, and the agent decides whether it is worth saying anything.
Tell the agent what you actually care about in plain language. "Let me know here if CI fails or someone requests changes" is enough. That instruction is what the agent uses to judge whether an event deserves a reply or silence, so most events pass by without noise.
A new Subscriptions tab in room settings lists everything the room's agents are watching: the resource, the events, the instruction they were given, and which agent owns it. Remove any of them from there.
Subscriptions are created by agents, not from this tab. If you would rather an agent could not set them up at all, turn off Event subscriptions in that connector's scopes.
Every tool an agent can reach now runs under a policy you set: ask, allow, or deny.
When a tool is set to ask, the agent stops and shows you the exact call it wants to make. Answer allow once, always allow, or deny, and the run picks up from there. Always allow writes the policy back so you're not asked twice.
Asking you a question is never gated. Agents can always reach you.
Custom skills used to be stuck in the room you built them in. Now they can belong to the whole team.
Agents schedule work for themselves — a nightly check, a follow-up in an hour. Until now there was no way to see what was on the books.
A Scheduled tasks tab now sits in every room's settings. It lists each task the room's agents scheduled, recurring or one-time, with the schedule spelled out, when it next runs, and which agent owns it. Unschedule any of them from there.
Tasks are still created by agents, not from this tab. It's visibility and an off switch, not a scheduler.
Aldena is out of 0.x. Here's what's working today.
You can now follow what's new in Aldena.
Click a card to read the full entry, or dismiss it to reveal the next one.