September 2026 · ONYT-2026-09 · m-2026-09-v1 · Protocol 1.0 · United States

How the indexes are scored

Onyt is not a model leaderboard. LMSYS, SimpleQA, and coding benches answer a different question: how good is the mind in the box? We ask whether the product behaves like help — whether it can take a real request and move it toward done for a real person, often inside a household.

The market is the United States. School paper, youth sports, iMessage and WhatsApp, Family Sharing, family court, Instacart, Echo Show on the counter. We do not rank products that cannot be used by a U.S. household this week, and we flag sunsets and store-listing ghosts instead of repeating dead roundups.

Each edition, editorial reviews public product surfaces, documentation, and hands-on use. Scores are 0–100 integers per dimension, then combined with fixed weights. We would rather be wrong in public, with a rubric you can argue with, than hide behind a proprietary “quality score.”

Assistant Index

Onyt ranks AI personal assistants — systems that take work off your plate — not chat models. The Assistant Index weights execution and household coordination more than raw conversation quality. The Family Life Index (U.S.) is a separate ranking of calendars, lists, chores, displays, co-parenting tools, and concierges. Scores are editorial, 0–100, refreshed each edition.

  1. 30% weight

    Execution

    Does it finish the work, or only talk about it?

    We score whether the assistant can take a real request — RSVP, order, book, remind, follow up — and carry it to a completed outcome with human approval, not just draft a reply.

  2. 20% weight

    Household

    Can more than one person share the load?

    Personal assistants live in families. We score multi-person coordination, shared lists and calendars, age-appropriate involvement, and whether one person still has to be the manager.

  3. 15% weight

    Memory

    Does it remember what matters next week?

    Persistent context across days: school calendars, preferences, open loops, and the difference between a chat history and a working memory of a household.

  4. 15% weight

    Privacy

    Who sees the school email and the grocery list?

    Data handling, on-device options, control over sharing inside a household, and whether the product is clear about training, retention, and human review.

  5. 10% weight

    Availability

    Is it where life already happens?

    Channels you already use (messaging, voice, email), always-on access, and whether you have to open a new app every time a plan changes.

  6. 10% weight

    Value

    Is the time it returns worth what it costs?

    Price against hours saved, free tiers, and whether paid plans buy execution or only more tokens.

Family Life Index

Onyt’s Family Life Index ranks U.S. apps that help households run everyday life — shared calendars, lists, chores, kitchen displays, co-parenting tools, and concierges. Weights favor follow-through and multi-person coordination, not feature checklists. Scores are editorial, 0–100, U.S. market only.

  1. 25% weight

    Follow-through

    Does it finish the errand, or only store the plan?

    Most family apps are filing cabinets. We score whether a product can RSVP, order, book, confirm a ride, or close a loop after a parent approves — not whether it can hold a to-do.

  2. 25% weight

    Coordination

    Can two adults and the kids share one working picture?

    Color-coded calendars, assignments, permissions, and whether one parent is still the secret project manager. U.S. households are the unit, including blended and two-home families.

  3. 15% weight

    Capture

    Can a school flyer become an event without retyping?

    Photo of a paper schedule, forwarded school email, WhatsApp or iMessage invite, voice dump. The mental load starts at intake.

  4. 15% weight

    Household ops

    Does it cover the rest of the house — meals, lists, chores, location?

    Grocery lists, meal plans, chore rotation, location sharing, documents. A calendar that cannot hold dinner is half a system.

  5. 10% weight

    Privacy

    Who sees the children’s schedule, and is it funded by ads?

    COPPA posture, ad-supported free tiers, court-record products, on-device options, and whether a U.S. family can explain the data story to the other parent.

  6. 10% weight

    Value

    In U.S. dollars, is the time back worth the price?

    Free-with-ads versus hardware-plus-subscription versus concierge. We price against a working parent’s hour, not against a chatbot token.

Test protocol

Six questions asked of every assistant this edition. Full inclusion rules, evidence grades, and the census live on the research desk.

  1. T1 · Intake

    Can a school flyer, forwarded email, or chat message become a structured next step without retyping?

  2. T2 · Consideration

    Does it present options a second adult can accept, reject, or reassign?

  3. T3 · Assignment

    Can the next step have an owner who is not the person who asked?

  4. T4 · Follow-through

    After approval, does anything leave the app — RSVP, order, calendar hold, message?

  5. T5 · Household graph

    Is the user a person, or a family with permissions?

  6. T6 · Privacy read

    Can a careful parent explain training, ads, kids’ data, and human review from the public policy?

This edition

  1. 2026-09-08

    Meta launches Muse in the U.S.

    Muse is Meta’s personal AI agent: WhatsApp, iOS, Android, and muse.ai, with a Secure VM, app connectors, and paid Power ($20) and Maximum ($100) plans. It is a one-person agent, not a household graph. Ranked separately from Meta AI, the chat bot already inside Instagram and Facebook.

  2. 2026-09-03

    OpenAI ships GPT-6 Astra

    ChatGPT’s frontier stack moved again. Agent-style computer use is real; household coordination is still a workaround. We score the product families use, not the bench table.

  3. 2026-09-01

    Claude Fable 5.1 and the Manus unwind

    Anthropic’s Fable 5.1 is the careful long-context leader. Manus resumed independent operations after its Meta deal was forced to unwind — still an agent for computer tasks, not a family product.

  4. 2026-08-27

    OurHome is gone; Maple is winding down

    The chore app every roundup still names is no longer installable; another developer has taken the name. Maple, a serious all-in-one family OS, has announced a December 31, 2026 shutdown. We keep both on the Family Life Index so searchers are not sent into a dead store listing.

  5. 2026-03-30

    Co-parenting apps went paid

    TalkingParents retired its free plan. AppClose, free for a decade, moved to a subscription in January. OurFamilyWizard remains the name U.S. family courts order. Split households are part of the U.S. market; we rank them on their own lane.

Partners and independence

Zazu is a partner of Onyt. That relationship funds the indexes. It does not set the weights. Partner products are labeled on every profile and in the ranking tables. If another product outscored Zazu on a published rubric, it would be #1. The weights favor execution and household coordination because that is the category we chose to rank — personal assistants and family operators, not chatbots.

Read the about page for the commercial disclosure, or cite the Index with the edition date.