The next interface for work

It should move at the speed of thought, begin with intent,
and understand the conversations that make work mean something.

00 / Introduction

For as long as there have been computers, we have been the ones adapting.

We learned to type, then to point, then to tap. We learned where the files live, what the buttons mean, and which order the steps go in. It made sense, because software could only follow exact instructions, so we learned to give them. We got so good at it that we stopped noticing how much of every day goes into explaining ourselves to a machine.

Ideas do not arrive that way. They arrive halfway through a sentence, in a meeting, on the walk back to your desk, in a conversation with someone who already knows what you mean. We think out loud and change our minds as we go. We say “the earlier plan” and trust the person listening to remember which one, and what happened to it since.

Software can now understand everyday language, yet most voice products stop at the words. Dictation types them. A command replaces a click. Both are faster, but the job has not moved: you still organize the thought, supply the background, and decide what the machine should do with it.

Hearing the words correctly is becoming easy. Say “Tell her we’re going with the earlier plan,” and the hard part begins after the last word. Which of the two plans from yesterday’s meeting is the earlier one? Is this a decision, or a thought still in progress? A good ear cannot answer either question. It takes a memory of yesterday, and the judgment to know which part of it still counts today.

Candor is a voice interface for work built around exactly that. You dictate a plan in the morning without breaking your train of thought. At noon the team pulls it apart in a meeting Candor is part of, and settles on something else. In the afternoon you say, “Tell her we’re going with the earlier plan,” and the follow-up starts from the right one, because the same intelligence heard both conversations. There is no dictation tool, meeting recorder, and chatbot to brief separately.

That only works if the memory underneath can be trusted. If yesterday’s proposal still outranks today’s decision, it has remembered the wrong thing. So it keeps what was said apart from what was decided, carries your corrections into tomorrow, and asks when it does not know. It grows with your team, and it follows the conversation off the screen and into the room.

Work can begin with a thought, spoken the way you would say it to a colleague. Understanding what you meant, remembering it accurately, and carrying it forward becomes the job of the software.

That is the idea behind Candor.

01 / Beliefs

Our core beliefs

The five beliefs below form one argument. Each begins where the previous one becomes insufficient, moving from a voice that hears words to an intelligence that can understand what still matters across people, places, and time.

01

Voice is not first-class until it understands intent.

Voice is often called an interface when it is really being used as a faster keyboard. You speak a paragraph, the computer types it, and the cursor moves on. The interaction is quicker, but the person still has to arrive with a finished thought and supervise the machine word by word.

Commands create the same illusion. “Open the plan” works only when there is one obvious plan and one obvious action. Real work sounds more like, “Send Molly the version we agreed on yesterday, but leave out the pricing change until I have spoken to her.” Every important word depends on something outside the sentence.

To understand that request, the computer has to follow the conversation behind it. It needs to know which Molly the person means, which draft survived yesterday's discussion, whether “her” still refers to Molly, and whether the pricing change was rejected or merely withheld. It also has to recognize the boundary inside the request: prepare one version now, preserve another decision for later.

The example reveals why intent is not a speech-recognition problem. Yesterday's discussion has to survive long enough for Candor to find the moment the team settled on a draft, connect it to Molly, and judge whether the speaker has authority to send it. At work, no single person necessarily holds that entire history. The words become useful only when those scattered pieces can be understood together.

The same words can also carry different intent depending on how they are spoken to the system. While dictating, the person may be composing the message to Molly. Asked as a question, the sentence may be a request to explain which version survived. Spoken while thinking aloud, it may require no response at all. A request to send the message crosses another boundary again. Until the interface can tell those moments apart, it understands the sentence but not the person.

Uncertainty is part of understanding, not evidence that the interface failed. If yesterday's meeting never settled which version to send, the intelligent response is a question. If the speaker lacks the authority to withhold the pricing change, the system cannot manufacture that authority from a confident tone. Knowing when not to proceed is part of knowing what someone intends.

The user should not have to think about those layers. They speak as they would to a capable colleague, with references, corrections, and unfinished thoughts intact. If the machine requires them to reconstruct the context and translate the request into exact commands, voice is still only decorating the old interface.

Recognition turns sound into words. Intent turns words into work.

02

A second brain is not a product. It is infrastructure.

Most people do not want to manage a second brain. They want to stop repeating themselves. A person may dictate three questions before a customer call, discuss them with the customer an hour later, and ask for a follow-up that afternoon. Today, each step often begins by briefing a different tool.

The first tool holds the questions. The meeting recorder holds the conversation. The writing assistant sees an empty message and waits for the person to explain both. Nothing was forgotten in the literal sense, yet the person still has to rebuild the day before continuing it.

A useful memory removes that reconstruction. The follow-up begins with the questions that led into the call, the answers the customer gave, and the correction the user made when the summary misunderstood one of them. The person stays inside the work instead of becoming the courier between three archives.

Continuity still requires judgment. A transcript proves that words were spoken, while a summary is an interpretation of them. The system can use both without pretending they carry the same weight. When the user corrects the summary, that correction has to reach the next interaction instead of dying with the document.

The memory can stay quiet in everyday use, but it cannot be opaque. If the follow-up contains a claim the customer never made, the person needs to see where it came from, correct it, and prevent the same mistake from returning tomorrow. Invisible effort is valuable. Invisible judgment is dangerous.

Trust changes what “remember” means as well. Permission to capture the call is not permission to retain every sentence forever, and retention is not permission to act. Those boundaries belong inside the experience, not in a policy page discovered after the history has already spread.

The second brain is the layer underneath the work: the part that lets one conversation begin where the last one ended. If the user has to visit and organize it all day, it has become the product again.

The best second brain is the one you stop having to visit.

03

Work intelligence should be multiplayer.

Most assistants become useful one person at a time. They learn how I write, remember the meetings I attended, and answer from the history I can see. That model works until the question belongs to more than one person, which is where most important work begins.

Imagine a customer renewal that has become uncertain. The account executive heard the customer ask for one concession. Support knows the problem behind the request is still unresolved. Product rejected the promised workaround in a meeting the account executive missed. Leadership approved a narrow exception, but only for this customer and only until the end of the quarter.

No individual has the whole account. An assistant limited to one employee can produce a polished answer from a partial history and be confidently wrong. A multiplayer intelligence can connect the authorized pieces, show where they disagree, and explain which commitment still governs the conversation.

What it remembers cannot be limited to final decisions. The rejected workaround matters because it prevents another team from offering it again. The temporary exception matters because a true answer this quarter becomes a false one next quarter. The unresolved support issue matters because it changes how the customer's request should be understood. Organizations run on this negative knowledge as much as they run on facts.

Sharing cannot erase the boundaries that made the conversations possible. A private call may help the system understand the account executive's intent without becoming visible to the rest of the company. A leadership decision may shape an answer without exposing the discussion behind it. Personal memory, shared context, and organizational knowledge are not one undifferentiated pool.

The network effect is simple to test. When Candor reaches support after sales, support's authorized context should make the account executive's next answer better. When product joins, the earlier users should gain an understanding they did not have before. If the second employee only adds a second subscription, there is no network effect.

Work intelligence becomes multiplayer when another person's participation improves what everyone else can understand without giving everyone access to everything.

No one holds the whole account. The organizational fabric should.

04

Context cannot stop at the screen.

Some of the conversations that change work never enter a work system. A customer catches someone after a meeting and says the deadline matters less than preserving one feature. Two colleagues settle a disagreement while walking back to their desks. The next morning, an email thread continues as if the old assumptions still hold.

The words on the screen are not wrong. They are incomplete. An intelligence reading only the email will optimize for the deadline because the conversation that changed its meaning disappeared when everyone left the room.

Software is the right place to begin. Native dictation captures what a person tells the computer, and native meetings capture what people decide together online. Selective access to messages can later fill a specific gap. That already covers a large part of professional conversation without requiring thousands of integrations into every database and workflow.

Physical work leaves a different kind of gap. Candor may not be present when the customer clarifies the tradeoff, and opening a laptop in the middle of that exchange may make the conversation worse. Those conversations are too important to remain outside the system, which is why hardware belongs in Candor from Phase I.

A meeting-room device, a desk surface, or portable professional capture can bring the same interface into places software alone cannot reach. The object is not the ambition. Its value comes from preserving a conversation that would otherwise vanish and returning it to the same memory, permissions, and intelligence as everything else.

That is why a tightly integrated hardware and software ecosystem matters to owning the interface. The software carries understanding across digital conversations. Hardware extends that continuity into physical ones. Neither becomes a separate product story.

The hardware commitment does not give every possible device a reason to exist. Each surface still has to close a specific capture or access gap better than the hardware people already have. Phase I begins bringing the interface into the room; later phases extend how completely it can be present there.

Most important conversations still happen outside a laptop.

05

Relevant recall is not enough. It has to be trustworthy.

A memory system can retrieve the perfect sentence and still give the wrong answer. Imagine a team discussing a United States launch in January and deciding on Europe in March. Ask about the launch region in April, and both conversations are relevant. Only one still governs the work.

Similarity cannot settle that question. The system needs to know that January was a proposal, March contained the decision, and the person who made it had the authority to do so. It needs the source of the change and the period in which it applies. Otherwise the oldest plan can keep resurfacing simply because it was stated more often.

The hard part begins before anything is retrieved. A frustrated remark after one bad call may not deserve to become a permanent preference. A confident sentence in a meeting may still be speculation. A model-generated summary may sound final even when the people in the room remained divided. Once one of those inferences becomes durable, it can return through every future answer wearing more authority than it ever earned.

Trustworthy recall preserves the difference between what was said, what was inferred, and what was decided. When Europe replaces the United States, the earlier plan does not need to disappear, because the history may still matter. It needs to lose the right to govern the present.

Sometimes there is no single current answer. Two teams may still disagree, or the person asking may not be allowed to see the conversation that would resolve the question. In those moments, confidence is not clarity. The useful answer is that the accounts conflict, what each one rests on, and what remains undecided.

Correction completes the loop. A person must be able to dispute a memory, replace it with a newer one, or revoke it entirely. The correction then has to reach the next answer without turning every local edit into a permanent rule.

The test is not whether Candor can find something related. It is whether a person can understand why this memory applies now and trust what will happen when it stops being true.

An old decision cannot keep wearing the badge of a current one.

02 / Interface

One intelligence, not a collection of products.

Imagine speaking a rough idea into an email in the morning, hearing the team reshape it in a meeting at noon, and asking for the follow-up before the day ends. Three different moments, but one continuous piece of work. The interface remembers where the thought began, what changed in the room, and what the person is now trying to say.

Three products. Nothing carries over. You explained it 0 times
  • Dictation appYou dictate the plan. It types the words, then forgets them.
  • Meeting recorderThe team changes the plan. It saves a transcript. It never knew there was a plan.
  • AssistantYou ask for the follow-up. It has seen neither the plan nor the meeting. You explain both, again.
Candor. One memory, all day. You explained it once
  1. MorningYou dictate the plan. Candor keeps it.
  2. NoonThe team changes it. Candor was in the room.
  3. EveningYou ask for the follow-up. It already knows both.

Three moments. One piece of work. It feels less like moving between tools and more like continuing the same conversation. That continuity is the opportunity, and deciding what deserves to survive it is the responsibility.

03 / Approach

Our approach

Every interaction creates evidence, but not every word becomes permanent. The intelligence identifies what may matter; confirmation, correction, source, and permission determine what carries forward. A correction to a customer's name may belong in future meetings. A rewrite made only to soften one email does not need to become a permanent preference.

That is the interaction loop. People speak because Candor helps immediately. The conversation gives it evidence. Governed memory turns the right evidence into continuity, and that continuity makes the next interaction require less explanation. Continued use creates more opportunities to confirm and correct what it understands. The memory becomes more trustworthy, not merely larger.

First-party capture is central to the strategy. Native dictation and native meetings let us shape the entire path from what was said to what is remembered, including the moment a person corrects the result. Broad integrations give access to more data, but they also pull the company toward shallow coverage of systems we do not control. We will add outside sources when they resolve a specific conversational gap, not because integration count looks like progress.

Conversation is where intent becomes visible, but it is not always the final record of consequential work. A confirmed delivery date may belong in a project plan, and an approved exception may belong in a customer system. Our role is to preserve the reasoning, keep the evidence connected, and let a confirmed outcome reach the appropriate record without becoming another universal work tracker.

Control follows the history wherever it goes. People need to see what Candor remembers, remove what it should not retain, and eventually carry their history elsewhere. They should stay because the intelligence understands them better, not because leaving means forgetting their own work.

The order is deliberate: earn frequent use through software and hardware in Phase I, prove that the right context survives accurately, extend that understanding across a team, and then broaden its physical reach and initiative.

04 / Plan

The master plan

Phase I

Build a voice interface people actually want to use all day.

The larger vision begins with a day that feels better from the first interaction. A person dictates an idea in the morning without breaking their train of thought. They enter a meeting where the idea is challenged and changed. That afternoon, they ask a follow-up question and the answer begins from the conversation Candor already witnessed.

This is the wedge: one voice intelligence across conversations with the computer and conversations with other people. Dictation creates frequency because people write throughout the day. Native meeting capture provides the explanations, decisions, disagreements, and commitments that give later requests their meaning. Memory connects the two so the day does not reset after every interaction.

The meeting capability is native from day one. We are not recreating the entire meeting-notes category, and the experience does not depend on a user importing history from another recorder. The conversation, its participants, and the moments that may matter later enter the same intelligence as dictation.

Capture alone is not enough. A transcript can preserve every word while losing what was actually decided. If the team challenged the morning idea but postponed the choice, the afternoon answer needs to preserve that uncertainty. When the user corrects the interpretation, the same mistake cannot return in tomorrow's message.

Phase I focuses on context we capture ourselves: dictation, native meetings, and direct conversations with the intelligence. Slack, email, and other sources can follow when they close an obvious gap. Arbitrary application data, broad screen understanding, and VoiceOS-style commands are not the launch product. Thousands of integrations cannot rescue an interaction that has not earned daily use.

Hardware is also a required part of Phase I. The first physical surface brings the same intelligence into in-person conversations and makes it available without asking someone to leave the room for a laptop. Its form must solve a real capture or access problem, but its place in the phase is not optional. Phase I establishes one system across conversations on the computer, conversations with other people online, and the important moments that happen away from both.

Trust begins at the same time as memory. Recording a meeting does not quietly authorize permanent retention of every sentence, and remembering a detail does not authorize action. Those distinctions need to feel clear without turning Candor into a memory settings dashboard.

The phase succeeds when the person can dictate before a meeting, participate in it, and continue the work afterward without moving context between products or briefing another assistant from the beginning. The first metric is not how many words we capture. It is how often the user no longer has to explain the day again.

Phase II

Turn personal intelligence into organizational intelligence.

Once Candor works for one person, the next change is not simply adding seats. It is allowing separately held conversations to improve shared understanding without dissolving the boundaries between them.

Imagine a customer renewal moving through the company over several months. Sales remembers a promise made on the first call. Support knows the problem behind it is still unresolved. Product rejected the proposed workaround, and leadership approved a temporary exception while the account executive was on leave. When the renewal appears in a planning meeting, nobody in the room holds the whole history.

Instead of pausing while five people search and message the absent account owner, the team asks what the company has committed to. The intelligence connects the conversations they are allowed to use, shows where each claim came from, and distinguishes the rejected workaround from the approved exception. If product and sales still understand the promise differently, it preserves the disagreement instead of inventing consensus.

The account executive returns a week later. A list of missed messages would show activity without restoring the mental model. Here, they can ask what changed, why the exception was approved, which customer commitment moved, and what remains unresolved. They return to the account, not merely the inbox.

Months later, a new employee inherits the renewal. The useful history is not only the final exception. It includes why the obvious workaround failed, which promise the company must not repeat, and who still holds context the system cannot share. Their first conversation with the team begins further ahead because the earlier reasoning survived.

If the original account owner eventually leaves, that same history becomes the handover. The organization keeps the professional context it is entitled to retain while private and personal material remains outside the record. Knowledge transfer becomes a continuing conversation rather than one hurried document written during a final week.

One customer history now connects a meeting, a return from leave, onboarding, cross-functional disagreement, and departure. The value comes from continuity across people, not from pooling everything everyone has ever said. Sometimes the most intelligent answer remains that the organization has not decided yet.

This is where the second brain as infrastructure comes into full force. Nobody has to visit a company memory and assemble the account by hand. The history sits underneath the conversation, bringing forward the right decision, disagreement, or missing piece when someone asks.

Phase II works when the second person makes the first person's experience better. If adoption increases the seat count but not the quality of shared understanding, the network effect is imaginary.

Phase III

Make intelligence available wherever work happens, then let it take more initiative.

By this point, Phase I hardware already carries the interface into selected in-person conversations. Two larger gaps remain: physical coverage is still incomplete, and useful context still waits for someone to notice that it matters. Phase III expands the hardware ecosystem and gives the intelligence more initiative without giving up the trust earned in the first two phases.

The first gap is physical coverage. A customer changes a requirement in a conference room, a team settles an issue at a whiteboard, or two people make a commitment away from their laptops. The initial Phase I hardware cannot cover every room and working style. A broader ecosystem of meeting-room, desk, and portable surfaces extends the same interface into those moments, all returning to the same memory and permissions rather than creating isolated archives.

Building hardware in Phase I does not remove the need for restraint. Every additional surface must preserve a valuable conversation or make the intelligence easier to reach without pulling attention away from the people in the room. The ecosystem grows by closing known gaps, not by producing objects for their own sake.

The second gap is initiative. Suppose someone promises revised numbers by Friday, then a pricing assumption changes on Thursday. The system already understands the promise and the change. Waiting for the person to rediscover the conflict and formulate a perfect question wastes that understanding.

It can surface the contradiction when the person begins the follow-up, explain what changed, and help prepare the response. It cannot silently decide which promise the company will honor. Recognizing a problem is not the same as having authority to resolve it.

That distinction keeps initiative useful. Permission to hear, remember, suggest, and act cannot collapse into one switch. Action begins with narrow, reversible work whose source, authority, and expected outcome are clear. A confident memory may justify a timely question. It cannot grant itself permission.

Phase III succeeds when the intelligence becomes more present without becoming more demanding. Fewer important conversations disappear, contradictions surface before they become mistakes, and people spend less time reconstructing the past before deciding what happens next.

05 / Say no

What we say no to

A bundle is not a new interface. Putting dictation, meeting notes, a chatbot, and a memory page in one menu may reduce the number of subscriptions. It does not reduce the distance between a thought and a useful response. Every part we build has to strengthen the same conversation across time.

The failure is easy to picture. Someone leaves a recorded meeting with a perfect transcript, opens the dictation tool to write the follow-up, and has to explain the meeting again before the tool can help.

Every feature worked. The interface did not.

We do not optimize for maximum capture. More transcripts can create more noise, more stale assumptions, and more confident mistakes. What matters is whether the right evidence survives with its source, status, and boundaries intact.

We do not build isolated point products that end at the transcript, message, or archive. We do not build a generic search, memory-infrastructure, or work-management layer that tries to understand everything about a company. And although hardware begins in Phase I, we do not build any physical surface as a disconnected product or without a necessary role in the same interface.

We will not compete with frontier labs by building another general model, thousands of integrations, or broad VoiceOS-style application control. Their advantage is breadth. Ours has to come from depth: owning the interface through which people express intent and the first-party conversational history that makes the next interaction intelligent.

Other models may perform parts of the reasoning, and other systems may hold the final operational record. Our place is the conversational layer between them, where people speak, context accumulates, and what carries forward is decided.

If a new capability does not make voice more useful now, improve what can be trusted later, or strengthen shared understanding, it does not belong in the flywheel.

06 / Why

Why will this work?

The behavior already exists in fragments. People speak when they want to move faster than a blank page allows. They record meetings because nobody wants to write down everything while also participating. They search old threads because the reason behind a decision has vanished. Then they brief each assistant again because none of those tools remembers what the others witnessed.

The timing has changed because capabilities that once looked scarce are becoming inputs any company can buy or build on. Recognition keeps improving, general models keep absorbing more kinds of work, and actions that required specialized software are becoming easier to generate. Accuracy still matters. It simply cannot be the whole advantage.

What compounds is the understanding created between one interaction and the next. Over time, Candor learns which decisions changed, which corrections mattered, how people refer to one another, and what a team is allowed to know. Another microphone or model can reproduce the first interaction. It cannot make the person live the intervening history again.

That advantage only exists if the history remains trustworthy. Native capture is not a moat when it produces a larger transcript pile. Memory is not a moat when old decisions remain active, private conversations leak across boundaries, or leaving Candor means abandoning your own history. A polished interface is not a moat when every conversation still begins from zero.

The thesis therefore carries its own test. The second conversation has to begin further ahead than the first without weakening the person's control over what is remembered. When a coworker joins, shared understanding has to improve without giving everyone access to everything. A new surface earns its place only by closing a real gap in that same loop.

If we meet those conditions, the change will be visible in ordinary work. A person will leave a meeting, begin the follow-up, and speak without reconstructing what everyone said. The intelligence will understand which decision still counts, ask when it does not know, and carry the correction into tomorrow.

For eighty years we learned
the machine’s language.
Now it learns ours.

Back to top