Blog

2026-10-02

Blog

Ask after the agent resolves the task

A chatbot NPS at first token measures the wait, not the job. Ask at true resolution, bind the rating to the session you already trace, and read it honestly.

Thursday, 14:07. Lena owns the support copilot at a 28-person SaaS. Ticket #6120 is a $48 refund. The user asked at 14:02. At 14:03 the agent replied that it was looking up invoice INV-8841. Eight seconds later a 0-10 docks in the chat: how likely are you to recommend our assistant. The user scores 3 at 14:04: it has not done anything yet.

At 14:11 the refund posts. The user says the credit landed. Friday's copilot review quotes assistant CSAT 2.4 (n=11). Thursday's sample is people who were asked while the tool was still running.

The weak version, and what it costs on Friday

The program that produced that review looks like this:

The same auto-trigger widget as the marketing site. Delay 8,000 ms. Every chat session, including the first assistant message and the turn where a tool is still open.

The same 0-10 as product NPS, mixed into the product headline, because one score is simpler.

A model-requested ask-them-now tool left on, so the agent can prompt when it is unsure, and those rows sit in the same number.

The version that scores the job looks like this:

Call resolveTask when the user's outcome is confirmed. For ticket #6120 that is 14:11, refund applied, not 14:03, looking it up.

One inline ask then. Optional comment. Never re-prompt after an answer.

Mint the show-token on the server and bind it to the same sessionId you already send to Langfuse.

Leave the model-requested tool off. If you turn it on, those captures stay out of the headline.

The weak version costs you Friday. You will ship a P1 called copilot CSAT collapsed. You will not notice that INV-8841 refunded in nine minutes. Next week you will add a faster first message, which makes the first reply shorter and the score worse, because you still ask before the job is done.

First reply is not resolution

An idle delay on a chat page is the same bug as an idle delay on first paint. The timing post is that rule for tickets versus onboarding. The anti-pattern post is that rule for password reset and the error page. A copilot is the same clock with a friendlier surface: the person is mid-task, the agent has not finished, and you measured the wait.

True resolution is the confirmed outcome. Refund applied. Ticket closed. File exported. Appointment booked. Handed to a human is not done if the user came for the refund. The agent quickstart is the call: resolveTask({ outcome: "refunded" }) on the client, after the outcome is real. Until that fires, the task stays open and the ask stays dark.

If the outcome lands off-platform days later, mint a deferred ask on the server and deliver it on your channel. UserVane does not send that email for you. The open-source SDKs include a clone-and-run chat example so you can see the bind and resolve order before you wire production.

Bind the rating to the session you already have

Your traces already have a session id. The useful rating is the one that cannot be pointed at a different session after the fact. Mint the token on the server with the secret key, thread the show-token to the client, and call bind before resolve. Submit then carries that session. A client-altered id that does not match the token is rejected.

That is a capture layer on top of the tracer you already run, not a second tracer. UserVane does not record the agent's tool tree. It stores the user's score on the session you named. The agents page is the product surface for that loop, and the correlation guide is the binding rule.

Without bind, bootstrap still works and those rows store as unlinked. Unlinked rows never enter the Langfuse push queue. If you want the number on the same session you already watch, bind it.

Read the number in the assistant you already pay for

Collection happens in the product, at resolve time. Reading it can happen in Claude, Cursor, or any MCP client on a paid UserVane org. The MCP page is that connection: OAuth, then get_results and list_responses. Thin samples stay suppressed. The margin comes with the score. The assistant you already use can pull Thursday's comments without a CSV.

MCP does not decide who gets asked. It does not let the model write itself a better headline. Create the post-task survey there if that is how you work; keep resolveTask in the copilot as the clock.

Eleven mid-task 3s are not a copilot grade

A linked rating proves the score belongs to that session. It does not prove the agent is safe, aligned, or that the eleven people who answered represent the ones who closed the tab. That is the limit. Check the slice on the sample size calculator before Friday's review treats 2.4 as a climate. The older post on why a naked score lies still applies, per survey, including the copilot one.

Keep the copilot ask as its own survey. Do not mix it into product NPS. Staff who dogfood the agent belong behind an employee trait, the same way they stay out of the relationship sample. One primary ask per person per quiet period, so a refund CSAT at 14:11 does not chain into a product NPS at 14:12.

Bottom line: ask when the job is done, on a token bound to the session you already trace. Leave first reply, open tools, and model-requested prompts out of the headline. Friday should be a decision about ticket #6120, not a score of the wait.