What I Learned Building an Internal AI Tool Used by 200+ People
Building an internal AI tool for more than 200 people changed the questions I paid attention to.
What I built was Promptly, an AI suite for editorial and marketing teams.
At the beginning, it was easy to focus on the obvious things: which model to use, how to write the system prompt, what the interface should look like. Real adoption brought a different set of concerns. Would a helpful constraint today become a limitation after the next model release? What happened when a feature improved the answer but made the experience worse? How should we think about costs once a small group of power users began using the tool much more than everyone else?
None of the insights below is a universal rule. They are the product decisions that held up best as the tool moved beyond a demo and became part of people's actual work.
Build around a moving model, not its current limitations
Every AI product has to decide what the model should handle, what the product should enforce, and what should remain under the user's control.
The tempting response to unreliable model behavior is another prompt rule or a tighter harness. That may fix the immediate problem. It can also reduce the model's usefulness elsewhere. A constraint introduced for one edge case quietly becomes a constraint on every request.
Planning made this trade-off especially clear. Many AI products create a plan as soon as the user submits a task. The idea makes sense: give the model a sequence to follow so that it does not miss an important step. I chose not to make that a fixed part of the product.
The problem is that the right steps are not always knowable at the start. Action produces information. A web search may reveal something that changes the task completely. A plan written too early can make the process look orderly while preventing the model from responding to what it learns.
That choice was a bet. I trusted model providers to improve task completion at the agent level instead of permanently patching the limitation in my own product. Models are a moving foundation. A workaround that is useful today may be redundant one generation later, yet remain in the product as another layer of complexity.
Users cannot wait for the next model, of course. The practical answer is to be selective. Add model-level constraints when they clearly earn their cost, and make them easy to remove. When a sensible default, a UI decision, or explicit user control solves the same problem without narrowing the model, I prefer that.
Where each decision lives
My bet was leverage, not replacement
For this tool, the useful promise was not: give us the work and we will automate it away. It was: keep control of the work, but do it with more range and less friction.
That distinction mattered. Other internal tools made different bets about how much work people wanted to hand over. Some reached implementation, then disappeared when those assumptions did not survive real use. Editorial teams had seen automation produce drafts that took more effort to repair than a collaborative process would have required in the first place.
An assistant was easier to trust because it respected the user's expertise. The tool could help someone explore a topic, draft, analyze, and revise without pretending that judgment had become unnecessary.
My bet was that this would lead to better work and more durable adoption. In the workflows I observed, people got more value when they directed the process, judged the output, and changed course after a weak first attempt. The tool grew beyond 200 users while keeping that collaborative model intact.
I reinforced the bet in the interface. I framed the product as a workspace or toolbox, removed low-value friction, made iteration feel normal, and kept the user visibly in charge.
That outcome does not make the same approach right for every AI product. A stable, repetitive workflow may reward a much stronger automation bet. For this tool and these editorial workflows, preserving user agency matched how people actually wanted to work. Other approaches came and went. This one lasted.
The bet that survived real use
Treat feedback as evidence, not a specification
Users often describe a solution: a button, a workflow, a setting they want. My job was to work backwards from that request. What problem is this person actually trying to solve? Does it recur? Can one capability solve it for more people?
A single request may be a passing thought. A second is a reason to investigate. A recurring pattern may point to a real product opportunity. Even then, the best solution can look quite different from what users originally proposed.
I tried to make the tool more composable instead of more bespoke: a small set of reliable capabilities that work on their own and combine into more complex workflows. Over time, more requests should be answerable with, “Here is how the existing pieces can help you do that,” rather than, “We need another special case.”
Power users deserve particular attention. They cost more to serve because they use the tool more intensely. They also expose limitations earlier, bring unfamiliar workflows, and often ask for capabilities that other users will want later.
But power users are a discovery channel, not the entire product. Their requests still need to be interpreted in the context of the wider user base. Lighter users need an experience they can understand without adopting a power user's habits. They are also what turns enthusiastic early use into durable internal adoption.
Do not return work that users thought they had delegated
One feature I expected to help let the chat ask a clarifying question before producing an answer. In theory, better input should have led to a more precise result.
In practice, the feature often annoyed people.
They returned to the chat expecting completed work and found another question waiting for them. The possible gain in answer quality did not make up for the broken expectation. They thought they had handed off a task, but the product handed part of it back.
This taught me to think about every feature as part of an interaction contract. A clarifying question can be useful, but the moment and the reason have to be clear. Otherwise, the product feels less capable even when its eventual answer is technically better.
The same applies to latency. A modest improvement in output is not automatically worth a response that takes several times longer. Quality includes usefulness, speed, friction, predictability, and the feeling that the work is moving forward.
That is why I now treat changes like this as experiments, not automatic progress. If real use shows that a well-intentioned feature makes the overall experience worse, remove it. Reversing a decision is not a failure. It is the product responding to evidence.
The interaction contract
Manage the economics before they manage the product
An AI tool does not have one stable cost curve. Adoption can grow. Existing users can become heavier users. Providers can change their prices. A new feature can suddenly make each session more expensive. Several variables move at once.
A fixed price per user is simple, but it can push a product towards artificial limits or unsustainable economics. Usage-based pricing kept more options open for me. It also required an ongoing conversation with the company: forecast demand, agree on a budget, watch actual use, and revise the decision openly.
I found staffing a more useful comparison than conventional software licensing. Spending should follow demonstrated demand and value, not the desire to say that the company has adopted AI.
Cost control works in both directions. The product needs enough budget to support valuable work, and it should use that budget efficiently. But the lowest possible bill is the wrong target. A cheaper workflow that produces weak results, causes more retries, or erodes trust may cost less per run while wasting more money overall.
The better target is useful outcomes per unit of spend. To measure that, the product needs a recognizable unit of value. In my case, one useful unit was a saved prompt that people ran repeatedly to produce a mostly complete artifact. Each run had a calculable cost, and the output cleared a recognizable quality threshold. Another product might use a research brief, a resolved support request, or a draft ready for human review.
Only then can an efficiency change be judged properly. It is an improvement if it reduces cost without pulling the result below the level people actually need.
The target is not the smallest bill
Once people depend on it, operate it like infrastructure
An internal AI tool may begin as an optional extra for a small group of enthusiasts. Once people build it into real workflows, the standard changes. A failure no longer interrupts an experiment. It blocks someone's work.
Support questions become more urgent. Criticism becomes more direct. Users may blame the product even when their own workflow or understanding contributed to the problem.
Blaming them back is useless. Confusion is product evidence. It points to something the interface, the behavior, or the communication failed to make clear.
The operating model has to mature with adoption. Protect user progress as a core reliability requirement. Test changes to foundational behavior more rigorously than ordinary feature additions. Use feature flags and small cohorts so that a new capability does not reach everyone at once. Power users are often good early adopters because they are willing to explore a change and can explain where it breaks down.
Failures will still happen. Trust then depends on the response: acknowledge the problem plainly, fix it quickly, and involve affected users in the recovery when that is useful. Asking them to assess the improvement gives them real influence over a tool they now rely on.