THE PROJECT STORY
A project story about hiring trust, privacy, context, controls, and learning the build side without pretending to be an engineer.
Years ago, I tried to build a technology business through Screenager. The idea had merit, but it did not become the business I believed it could be.
The limitation was not that I could not see the problem. It was that I was dependent on an engineering team to translate the problem into a working product. I could explain the customer need, describe the experience and make commercial decisions, but I could not stay close enough to the build to inspect it, test it and change it myself.
When the translation between the idea and the technology broke down, I did not have the ability to close that gap.
I later watched businesses built around similar ideas succeed, often led by founders who could work much closer to the technology. That stayed with me. A good idea was not enough. The ability to test it, challenge it and keep changing it mattered just as much.
A few months ago, I decided to find out whether generative AI had moved that barrier.
I did not start because I wanted to call myself technical. I wanted to understand whether someone with operating and business experience, but without an engineering background, could now take an idea beyond a presentation, a requirements document or a prototype designed by somebody else.
Verlinko became that test.
I have spent most of my career in operations and transformation. In that world, a process is only as strong as the assumptions nobody has thought to challenge.
Remote hiring rests on a large one.
We verify qualifications, references, right to work and employment history. But we usually assume that the person attending the final interview is the person who completed the earlier stages, and that the person who arrives on Day One is the person who earned the role.
AI makes that assumption easier to exploit. Convincing identities, manufactured histories, real time assistance, synthetic video and proxy interviewing are no longer specialist capabilities. The problem is not that employers have suddenly become careless. The problem is that most hiring processes were never designed to maintain continuity of the human being across the journey.
My starting question was therefore not, “Can AI detect another AI?”
That looked like an arms race. Every detector would eventually face a better generator.
The question I chose was simpler:
Can an organisation establish that the person it is speaking to now is the same person it met earlier?
That shifted the problem from detection to continuity.
Most identity checks happen during onboarding. By then, the organisation has already made the decision.
If a proxy completed the assessment or attended the interviews, verifying an identity document during onboarding does not reconnect the person who arrives with the person whose capability was evaluated. It proves that an identity exists. It does not prove continuity across the hiring process.
I wanted the anchor to be created at the first meaningful interaction and then checked at each point where substitution could occur: later interviews, Day One and, for genuinely sensitive roles, brief moments during tenure.
The second constraint mattered just as much.
Could the system establish that continuity without becoming another repository of personal information?
The obvious answer to an identity problem is to collect more identity data. I wanted to test the opposite approach.
No stored interview recordings. No retained identity documents. No candidate names or email addresses in the verification system. No library of faces to search. One encrypted biometric template, treated as sensitive biometric data, used only to compare one person with their own earlier anchor.
That combination became the core of Verlinko:
I believed in it enough to file a patent on the approach.
I’m saying that because most of what follows is me finding out which parts of that belief were wrong. Those two things sit together fine, I think. The patent was a bet that the problem is real and worth solving. Removing three of my own features later was just accepting that the measurements were right and I wasn’t.
Verlinko is now a working prototype rather than a concept deck.
A candidate opens a link and gives consent. A face template is created and encrypted. Their phone becomes a second camera, providing an independent view of the room and the person from another angle. An authorised interviewer approves the session before it begins. Silent questions can be shown on screen so that an earpiece or remote listener has no spoken question to relay. Later continuity checks compare the person with the original anchor.
The design deliberately avoids making hiring decisions. It records signals and evidence for human review. A failed comparison asks for another check and then routes to a person. It does not reject a candidate.
The prototype also records the rule used at the time of each comparison: the model, threshold and method. That matters because a result without the rule that produced it is difficult to interpret later.
The public site includes the problem and architecture, the primary evidence behind the problem, a local demonstration, and a build status page that lists the gaps alongside the implemented controls.
I need to be clear about what “I built” means.
I did not personally write every line of code. AI generated much of the implementation. I used it for research, architecture, coding, testing, documentation and review, primarily through Claude and Codex.
My role was to define the problem, set the boundaries, make the product decisions, challenge the proposed solution, test what had actually been built and decide what needed to change or be removed.
That distinction matters. I have not become a software engineer. What I developed was the ability to move much closer to the build and to take responsibility for the decisions around it.
This was my first serious attempt to build in this way. I did not know what a good AI assisted development process looked like. I had barely encountered the term “vibe coding.”
At the beginning, the experience was intoxicating. I could describe a feature and see something appear minutes later. Every session felt productive because there was always visible output.
I was not yet building a coherent system. I was accumulating one.
Features arrived faster than my ability to understand how they connected. Files landed wherever the AI placed them. Names made sense in the conversation that created them but not necessarily to the next session. Decisions changed while the documents describing the previous decision remained. Several files could each appear to be the current source of truth.
The product kept getting larger, so it looked like progress. Underneath it, the cost of understanding the project was rising.
At first, I used one AI as an adviser and another inside the code editor. They could not communicate with each other.
I became the wire between them.
I would describe a problem to the adviser. It would write a complete prompt for the coding tool. I would copy that prompt into the editor, wait for the result, paste the result back, take screenshots of what happened and repeat the process.
Every coding prompt began with another explanation of what Verlinko was, what we were building and why. Every new conversation required me to paste the same context again.
Without realising it, I had already encountered one of the central problems of working with AI: context does not automatically persist, and missing context is reconstructed differently every time.
My answer was manual. I carried the context from one system to another, one paste at a time.
That worked until the project became too large for me to hold the integration together reliably.
My first response to context loss was to provide more context.
More background. More instructions. More notes. More documents explaining the same system from different moments in its development.
At one point, the project had accumulated 51 documents describing itself.
It was not better informed. It was more confused.
With 51 documents that partly disagree, the answer depends on which document gets read first. A carefully written document can be actively harmful if it describes a state that no longer exists.
We eventually reduced those documents to 16 and made one file the router. That file did not try to contain every answer. It explained where the authoritative answer lived and what should win when two sources disagreed.
The hierarchy became simple:
That small hierarchy changed the quality of the work more than another large context document would have.
The lesson was not “write more documentation.” It was:
Structured context beats abundant context.
A document that nobody is required to read has no operational effect. The real skill is making sure the right source is read at the moment someone, or something, is about to act.
One of the more uncomfortable discoveries was that I had written many of the right rules early in the project.
I created a Build Methodology document that said:
The rules were sound. I still stopped following them consistently.
They were attached to a document that gradually became outdated. As confidence in the document fell, the discipline contained inside it disappeared too. There was no decision to abandon the method. It simply stopped being part of the active workflow.
I later moved that document into the archive and made it required reading through the current rules file.
The archive is not there for nostalgia. It is where the expensive lessons live. I had already paid for them once. I did not want to pay for them again.
Writing down a lesson is not enough. A system has to make the lesson reappear when it is relevant.
At one point, I discovered that the live product had been serving code from a branch the project no longer used.
It remained available. Nothing visibly crashed. But the fixes and improvements I believed had reached it had gone nowhere.
The technical correction was straightforward. The important failure was operational.
I had no mechanism comparing what I believed was deployed with what was actually running, so the gap could grow without producing a single symptom. My habit was to inspect the work I had completed, not the system a user was actually using.
That led to one of the rules I now trust most:
Documents describe intent. Code describes intended behaviour. Running systems describe reality.
When those sources disagree, reading more documents does not resolve the disagreement. You have to inspect and measure the running system.
Some of the manual scaffolding I built became unnecessary as the tools improved.
Context that I once pasted manually could be loaded from a project file. The AI could read the repository, run commands, inspect the database and check the browser directly. Work that had to be carried between two tools could increasingly happen within one controlled workflow. Independent agents could examine different parts of a problem and return their findings.
I learned that working with a fast moving tool also means noticing when one of your own workarounds has become obsolete.
That is harder than it sounds. A workaround that solved a real problem and took effort to create is easy to keep using long after it has stopped earning its place.
But the important distinction was this:
The scaffolding became obsolete. The discipline did not.
Research before building. Make the constraints explicit. Test one change at a time. Measure the result. Ask what could break. Do not accept a reassuring label as evidence. Those principles survived every change in the tools because they were never really about the tools.
They were about the danger of believing something that had not been checked.
Before building a significant feature, I now ask for a separate review whose only purpose is to make the plan fail.
Not to improve the writing. Not to make the solution more elegant. To find missing controls, unsupported assumptions, unintended consequences and dependencies the plan has quietly dropped.
One review found that safety checks described in the plan would not actually run on the version of the code being built. We could have continued believing that the safety net was active while all the work happened outside it.
Then I gave the same plan to a completely different AI and asked it to attack the first review as well as the plan.
It returned 31 findings. Two were wrong. Several were already handled. One exposed a genuine design problem in how company access would work. I checked the finding against the identity provider’s own documentation and changed the design before it became expensive to reverse.
Different systems have different blind spots. An independent second model is not a substitute for expert review, but it is inexpensive insurance against accepting the first model’s confidence.
That process also taught me not to judge a review by whether every finding is correct. A review can contain several false alarms and still be valuable if one finding prevents a serious mistake.
My early method of understanding the impact of a change was to make the change and see what happened.
I now ask three questions first:
The third question became the most useful.
Small systems make structural decisions feel cheap. Changing a database with test records is easy. Changing the same structure after it carries real organisational data is a programme of work.
Impact analysis written after the decision often becomes a justification for what has already been chosen. I needed it before the decision, when it could still change the answer.
The project became most useful when it started telling me that my ideas were wrong.
I built a feature intended to detect whether someone was reading from a script by tracking eye movement. Across the available sessions, it fired 51 times while people were sitting still. It was detecting noise in its own measurement rather than human behaviour.
A companion rule intended to detect somebody looking away never fired in 255 scoring windows. Its threshold required an eye to move further than the eye socket physically allows.
Another model was meant to identify phones, screens and papers visible to the second camera. It returned zero detections across 132 samples, including frames where another model could clearly see a face. It was blind at that angle.
The tempting response was to keep tuning. Instead, the features were removed.
That decision mattered more than building them. A weak signal does not become safe merely because it has a sophisticated name or sits inside an AI product.
The build status page publishes those failures because a system that only records successful experiments teaches the next person nothing.
The most worrying AI failures were not obvious errors. They were plausible controls that looked complete.
A function had a name suggesting that it checked ownership, but performed no meaningful ownership check. A field named “encrypted” defaulted to true whether or not encryption had actually occurred. A shared component was created to stop several screens from disagreeing, while each screen continued to maintain its own copy of the data.
Each looked reassuring in a summary. The names created confidence before the behaviour had been verified.
I found the same pattern at a more serious level.
A privacy commitment said that records would be deleted after a defined period. The records carried an expiry date. The design described the deletion process. But nothing acted on the expiry date. The promise existed in the documentation and on the website, not in the running system.
A safety gate intended to prevent an ordinary change from altering the live database was described as though it existed. The setting required to make it real had never been enabled.
Nothing in either description sounded unfinished.
That taught me to separate four claims that are often collapsed into one:
They are not synonyms.
Code being accepted proves that the code was accepted. It does not prove that the intended outcome occurred.
This sounds obvious, but it was one of the hardest habits to build because both people and AI are naturally drawn to completion signals. The file changed. The test suite passed. The pull request merged. The task was marked complete.
The user does not experience any of those things. The user experiences the running system.
The practical rule became: after a meaningful change, verify the actual behaviour through the same surface a user would encounter. If the change concerns stored data, query the data. If it concerns privacy, try to violate the privacy rule. If it concerns deletion, prove that the record disappeared and that the failure path is visible when it does not.
I now treat “built” and “proven to work” as separate statements, and I try to say which one I mean.
Verlinko is intended to create an evidence record around identity continuity. Building it forced me to think carefully about what makes evidence credible.
A record that can report only success is not evidence. It is marketing generated by a system.
The difficult decisions were about what the system was permitted to assert. “No person reviewed this session” must remain different from “a person approved it.” A comparison that could not complete must not quietly become a failed match. A deletion request must not be recorded as completed merely because it was sent.
The record has to preserve the uncomfortable state as accurately as the successful one.
That principle transfers well beyond hiring technology. Operational dashboards, AI quality scores and transformation reports all become dangerous when the measurement is designed to make the programme look healthy rather than to reveal what is true.
Initially, I assumed AI cost was driven mainly by the size of the task. A large feature would cost more. A small correction would cost less.
What I observed was different.
Much of the effort was determined by how difficult the project was to understand. The AI would search for the relevant files, read an outdated explanation, begin from a slightly wrong picture and then need to be corrected. The useful work began only after orientation.
Once the project had a clearer structure, a reliable source hierarchy, current decision records and rules loaded at the start of every session, the number of corrections fell materially.
The model had not become more intelligent. The conditions around it had improved.
That made me think differently about technical debt and, more broadly, context debt.
A messy environment always had a cost. It was usually absorbed through employee frustration, slower onboarding and rework. With AI, part of that cost can appear repeatedly in usage, latency and failed iterations. Every session may have to purchase the same understanding again.
The relevant measure is therefore not simply total AI usage. It is the cost of a useful loop: how much effort is required to move from a question to an output that can be trusted, and how many corrections are required before it becomes usable.
Quality came from iteration. When each loop became cheaper and more reliable, I could afford to challenge the work again rather than accept the first plausible result.
The same context problem exists across large operations, even when no code is involved.
It is the procedure that was updated in one location but not another. The exception explained in an email but never added to the knowledge base. The policy that differs by market. The process map that describes how work was designed while experienced employees follow a different process to make it function.
A new employee spends months learning which source to trust. An AI agent can encounter the same confusion in seconds and produce the wrong answer with complete confidence.
Consider an AI assistant supporting a customer refund. It may need the current policy, the customer history, the contractual exception, the authority limit, the regulator’s requirements and the record of what has already been promised. If those sources disagree, the model does not remove the ambiguity. It makes a prediction from it.
The same applies to claims, finance operations, HR services, compliance reviews and any back office process where the real procedure lives partly in systems, partly in documents and partly in people’s heads.
That is why I no longer see enterprise AI transformation primarily as a model selection exercise.
The larger work is:
AI can compress the distance between an operating problem and a working system. It cannot decide which tradeoffs an organisation should accept. It cannot turn contradictory policies into a coherent operating model merely by reading more of them. It cannot own the consequences of a decision.
Those remain transformation and leadership responsibilities.
Verlinko did not turn me into a software engineer, and I do not want to make that claim.
It did give me practical experience of moving from a problem hypothesis to a functioning system. I learned to work with AI as a researcher, implementer, reviewer and adversary. I learned to make constraints explicit, test assumptions, distinguish implementation from evidence, remove features that failed measurement and identify where a confident technical explanation did not match reality.
Most importantly, I learned enough about the build side to ask better questions earlier.
That matters because many AI programmes do not fail because the model is incapable. They fail because the problem was framed badly, the process was never made explicit, the data carried contradictions, the exception path was ignored or nobody established what evidence would prove that the system was working.
My value is not that I can write code faster than an engineer. I cannot.
It is that I can combine operating experience with enough practical build understanding to keep the business problem, the user journey, the controls and the running system connected.
That is the capability I wanted to test when I began Verlinko. The prototype is one result. The more important result is the way I now think about building with AI.
Verlinko is an independent portfolio project. It is not currently offered as a commercial service.
The public site shows the architecture, the evidence behind the problem, measured results from the prototype and the controls that remain incomplete. Some parts work. Some are still being strengthened. Several ideas were built and removed because the measurements did not support them.
That unfinished state is part of the point.
If I published only the polished result, I would hide the most useful part of the work: how often a reasonable idea failed contact with the running system, and how much discipline was required to discover that before asking someone else to trust it.
If you work in security, privacy, hiring technology, AI governance, CX or back office transformation and something in the project looks wrong, I would genuinely rather hear it than not.
You can explore how Verlinko works, review the evidence behind the problem, inspect the current build status and known gaps, or try the local matching demonstration.
AI disclosure. I used AI extensively to build Verlinko and to help structure this article. The product decisions, operating judgements, experiences and opinions are mine. AI generated much of the implementation, but it did not own the consequences of any decision. I did.
Questions about the project: hello@verlinko.com