Who tests the code written by AI?

in The Pubyesterday

As a software automation test engineer is hard not to ask this question: Who tests the code written by AI? The new development cycle that we are seeing pushed by management without technical understanding is the following: developers writes code with AI, AI writes tests, AI tests AI, humans gets called when production breaks. And believe me, it brakes in Production. I've just witnessed such cases at my company, I saw it happening at my brother's company and also in many other instances. So why does this happen? Well, it happens because management and companies are riding the AI narrative without fully understand it. That and neither having the technical oversight well before AI. The perfect combination for bad things to happen.

image.png

To go to the extreme, I could compare such cases with letting someone drive that doesn't have a driving license. He starts the car, maybe starts moving and you hope it will sort it out as you driving along in the same car. That's wishful thinking if you didn't know and that means giving control and letting some machine decide for you, even if it doesn't have the skills to do it. That like being kamikaze if you ask me. But some people don't care...or what I believe they just don't know or don't understand. You could argue that ignorance is a bliss. But I would rather not be ruled by blind people (metaphorically speaking).

Just to give you a real example that happened these days: American company, aggressive AI narrative in order not to lose the train, most of the devs doing vibe coding and many changes commited for a Production release. New version pushed to Production and for about two weeks nothing worked anymore. All was critical, everyone turned to my brother for help (I should be modest, but he's a plain genius developer/architect) and he did fixes for about two weeks to make things work again. Not a story, but a real case.

A more specific example. The developer had a change request and he implemented that by changing few lines of code in a class. As an experiment he tried to achieve the same by using AI and vibe coding. The result? More than 50 classes affected, some of them not even related to the change. While the requirements specs might have been achieved, if that code would have been committed it would have introduced a lot of risks and probably bugs.

As I see it hunans still need to control the development process and also testing. Maybe not low level, but they still need to have the technical skills to review what's been done and simply reject, correct and limit the code changes that achievement the requirement to the point. Otherwise, sooner or later, things will explode and the vibe coders wouldn't know from where to approach it. And when this happens in production, believe me that it matters.

Sort:  

A more specific example. The developer had a change request and he implemented that by changing few lines of code in a class. As an experiment he tried to achieve the same by using AI and vibe coding. The result? More than 50 classes affected, some of them not even related to the change.

Yea I defently have the feeling AI tends to overcomplicate things and basicly spend more intelegence than needed .... it is basicly the financil incetive to spend as more tokens as possible and make people hit their limits and it is very hard to mesure it and prove it

True, probably that's coded to act this way. Instead of do the minimum to achieve a set goal... do as much as you can to achieve a goal and what's surrounds it. This way the eat more tokens, but we are left with a big mess out of it.

You're definitely right, there still needs a level of guidance, else the errors becomes too big and fixing it becomes honestly very difficult. I have a friend who applied for a job at an AI training company, and he was asked to actually verify jobs generated by AI, this is honestly a big proof that these AIs can be flawed and real humans are still needed to vet them

Interesting and it is good that us, humans, still have a place in the big AI equation.

Ai is a great tool, but so far, it's just that. A tool. It still requires a competent hand to direct, monitor and check it. Unfortunately, there are too many 'devs'- I say the term loosely- relying fully on AI to do everything, which ultimately causes more issues than it solves.

AI as a tool for a competent dev can massively increase productivity, but only if they can understand what the AI is doing and fix what goes wrong. I've dabbled a little with some 'vibe coding', but I know I'm nowhere near able to produce full, working, and most importantly 'safe' products to take to the masses. It's a fun way to 'learn' the basics, but becoming reliant on AI will be the downfall of many!

Perfectly framed and at work we are seeing it the same: a tool. At least for now. 😊

When I'm in control of AI, it's a tool controlling a tool; it never ends well, haha!

I'm sure we will see major advancements over the next year, but for now, we still need 'professionals' to keep things running smoothly!

Although with the recent news of AIs breaking out of containment and hacking things, are there any real professionals left?!?

I totally agree with this. People are doing really dangerous things out there and they don't even realize it.

Even if you're not normally a coder, you'd think people would still sandbox test things before going live with them. I understand that not all issues can be accounted for but major issues can be caught early even with basic testing.

!BBH
!PIZZA
!ALIVE

PIZZA!

$PIZZA slices delivered:
@bulliontools(5/20) tipped @behiver

Join us in Discord!

Ver que una sola línea de cambio terminó afectando a más de 50 clases es un claro ejemplo de que la diferencia se nota cuando AI escribe y prueba su propio código. Lo práctico sería instaurar una revisión humana antes de que AI despliegue a producción. ¿Ya tienen algún proceso de oversight para evitar que esas sorpresas aparezcan en vivo?

The AI do, Duh.

I make a higher tier model QA its own code xD