Come Claude ha lavorato per me mentre dormivo: reverse engineering di un protocollo industriale How Claude worked for me while I slept: reverse engineering an industrial protocol

Davanti a un lavoraccio di interpretazione di un protocollo di comunicazione industriale proprietario, ho pensato di delegare il lavoro a Claude. Ecco come, e perché ha funzionato. Faced with the tedious task of interpreting a proprietary industrial communication protocol, I decided to delegate the work to Claude. Here is how, and why it worked.

TL;DR Un impianto industriale critico gira 24 ore su 24. Il fornitore del software di controllo è sparito, nessuno ha i sorgenti né la documentazione, e il proprietario ha bisogno di funzionalità nuove. Riscrivere tutto da zero “e basta” vorrebbe dire installare il nuovo software, e installarlo vuol dire fermare l’impianto. Non è un’opzione. L’unica strada percorribile che rimane è costruire l’impianto nuovo sopra quello vecchio, in modo da poterlo sostituire gradualmente e senza downtime. Purtroppo, per farlo, serve fare reverse engineering di un protocollo di comunicazione sconosciuto e proprietario. Ho raccolto le catture di rete, ho preparato il contesto, ho messo Claude Code in modalità autonoma e sono andato a dormire. Sei ore e mezza dopo il livello di trasporto era verificato bit per bit su oltre mezzo milione di pacchetti, ventuno comandi su ventiquattro erano decodificati, il layer applicativo era aperto, e c’era una lista prioritizzata di cosa mancava e di come ottenerlo. Giornate di lavoro, risolte nel sonno. Le settimane successive sono state in discesa, e con la stessa tecnica sono arrivato a decifrare tutto. Mesi di lavoro manuale ridotti a qualche sessione di raccolta dati mirata.

La scena che non dimenticherò

Mi sono svegliato verso le sette. Curioso di vedere come se la fosse cavata, la prima cosa che ho fatto è stata guardare il processo di Claude Code.

Concluso. C’era un file Markdown nuovo. Specifica tecnica consolidata. Aperto.

Migliaia di parole e varie tabelle, con in cima una sezione intitolata “esito sintetico della verifica”: 532.189 frame analizzati, 532.189 con checksum valido, zero errori. Un catalogo completo dei comandi decifrato. Il layer applicativo, che il giorno prima erano solo byte pseudocasuali, spacchettato campo per campo. In fondo al file un inventario esplicito di tutto ciò che restava da decodificare, in ordine di priorità, con per ognuno la singola cattura da fare per chiuderlo.

Non un vago “fatto”. Una to-do list ingegneristica, scritta da qualcuno che sembrava aver capito il problema meglio di come l’avevo lasciato la sera prima. In sei ore e mezza di computazione. Inutile dire che mi abbia lasciato di stucco. Immaginate cosa significa analizzare mezzo milione di pacchetti di rete catturato su Wireshark, ognuno da 6/20 byte, senza un preciso significato, senza codice. Ci avrei messo mesi, sicuramente, e invece mentre dormivo era praticamente risolto. Ovviamente non completamente, ma il 90% del lavoro era concluso e le settimane successive sono state un perfezionamento del processo, filato liscio, senza intoppi né difficoltà particolari.

La parte interessante è perché, in questo caso preciso, ha funzionato così bene. Capirlo ha cambiato il modo in cui lavoro.

Il problema: una fabbrica che non si può spegnere

Per motivi ovvi di segreto aziendale, userò un caso fittizio con le stesse caratteristiche del caso reale.

Poniamo quindi che esista un impianto di smistamento industriale, a un grande sistema di convogliatori che instradano contenitori da un punto all’altro di uno stabilimento, attraverso decine di linee, scambi motorizzati, attuatori pneumatici, sensori, stazioni di carico e scarico. Un PC centrale fa da cervello: interroga di continuo i gateway di ogni linea, manda comandi (muovi quello scambio, instrada quel contenitore, attiva quell’attuatore), e ricostruisce in tempo reale lo stato e la topologia dell’intero impianto.

Il sistema è critico e gira ventiquattro ore al giorno, sette giorni su sette. Se si ferma, si ferma la produzione, e ogni ora di fermo è un costo che si misura in migliaia di euro più la catena di disservizi a valle.

I problemi, in fila:

  1. Il fornitore del software non c’è più. Lo sviluppatore originale è uscito di scena senza lasciare né i sorgenti né uno straccio di documentazione. Il software di controllo è di fatto una scatola nera, e così anche il modo con cui comunica alle macchine e ai singoli firmware.

  2. Il proprietario ha bisogno di funzionalità nuove. Integrazioni, logiche di instradamento diverse, reportistica, controlli che il vecchio software non offre e che nessuno può aggiungere, perché nessuno può più toccare quel codice. Senza entrare nel dettaglio, “sono da fare modifiche al sistema”.

  3. Non si può riscrivere tutto, buttare il software vecchio e installare quello nuovo. Questo perché tutte le schede delle stazioni, delle linee etc. sono anch’esse programmate e comunicano con il software vecchio. Riscrivere tutto in questo senso significa per forza cambiare anche le schede. La sostituzione fisica richiederebbe un tempo inaccettabile. Risulta quindi necessario creare un software nuovo “sopra” quello esistente, che coesista e parli la stessa lingua.

Il codice di per sé è classica programmazione di sistemi industriali. Di per sé neanche eccessivamente complicato. Il problema risiede nei vincoli: continuità operativa assoluta, zero documentazione, zero sorgenti.

La strada obbligata: costruire sopra il sistema, senza toccare sotto

Se non puoi sostituire il sistema, e non puoi spegnerlo, devi convivere con lui. Devi imparare la sua lingua.

Il PC centrale e le macchine si parlano in continuazione su rete. Quel dialogo è un protocollo binario proprietario: un flusso di pacchetti in un formato ignoto. Se riesco a intercettarlo, a capirne la sintassi e la semantica, posso fare due cose enormi:

  • Leggere lo stato dell’impianto in tempo reale, ricostruendolo dagli stessi messaggi che le macchine già si scambiano, senza chiedere niente a nessuno.
  • Scrivere, cioè comporre io stesso comandi validi e iniettarli, parlando alle macchine esattamente come fa il software originale, al punto che le macchine non distinguono i miei comandi dai suoi.

A quel punto il software nuovo non sostituisce il vecchio, ci convive, lo osserva, e agisce al suo fianco. Il software nuovo è come una scatola che incapsula il programma esistente, ci aggiunge le funzionalità e le spedisce al sistema interno parlando la sua lingua.

È elegante. Ed è, come ogni cosa elegante, parecchio più difficile di come suona.

Reverse engineering di un protocollo, in parole povere. Immagina di origliare una conversazione fittissima tra due persone che parlano una lingua che non conosci, di cui non esiste vocabolario, e che non puoi interrogare. Non puoi chiedere “scusa, cosa significa questa parola?”. Puoi solo ascoltare migliaia di scambi, osservare cosa succede nel mondo mentre certe frasi vengono pronunciate, e da lì ricostruire grammatica e significato. Con una differenza crudele rispetto a una lingua umana: qui un singolo bit sbagliato e la frase non vuol dire “andiamo a sinistra”, vuol dire niente. O peggio, vuol dire il contrario.

La materia prima: catture e ground truth

Mi sono piazzato nel PC centrale, l’unico nodo che parla con tutti i gateway contemporaneamente (e per fortuna che il sistema era a stella e non a grafo interconnesso). Lì ho registrato il traffico con gli strumenti standard di analisi di rete. Pacchetto per pacchetto, il dialogo grezzo tra il cervello e le macchine. Più di quaranta sessioni di cattura distinte, ciascuna mirata a un’azione precisa.

Il secondo è il punto cruciale: la verità di riferimento. Il software originale, per quanto chiuso, ha un’interfaccia che mostra un log degli eventi. Una finestra dove scorre, in linguaggio “circa” umano, cosa sta facendo l’impianto. “Motore 1 in posizione 2.” “Attuatore pneumatico: aspirazione.” “Scambio instradato sul ramo 3.” Quelle righe sono oro. Sono la traduzione ufficiale di ciò che, sul filo, è solo una sequenza di byte.

Così ho lavorato come un linguista da campo. Premevo un pulsante nel software (muovi quello scambio) e contemporaneamente registravo due cose: i byte che passavano in rete, e lo screenshot del log che diceva, in italiano, cosa era appena successo. L’idea, ovviamente, è che il log è la verità, non la cattura. Non si adatta mai la decodifica al log; si decodifica, e poi si controlla contro il log. Un po’ come i sistemi di machine learning, che senza il validation set non fanno altro che overfitting sui dati, senza imparare nulla.

Quindi dopo aver catturato i pacchetti di rete e screenshottato i log di sistema, ho suddiviso i dati nell’equivalente di “train test” e “validation test”. Il train test serviva a permettere di fare ipotesi sull’associazione corretta tra byte e righe di log, ipotizzare la struttura del binario, verificare l’esistenza di codici di controllo errori così via. Il validation test serviva a fare test alla cieca. A partire dai byte, predire cosa avrei dovuto trovare sui log, e viceversa, poi verificare. Una predizione azzeccata alla cieca vale dieci letture fatte con la risposta già davanti. Uno studente che ha capito contro uno che ha copiato il compito.

Il muro della sintassi

Ogni protocollo ha strati. Il più basso è il trasporto: non la semantica, ma la sintassi, le regole grammaticali, come impacchetti qualunque cosa tu dica. Dove inizia un messaggio, dove finisce, come ti accorgi se è arrivato corrotto.

Questo strato è ostico ma ha una qualità meravigliosa: è deterministico, senza interpretazione. E quando lo indovini lo verifichi velocemente su milioni di casi.

Ho ricostruito un piccolo decoder (niente librerie esotiche, solo Python e uno strumento standard di analisi pacchetti) e ho fatto letteramente bruteforce delle possibili codifiche: dei marcatori che delimitano inizio e fine di ogni messaggio, un meccanismo di escape per quando quei marcatori compaiono per caso dentro i dati veri, e in coda un checksum a 16 bit a proteggere l’integrità.

Quando mi sembrava di aver azzeccato l’algoritmo, ho lanciato il decoder su tutte le catture. Oltre mezzo milione di pacchetti. Il decoder ricalcolava il checksum di ognuno e lo confrontava con quello sul filo. Zero errori. Mezzo milione di frame, mezzo milione di checksum giusti.

A quel punto sapevo separare con precisione assoluta i singoli messaggi dal flusso, distinguere chi parla (PC verso macchina, o macchina verso PC), leggere l’indirizzo del dispositivo destinatario. La busta era trasparente. Restavano le lettere dentro le buste, e lì il gioco cambia completamente.

Il vero muro: la semantica

La sintassi è meccanica. La semantica è interpretazione, ed è dove si consumano i giorni.

Sapere che un messaggio è fatto bene non ti dice cosa significa. Vedevo passare comandi come “12 01 02” o “80 84 00” e dovevo trasformarli in frasi di senso compiuto: questo è “muovi il motore 1 in posizione 2”, quest’altro è “attiva la pressione pneumatica”. Ogni mappatura del genere richiede:

  • isolare un’azione singola e pulita nella cattura,
  • allinearla al millisecondo con la riga corrispondente nel log degli eventi,
  • formulare un’ipotesi sul significato dei byte,
  • cercarne conferma in altre catture, su altre linee, con altri dispositivi,
  • assegnare un livello di confidenza onesto.

Su quest’ultimo punto mi ero dato una scala a tre gradini, perché in un lavoro così è fin troppo facile prendere un’ipotesi suggestiva per un fatto:

  • Fortemente probabile: c’è la verità testuale del log e/o una predizione cieca andata a segno.
  • Probabile: coerente su almeno due osservazioni indipendenti, senza test cieco.
  • Ipotesi: una sola osservazione, o un’inferenza indiretta. Da maneggiare con i guanti.

Il problema è che questo lavoro è lento, ripetitivo e seriale. Ogni comando è una piccola indagine. Moltiplica per decine di comando, decine di tipi di record che le macchine rimandano indietro, decine di tipologie diverse di dispositivo, i vari messaggi di errore etc.

Impostazione del lavoro con Claude Code

“Santa pigrizia”, oserei dire. Non avevo nessuna voglia di fare sto lavoraccio a mano, e allora mi sono inventato un modo di quantomeno provare a delegarlo alle macchine. Ho costruito intorno a Claude Code un ambiente in cui non poteva permettersi di inventare.

  1. Un file di contesto preciso. Non “decodifica il protocollo”, ma un documento che spiegava cos’è il sistema, com’è fatta una cattura, dove sta la verità di riferimento, quali convenzioni rispettare, cosa era già stato stabilito e con quale confidenza. Un’AI brillante con un contesto vago produce sciocchezze plausibili; la stessa AI con un contesto chirurgico produce lavoro da professionista.
  2. Strumenti deterministici come oracolo. Il decoder aveva i permessi di esecuzione automatica, strumento in mano a Claude: decodificare una cattura, contare i frame, verificare i checksum, estrarre i comandi allineati all’orario.
  3. Tutte le catture e tutte le verità di riferimento: gli screenshot dei log, le note delle azioni. Materiale grezzo abbondante, così che ogni ipotesi avesse contro cosa essere messa alla prova.
  4. Un harness rigido, esplicito e non negoziabile. Il log è verità, non ci si adatta mai. Allinea i comandi all’orario assoluto, al secondo, prima di affermare una mappatura. Tagga ogni affermazione con il suo livello di confidenza. Lancia subagent per pianificazione, esecuzione e controllo. “Le catture senza log sono test set”. Predici e verifica, non adattare mai il modello ai dati di test. Non modificare i documenti esistenti, aggiungi in file nuovi.

Poi ho dato l’obiettivo, ho attivato la modalità autonoma, ho controllato che il primo passo girasse, e sono andato a letto. Aspettative: caute. Mi ero detto che, nel migliore dei casi, avrei trovato qualche progresso e una bella confusione da riordinare.

Il trucco non è “lasciare lavorare l’AI da sola”. È costruirle intorno un mondo in cui l’unico modo di andare avanti è essere rigorosi. Un oracolo, una verità di riferimento, e regole che rendono l’onestà più facile della fantasia.

La mattina dopo

Sei ore e mezza di esecuzione autonoma. Questo è ciò che ho trovato nella cartella, e che ho poi verificato di persona prima di crederci.

Il livello di trasporto, verificato bit per bit su tutto. Non più “funziona sulle catture che ho guardato io”, ma su tutte le catture nuove. zero errori di checksum, zero anomalie sui contatori di sequenza. La fondazione, certificata.

Il catalogo dei comandi, quasi completo. Dei ventiquattro tipi osservati nell’intero corpus, ventuno decodificati, la maggior parte al livello di confidenza più alto, con la verità del log a sostegno. Movimento dei motori, instradamento degli scambi, attuatori pneumatici, modalità dei dispositivi, reset, controlli dell’interfaccia operatore. Ognuno con la sua sintassi, i suoi parametri, le sue varianti tra dispositivi diversi. Tre soli comandi rimasti davvero oscuri, e identificati con precisione.

Il livello applicativo, aperto. Quella parte che mi era sembrata rumore (il formato con cui il sistema descrive un’operazione complessa, tipo “crea un nuovo instradamento per il contenitore X verso la destinazione Y”) si è rivelata una struttura ordinata a campi etichettati. Spacchettata e verificata su più di una dozzina di esempi reali, con dentro perfino il timestamp e la data in chiaro. Zero byte residui non spiegati.

Oltre a questo, una specifica consolidata, unica, implementabile: tutti i round di lavoro sparsi fusi in un documento coerente, con le implicazioni esplicite per chi dovrà scrivere il nuovo software centrale.

La ciliegina sulla torta, un inventario esplicito di ciò che resta da fare, in ordine di valore. Non un trionfalistico “completato”, ma una lista lucida: “questo record di notifica d’arrivo non è ancora decodificato campo per campo, serve per tracciare l’arrivo del contenitore end-to-end, e per chiuderlo fai questa specifica cattura”. Aveva capito il problema abbastanza a fondo da dirmi quali esperimenti progettare per finire il lavoro.

Mi ha davvero stupito come non mai. Un conto è creare un frontend, con migliaia di template su cui è stata addestrata, un altro è lavorare per più di 6 ore di fila su un problema ignoto e arrivare a questo livello di precisione di decodicfica, risparmiandomi mesi di lavoro in una notte.

Cosa non ha fatto

Ovviamente il caro vecchio Claude non avrebbe mai potuto fare tutto da solo, non è onnipotente.

  • Tutto il lavoro poggia sulle catture e sui log che ho raccolto io, sul campo, con il metodo giusto e nel punto di rete giusto. Garbage in, garbage out: senza materia prima pulita non ci sarebbe stata alcuna decodifica, solo fantasie ben formattate. Il lavoro umano insostituibile è venuto prima.
  • Non ha chiuso ciò che richiede nuova evidenza. Lo scenario d’errore non era in nessuna cattura (il lavoro poi è continuato in modo simile per settimane), e nessun ragionamento può decodificarlo dal nulla.
  • Non ha deciso cosa contasse. Sono stato io a inquadrare il problema, a scegliere quali azioni catturare, a definire l’harness e a verificare le sue affermazioni prima di fidarmi.

In tutto questo ho anche avuto un bel po’ di fortuna a trovarmi i messaggi in chiaro, ad avere una topologia ad albero, ad avere dei log (per quanto un po’ scarni) come ground truth etc… Rimane il fatto che il mi ha stupito, e che mi ha risparmiato mesi di lavoro, facendo le cose al posto mio letteralmente mentre dormivo.

Se il racconto ti suona familiare (il sistema critico, la scatola nera, la richiesta di evoluzione che sembra impossibile) parliamone. La via d’uscita forse esiste. E sì: una parte del lavoro, probabilmente, la farò mentre dormo.

TL;DR A critical industrial plant runs 24 hours a day. The control software supplier has vanished, no one has the source code or the documentation, and the owner needs new features. Simply rewriting everything from scratch would mean installing the new software, and installing it means stopping the plant. It’s not an option. The only viable path left is to build the new plant on top of the old one, so it can be replaced gradually and without downtime. Unfortunately, to do this, we need to reverse engineer an unknown and proprietary communication protocol. I collected the network captures, prepared the context, put Claude Code in autonomous mode, and went to sleep. Six and a half hours later, the transport layer was verified bit by bit on over half a million packets, twenty-one commands out of twenty-four were decoded, the application layer was laid bare, and there was a prioritized list of what was missing and how to get it. Days of work, resolved in my sleep. The following weeks were downhill, and with the same technique, I managed to decipher everything. Months of manual work reduced to a few targeted data collection sessions.

The scene I won’t forget

I woke up around seven. Curious to see how it had fared, the first thing I did was look at the Claude Code process.

Finished. There was a new Markdown file. Consolidated technical specification. Opened it.

Thousands of words and various tables, with a section at the top entitled “synthetic verification outcome”: 532,189 frames analyzed, 532,189 with valid checksum, zero errors. A complete catalog of deciphered commands. The application layer, which the day before were just pseudorandom bytes, unpacked field by field. At the bottom of the file, an explicit inventory of everything left to decode, in order of priority, with the single capture to be made for each one to close it.

Not a vague “done”. An engineering to-do list, written by someone who seemed to have understood the problem better than how I had left it the night before. In six and a half hours of computation. Needless to say, it left me dumbfounded. Imagine what it means to analyze half a million network packets captured on Wireshark, each 6 to 20 bytes long, without a precise meaning, without code. It would have taken me months, surely, and instead, while I slept, it was practically resolved. Obviously not completely, but 90% of the work was finished and the following weeks were a refinement of the process, which ran smoothly, without hitches or particular difficulties.

The interesting part is why, in this specific case, it worked so well. Understanding it changed the way I work.

The problem: a factory that cannot be turned off

For obvious reasons of corporate secrecy, I will use a fictitious case with the same characteristics as the real one.

Let’s suppose then that there is an industrial sorting plant, a large system of conveyors that route containers from one point of a factory to another, through dozens of lines, motorized switches, pneumatic actuators, sensors, and loading and unloading stations. A central PC acts as the brain: it continuously polls the gateways of each line, sends commands (move that switch, route that container, activate that actuator), and reconstructs the status and topology of the entire plant in real-time.

The system is critical and runs twenty-four hours a day, seven days a week. If it stops, production stops, and every hour of downtime is a cost measured in thousands of euros plus the chain of downstream disruptions.

The problems, in order:

  1. The software supplier is no longer around. The original developer exited the scene without leaving either the source code or a shred of documentation. The control software is effectively a black box, and so is the way it communicates with the machines and individual firmwares.

  2. The owner needs new features. Integrations, different routing logic, reporting, controls that the old software doesn’t offer and that no one can add, because no one can touch that code anymore. Without going into detail, “modifications need to be made to the system”.

  3. You can’t rewrite everything, throw away the old software, and install the new one. This is because all the boards of the stations, lines, etc. are also programmed and communicate with the old software. Rewriting everything in this sense necessarily means changing the boards too. The physical replacement would require an unacceptable amount of time. It is therefore necessary to create a new software “on top” of the existing one, which coexists and speaks the same language.

The code itself is classic industrial systems programming. In itself, not even excessively complicated. The problem lies in the constraints: absolute operational continuity, zero documentation, zero source code.

The forced path: build on top of the system, without touching underneath

If you can’t replace the system, and you can’t turn it off, you have to live with it. You have to learn its language.

The central PC and the machines talk continuously over the network. That dialogue is a proprietary binary protocol: a stream of packets in an unknown format. If I can intercept it, understand its syntax and semantics, I can do two huge things:

  • Read the plant’s status in real-time, reconstructing it from the very same messages the machines are already exchanging, without asking anyone for anything.
  • Write, meaning compose valid commands myself and inject them, talking to the machines exactly as the original software does, to the point that the machines cannot distinguish my commands from its own.

At that point, the new software doesn’t replace the old one, it coexists with it, observes it, and acts alongside it. The new software is like a box that encapsulates the existing program, adds features to it, and sends them to the internal system speaking its language.

It’s elegant. And it is, like everything elegant, quite a bit harder than it sounds.

Reverse engineering a protocol, simply put. Imagine eavesdropping on a very dense conversation between two people speaking a language you don’t know, for which there is no dictionary, and whom you cannot interrogate. You can’t ask “excuse me, what does this word mean?”. You can only listen to thousands of exchanges, observe what happens in the world while certain phrases are pronounced, and from there reconstruct grammar and meaning. With a cruel difference compared to a human language: here a single wrong bit and the phrase doesn’t mean “let’s go left”, it means nothing. Or worse, it means the opposite.

The raw material: captures and ground truth

I positioned myself at the central PC, the only node that talks to all the gateways simultaneously (and luckily the system was a star topology and not an interconnected graph). There I recorded the traffic with standard network analysis tools. Packet by packet, the raw dialogue between the brain and the machines. More than forty distinct capture sessions, each targeted at a specific action.

The second is the crucial point: the reference truth. The original software, however closed, has an interface that shows an event log. A window where it scrolls, in “roughly” human language, what the plant is doing. “Motor 1 in position 2.” “Pneumatic actuator: suction.” “Switch routed on branch 3.” Those lines are gold. They are the official translation of what, on the wire, is just a sequence of bytes.

So I worked like a field linguist. I pressed a button in the software (move that switch) and simultaneously recorded two things: the bytes passing on the network, and the screenshot of the log saying, in Italian, what had just happened. The idea, obviously, is that the log is the truth, not the capture. You never adapt the decoding to the log; you decode, and then you check against the log. A bit like machine learning systems, which without the validation set do nothing but overfit the data, without learning anything.

So after capturing the network packets and screenshotting the system logs, I divided the data into the equivalent of “train test” and “validation test”. The train test served to allow hypothesizing the correct association between bytes and log lines, hypothesizing the structure of the binary, verifying the existence of error control codes, and so on. The validation test served to do blind tests. Starting from the bytes, predicting what I should have found on the logs, and vice versa, then verifying. One accurate blind prediction is worth ten readings done with the answer already in front of you. A student who has understood versus one who has copied the homework.

The syntax wall

Every protocol has layers. The lowest is transport: not semantics, but syntax, the grammatical rules, how you package whatever you say. Where a message begins, where it ends, how you notice if it arrived corrupted.

This layer is tough but has a wonderful quality: it is deterministic, without interpretation. And when you guess it, you verify it quickly on millions of cases.

I rebuilt a small decoder (no exotic libraries, just Python and a standard packet analysis tool) and literally bruteforced the possible encodings: the markers that delimit the beginning and end of each message, an escape mechanism for when those markers randomly appear inside the real data, and at the tail a 16-bit checksum to protect integrity.

When I felt I had nailed the algorithm, I ran the decoder on all the captures. Over half a million packets. The decoder recalculated the checksum of each one and compared it with the one on the wire. Zero errors. Half a million frames, half a million correct checksums.

At that point, I knew how to separate the individual messages from the stream with absolute precision, distinguish who is talking (PC to machine, or machine to PC), and read the destination device address. The envelope was transparent. What remained were the letters inside the envelopes, and there the game changes completely.

The real wall: semantics

Syntax is mechanical. Semantics is interpretation, and it’s where days are consumed.

Knowing that a message is well-formed doesn’t tell you what it means. I saw commands passing like “12 01 02” or “80 84 00” and had to transform them into meaningful sentences: this is “move motor 1 to position 2”, this other is “activate pneumatic pressure”. Every such mapping requires:

  • isolating a single, clean action in the capture,
  • aligning it to the millisecond with the corresponding line in the event log,
  • formulating a hypothesis on the meaning of the bytes,
  • seeking confirmation in other captures, on other lines, with other devices,
  • assigning an honest confidence level.

On this last point, I had given myself a three-step scale, because in a job like this it’s all too easy to take a suggestive hypothesis for a fact:

  • Highly probable: there is the textual truth of the log and/or a successful blind prediction.
  • Probable: consistent across at least two independent observations, without a blind test.
  • Hypothesis: a single observation, or an indirect inference. To be handled with care.

The problem is that this work is slow, repetitive, and serial. Every command is a small investigation. Multiply by dozens of commands, dozens of record types that the machines send back, dozens of different device types, various error messages, etc.

Setting up the work with Claude Code

“Holy laziness”, I dare say. I had no desire to do this dirty work by hand, and so I invented a way to at least try to delegate it to machines. I built an environment around Claude Code where it couldn’t afford to invent.

  1. A precise context file. Not “decode the protocol”, but a document explaining what the system is, what a capture looks like, where the reference truth lies, what conventions to respect, what had already been established, and with what confidence. A brilliant AI with vague context produces plausible nonsense; the same AI with surgical context produces professional work.
  2. Deterministic tools as an oracle. The decoder had automatic execution permissions, a tool in Claude’s hands: decode a capture, count the frames, verify the checksums, extract the commands aligned with the time.
  3. All captures and all reference truths: the log screenshots, the notes on the actions. Abundant raw material, so that every hypothesis had something to be tested against.
  4. A rigid, explicit, and non-negotiable harness. The log is truth, it is never adapted. Align commands to absolute time, to the second, before affirming a mapping. Tag every assertion with its confidence level. Launch subagents for planning, execution, and control. “Captures without logs are the test set”. Predict and verify, never fit the model to the test data. Do not modify existing documents, append in new files.

Then I gave the objective, activated the autonomous mode, checked that the first step ran, and went to bed. Expectations: cautious. I had told myself that, at best, I would find some progress and a nice mess to tidy up.

The trick is not “letting the AI work alone”. It is building a world around it where the only way forward is to be rigorous. An oracle, a reference truth, and rules that make honesty easier than fantasy.

The morning after

Six and a half hours of autonomous execution. This is what I found in the folder, and what I later personally verified before believing it.

The transport layer, verified bit by bit on everything. No longer “it works on the captures I looked at”, but on all new captures. Zero checksum errors, zero anomalies on the sequence counters. The foundation, certified.

The command catalog, almost complete. Of the twenty-four types observed in the entire corpus, twenty-one decoded, most at the highest confidence level, supported by the log’s truth. Motor movement, switch routing, pneumatic actuators, device modes, resets, operator interface controls. Each with its syntax, its parameters, its variants between different devices. Only three commands remained truly obscure, and precisely identified.

The application layer, opened up. That part that had seemed like noise to me (the format with which the system describes a complex operation, like “create a new routing for container X to destination Y”) turned out to be an orderly structure with labeled fields. Unpacked and verified on more than a dozen real examples, containing even the timestamp and clear text date. Zero unexplained residual bytes.

Beyond this, a consolidated, unique, implementable specification: all the scattered rounds of work merged into a coherent document, with explicit implications for whoever will have to write the new central software.

The icing on the cake, an explicit inventory of what’s left to do, in order of value. Not a triumphant “completed”, but a lucid list: “this arrival notification record is not yet decoded field by field, it’s needed to track the container’s end-to-end arrival, and to close it you need to do this specific capture”. It had understood the problem deeply enough to tell me what experiments to design to finish the job.

It truly amazed me like never before. It’s one thing to create a frontend, with thousands of templates it was trained on, it’s another to work for over 6 hours straight on an unknown problem and arrive at this level of decoding precision, saving me months of work in one night.

What it didn’t do

Obviously, good old Claude could never have done everything alone, it is not omnipotent.

  • All the work rests on the captures and logs that I collected, in the field, with the right method and at the right network point. Garbage in, garbage out: without clean raw material there would have been no decoding, only well-formatted fantasies. The irreplaceable human work came first.
  • It didn’t close what requires new evidence. The error scenario wasn’t in any capture (the work then continued in a similar way for weeks), and no reasoning can decode it from thin air.
  • It didn’t decide what mattered. I was the one who framed the problem, chose which actions to capture, defined the harness, and verified its claims before trusting them.

In all this, I also had a good bit of luck finding messages in clear text, having a tree topology, having logs (however somewhat sparse) as ground truth, etc… The fact remains that it amazed me, and that it saved me months of work, doing things for me literally while I slept.

If the story sounds familiar to you (the critical system, the black box, the request for evolution that seems impossible) let’s talk. The way out perhaps exists. And yes: part of the work, probably, I will do while I sleep.

Note: this article was automatically translated from Italian by AI.