The AI-DX project by Andrea Decarolis (K1FM) is an open-source Python application that can autonomously conduct real SSB QSOs on the KV bands - without an operator sitting behind the radio. It is not a chatbot simulating on-screen traffic. The AI-DX connects to the physical transceiver, listens to the band, recognizes the call, answers by voice, logs the connection to the ADIF, and calls CQ again. The project does not have a corporate background or a team of developers behind it - so far three stars on GitHub and one enthusiast who simply wrote it once and put it out there.
For the radio amateur community, AI-DX is interesting not only as a finished tool, but also as a reference implementation: it shows how modern AI APIs can be integrated with existing radio hardware without the need for a special hardware interface. Anyone interested in understanding how GPT-4o Realtime, WebSocket audio bridge and ADIF logging work together can simply run demo mode on any MacBook and watch what happens.
You will read in the article
K1FM: who is behind it

Author Andrea Decarolis has a brand K1FM. Je vývojár so záujmom o prepojenie moderných AI API s rádioamatérskou praxou. AI-DX nie je jeho jediným projektom v tejto oblasti – spolu s ním je na GitHube aj wfweb, webový server, ktorý sprístupňuje audio a PTT transceiver via WebSocket interface with TLS support. This architecture has the advantage that wfweb can run directly on a computer connected to the radio (for example, a Raspberry Pi in a station), while AI-DX runs on any other machine on the network - even a remote one.
Architecture: what's really going on behind the scenes

AI-DX builds on GPT-4o Realtime API from OpenAI – a special interface designed for real-time audio processing. The whole chain looks like this: wfweb WebSocket receives audio from the receiver (48 kHz PCM16), sends it to GPT-4o Realtime, where speech recognition (STT), language model processing (LLM) and response synthesis (TTS) take place - all on the OpenAI server side. The generated sound is sent back via wfweb to the transmitter and the model simultaneously toggles the PTT via WebSocket.
Dôsledok tohto dizajnu je zaujímavý: na lokálnom počítači nebeží žiadny STT engine, žiadny TTS engine, žiadna detekcia hlasovej aktivity. Additionally, the model does not hold any memory between QSOs - each CQ call starts with an empty context. This is not a bug, but a deliberate decision: it avoids the accumulation of state that could confuse the model during long operating sessions. The computing load is fully transferred to the OpenAI cloud server. This makes it possible to run AI-DX even on a Mac Mini M1 without any GPU. The practical trade-off is obvious: it only runs when you have internet and a valid OpenAI API key with access to the model gpt-4o-realtime-preview.
Kontakty sa sledujú cez function calling: when the model hears the call sign, operator name or QTH, zavolá interný nástroj update_contact() and continuously supplements the record. When the operator says goodbye, the model calls update_contact(closing=true), the connection is written to the ADIF log and AI-DX calls CQ again. No regular expressions, no text parsing - just native language understanding.
Installation and technical requirements
The project requires Python 3.10 to 3.13 and the uv package manager. Installation is simple:
git clone https://github.com/adecarolis/AI-DX && cd AI-DX && uv sync
The project is currently being tested on macOS Apple Silicon (M4). On other platforms, modification may be required - the author openly mentions this in the README. In addition to the Python dependencies, an instance of wfweb must be running on the network to physically control the transceiver. A demo mode is available without wfweb: AI-DX listens to the computer's microphone and responds through the speakers without any radio equipment - great for testing or getting to know the model's behavior.
Configuration: .env file instead of GUI
All settings go through the file .env in the root directory of the project. Required variables are OPENAI_API_KEY, callsign a WFWEB_URL. Optional variables cover Operator Name, QTH, Antenna, Power and Transceiver Type - the model uses these directly in the QSO speech. Frequency and mode are loaded live from wfweb status reports; the exact value is written to the ADIF log VFO v momente ukončenia QSO.
Zásadnou voľbou je štýl operátora, set by the variable OPERATOR_STYLE:
| Style | Behavior |
|---|---|
| CALLING_CQ | Calls CQ every N seconds, conducts regular QSOs |
| CONTESTING | Quick exchange of serial numbers, contest protocol |
| MONITORING | It waits for a direct call, CQ itself does not call, it issues an ID every 5 minutes |
| SWL | Receive only, never transmit |
The CONTESTING mode is interesting from a technical point of view: the model has to handle fast pile-up protocols, short exchanges and the order of call processing. The current capabilities of GPT-4o Realtime in this direction have not yet been systematically documented - this is an open space for experimentation. the duplicate detector (station skip detection) compares call signs using string similarity and prevents the model from blocking others by calling the same station in a loop.
Debug variables like VAD_THRESHOLD (voice recognition threshold, 0.0–1.0), VAD_SILENCE_DURATION (pause to end rap in seconds), CQ_INTERVAL_SEC a CQ_RESTART_DELAY_SEC dávajú operátorovi priamu kontrolu nad rytmom prevádzky. Na rušnom pásme s QRM low VAD_THRESHOLD will cause false activations; on the quiet band, too high a threshold can skip weak stations. Before the first deployment, it is worth calibrating these values in demo mode.
Terminal interface and ADIF log

AI-DX includes a full-featured terminal interface built on the Rich library, updated 10 times per second. It displays S-meter (values from wfweb, range −54dB to +60dB from S9), PTT indicator (RX / VOICE↑ / ON AIR), current frequency and mode, transmit power and SWR, and continuous recording of RX/TX transcripts. An active QSO with status, call sign and QTH is visible on the bottom line.
The ADIF log is saved to logs/contacts.adi. Demo mode creates a separate time-stamped file logs/demo_YYYYMMDD_HHMMSS.adi, so production records are never affected by testing. Each record contains fields QSO_DATE, TIME_ON, CALL, FREQ, FASHION, RST_SENT, RST_RCVD, NAME, QTH a NOTES. The frequency is read live from wfweb at the moment the QSO ends - not a fixed value from the configuration.
Limitations, legal liability and open questions

The author in DISCLAIMER.md explicitly emphasizes: the operator responsible for each broadcast is you, not the model. All legislative obligations applicable to autonomous systems under the regulations of the IARU, FCC and national regulators apply to the licensee. The project has a special warning section in the README, which is a condition for running.
Model gpt-4o-realtime-preview is recommended for strong ability to understand partial call signs, phonetic alphabet and weak HF signals. Cheaper gpt-realtime-1.5 is available as a backup, but the author notes a noticeably worse accuracy in busy airwaves. OpenAI API costs depend on session length and are not negligible for long contest events.
Native integration with an external DX cluster is currently missing, HamQTH callbook alebo LoTW – ADIF log existuje, ale jeho automatické odovzdanie je na používateľovi. Rovnako nie je implementovaná podpora CW; the model understands CW text when the transceiver decodes it and displays it as text, but native decoding of Morse code from audio would require an additional module. These are the spaces where the community could move the project forward.
A more serious practical limitation may be cloud API latency. GPT-4o Realtime usually responds within a few hundred milliseconds after the calling station's rep ends, but with a busy server or unstable internet, the pause before TX may be noticeable. On a contest lane with a dense pile-up, it can give the impression of a slow operator. Value VAD_SILENCE_DURATION=0.6 is the default, but reducing it to 0.4 s can speed up the response at the cost of occasional premature broadcasts.
Regulatory framework: who is responsible?
AI-DX openly relies on the fact that according to the regulations of most countries, including the Slovak Republic and the Czech Republic, a control operator - a holder of a valid amateur license, who actively supervises the operation - must be responsible for each broadcast. The autonomous system does not change this rule. If AI-DX mistakenly answers an illegal call, sends incorrect power or violates the bandplan, the operator, not the model, is responsible. The author states in DISCLAIMER.md that the software is provided "as is" without warranty and the user acknowledges all regulatory obligations.
Praktický dôsledok: AI-DX nie je určený na prevádzku bez dozoru. Realistický prípad použitia je napríklad contestová prevádzka, where the operator watches the terminal and possibly takes over manual control, or experimental testing of SSB pile-up processing without physical presence at the radio, but with the possibility of immediate disconnection. Mode SWL it is the safest from this point of view - it never transmits and serves exclusively for reception and analysis.
How to get involved
The source code is available at github.com/adecarolis/AI-DX pod licenciou MIT. Licencia je čistá: predchádzajúce verzie projektu využívali Hamlib (LGPL v2.1), the current version does not contain Hamlib - PTT control is exclusively via the wfweb WebSocket command setPTT. Záujemcovia o portáciu na Linux alebo Windows môžu otvoriť issue na GitHube – základ je čisto Python a asyncio, takže technická bariéra nie je vysoká. Chyby a návrhy patria do záložky GitHub Issues. Kto má skúsenosť s integráciou Hamlib, N1MM+ or external DX clusters, can contribute with pull requests - the project has a clean modular structure with separate directories for audio, AI, radio, core and UI. Testing in demo mode requires no radio or license - just Python, an OpenAI API key and a microphone.
Source code, installation procedure and DISCLAIMER are available at github.com/adecarolis/AI-DX. Demo mode works without radio hardware.
