← Back to Portfolio

iGlasses & TOKI AI

Developed an AI-powered smart glasses prototype to assist the visually impaired. Integrated a Raspberry Pi 5 and the TOKI AI voice assistant into a compact wearable. I specifically engineered the custom hardware power design and actively developed the device's Python-based AI voice assistant, utilizing speech-to-text, large language model query routing, and automated Geoapify location parsing. Secured 3rd place at Youth Tech Begin and won the Entrepreneurship Potential Award given by the Türkiye Technology Team Foundation (T3 Foundation).

Snapshot Info

Hardware Stack
Raspberry Pi 5 IMX219-83 Stereo Binocular Cameras INMP441 I2S Microphone PCM5102 I2S DAC Decoder PAM8403 Amplifier 8Ω 3W Speaker XL4016 Buck Converter 2S Li-Ion BMS 2S Lithium Charging Circuit USB-C Breakout Board
Role
Embedded Systems Design Edge Image Processing Voice Assistant Architecture Power Distribution

Architecture

Power Distribution Network

Primary Supply: Dual 3.7V 18650 Li-ion cells connected in series to establish a 7.4V nominal bus.

Battery Management & Charging: Integrated a 2S BMS to handle cell balancing, overcharge/over-discharge protection, and thermal safety, paired with a dedicated 2S charging module.

Step-Down Regulation: Integrated an XL4016 DC-DC buck converter to step down the 7.4V bus to a stable 5V output capable of powering the high-current demands of the Raspberry Pi 5.

Delivery Infrastructure: Routed power through heavy-gauge 18 AWG wiring to minimize voltage drop, terminating into a USB-C breakout board and cable for stable delivery directly to the Pi.

Audio Subsystem & Shared I2S Bus Routing

Microphone Input: Integrated an INMP441 MEMS omnidirectional microphone using digital I2S audio protocol for direct, noise-resilient voice capture.

Digital-to-Analog Conversion: Since the Raspberry Pi 5 lacks an onboard analog DAC, integrated a PCM5102 I2S DAC module to decode high-fidelity audio streams for the voice assistant.

Shared I2S Bus Architecture: Mapped both input (INMP441) and output (PCM5102) onto shared clock lines from the Raspberry Pi 5 header.

Amplification & Output: Routed the analog output from the PCM5102 into a PAM8403 mini-amplifier board to drive an onboard 8Ω 3W speaker for real-time voice feedback.

Computer Vision & iGlasses Assistive Navigation

Stereoscopic Hardware Interface: Integrated an IMX219-83 Stereo Binocular Camera module via dual MIPI-CSI connectors, utilizing Picamera2 to feed low-latency visual data directly to the Raspberry Pi 5.

OpenCV Depth Estimation: Engineered a custom Python-based processing pipeline that captures and converts dual visual arrays for OpenCV, enabling real-time stereoscopic depth estimation and spatial mapping.

Edge-Optimized YOLO Tracking: Deployed a YOLOv8 object detection model, heavily optimized through a compressed NCNN framework, to execute localized, high-speed object tracking entirely on the edge without cloud latency.

iGlasses Navigation Subsystem: Fused the OpenCV depth maps with YOLO bounding boxes to power the iGlasses feature, successfully translating complex environmental and obstacle data into actionable spatial awareness for the user.

Hardware prototype of TOKI system components
Initial Subsystem Prototyping

Software Architecture & Code Implementation

Module 1: TOKI Voice Assistant

PYTHON
import os
import re
import json
import subprocess
import time
import math
import requests
import speech_recognition as sr
from gtts import gTTS
from groq import Groq
from urllib.parse import quote
from ctypes import *
import threading
import pyaudio
from vosk import Model, KaldiRecognizer

# --- 1. HARDWARE LEVEL STDERR SUPPRESSION ---
class suppress_stderr:
    def __enter__(self):
        self.null_fd = os.open(os.devnull, os.O_RDWR)
        self.save_fd = os.dup(2)
        os.dup2(self.null_fd, 2)
    def __exit__(self, *_):
        os.dup2(self.save_fd, 2)
        os.close(self.null_fd)
        os.close(self.save_fd)

try:
    ERROR_HANDLER_FUNC = CFUNCTYPE(None, c_char_p, c_int, c_char_p, c_int, c_char_p)
    def py_error_handler(filename, line, function, err, fmt):
        pass
    c_error_handler = ERROR_HANDLER_FUNC(py_error_handler)
    asound = cdll.LoadLibrary('libasound.so')
    asound.snd_lib_error_set_handler(c_error_handler)
except Exception:
    pass

Purpose

Explains the architectural resolution of low-level runtime stream pollution in embedded Linux audio pipelines.

Problem

Linux sound architectures (ALSA) and underlying hardware abstractions on embedded platforms like the Raspberry Pi persistently output unhandled, non-critical C-level driver warnings directly to standard error (stderr). During real-time voice assistant stream initializations, this extraneous text corrupts standard output logs and degrades clean process monitoring.

Solution

suppress_stderr Context Manager: I implemented a deterministic, resource-safe Python context manager utilizing low-level OS file descriptors.

Redirecting stderr to /dev/null: I duplicated file descriptor 2 (stderr) using os.dup() and re-routed output streams via os.dup2() into a null device stream during stream handshakes.

Restoring stderr Safely: I designed the system to automatically restore original file descriptor bindings upon exit execution to prevent global stream degradation.

ALSA Custom Error Handler using ctypes: I bypassed OS stream boundaries by directly interfacing with compiled shared object symbols in user space.

Loading libasound.so: I dynamically bound the native ALSA client library at runtime via CDLL.

snd_lib_error_set_handler(): I overwrote the default error callback hook with a silent Python no-op trampoline function, completely muting native hardware spam.

Why this is important

Suppressing raw driver-level C diagnostics isolates high-level Python application telemetry from hardware noise. This guarantees production-grade console hygiene for headless embedded edge devices without compromising exception visibility or masking genuine runtime faults.

Key Takeaways

  • Low-level systems programming in Python requires manipulating OS file descriptors beyond standard runtime abstractions.
  • Native C libraries (libasound.so) can be safely intercepted using ctypes function signatures and custom error trampolines.
  • Context managers (__enter__ and __exit__) ensure deterministic resource acquisition and teardown for stream-level redirection.
  • Console noise suppression prevents log file bloat in long-running autonomous edge infrastructure nodes.
  • Bridging high-level speech recognition logic with bare-metal audio configurations is essential for robust embedded AI deployments.
PYTHON
# --- 2. CONFIGURATION & API KEYS ---
# SECURITY: Removed compromised hardcoded keys. Use environment variables.
GROQ_API_KEY = os.environ.get("GROQ_API_KEY")
GEOAPIFY_API_KEY = os.environ.get("GEOAPIFY_API_KEY")
GOOGLE_GEOLOCATION_API_KEY = os.environ.get("GOOGLE_GEOLOCATION_API_KEY")
ELEVENLABS_API_KEY = os.environ.get("ELEVENLABS_API_KEY")

# Telegram Bot Token
TELEGRAM_BOT_TOKEN = os.environ.get("TELEGRAM_BOT_TOKEN")

# --- 2.5 MEMORY CACHE ---
# --- 2.5 MEMORY CACHE & CIRCUIT BREAKERS ---
CACHED_LOCATION = None
LAST_LOCATION_TIME = 0
ELEVENLABS_DISABLED = False  # The new circuit breaker

# Map spoken names to their specific Telegram Chat IDs 
TELEGRAM_CONTACTS = {
    "omar": os.environ.get("TELEGRAM_CHAT_ID_OMAR"),
    "asadullah": os.environ.get("TELEGRAM_CHAT_ID_ASADULLAH"),
 
}

if not GROQ_API_KEY:
    raise ValueError("GROQ_API_KEY environment variable is missing.")

llm_client = Groq(api_key=GROQ_API_KEY)
http_session = requests.Session() # Optimization: Connection pooling
r = sr.Recognizer()

print("Loading local STT model...")
with suppress_stderr():
    vosk_model = Model("model")

category_map = {
    "supermarket": "commercial.supermarket", "süpermarket": "commercial.supermarket", "market": "commercial.supermarket", "markete": "commercial.supermarket",
    "cafe": "catering.cafe", "kafe": "catering.cafe", "kafeye": "catering.cafe",
    "restaurant": "catering.restaurant", "restoran": "catering.restaurant", "restorana": "catering.restaurant",
    "pharmacy": "healthcare.pharmacy", "eczane": "healthcare.pharmacy", "eczaneye": "healthcare.pharmacy",
    "hospital": "healthcare.hospital", "hastane": "healthcare.hospital", "hastaneye": "healthcare.hospital",
    "school": "education.school", "okul": "education.school", "okula": "education.school",
    "university": "education.university", "üniversite": "education.university", "üniversiteye": "education.university",
    "bank": "service.financial.bank", "banka": "service.financial.bank", "bankaya": "service.financial.bank",
    "atm": "service.financial.atm", "atmye": "service.financial.atm",
    "gas station": "service.vehicle.fuel", "benzinlik": "service.vehicle.fuel", "benzin istasyonu": "service.vehicle.fuel",
    "hotel": "accommodation.hotel", "otel": "accommodation.hotel", "otele": "accommodation.hotel",
    "bus stop": "public_transport.bus", "durak": "public_transport.bus", "durağa": "public_transport.bus",
    "train station": "public_transport.train", "istasyon": "public_transport.train", "istasyona": "public_transport.train",
    "park": "leisure.park", "parka": "leisure.park",
    "store": "commercial", "mağaza": "commercial", "dükkan": "commercial", "dükkana": "commercial"
}

Purpose

Establishes secure credential management, optimizes network and memory states, and defines the semantic mapping necessary for routing natural language to geographic APIs.

Problem

Hardcoding API keys exposes sensitive credentials in public repositories. Furthermore, voice assistants suffer from latency if they re-initialize connections or repeatedly query GPS hardware for every command. Additionally, relying solely on cloud Text-to-Speech (TTS) services introduces a single point of failure if quotas are exceeded, and unstructured natural language inputs for navigation must be standardized into strict API-friendly parameters.

Solution

Environment Variable Security (os.environ.get): I migrated all API tokens and Telegram Chat IDs to environment variables, preventing credential leaks and decoupling configuration from the codebase for safe version control.

Circuit Breaker Pattern (ELEVENLABS_DISABLED): I implemented a state flag to instantly bypass the premium TTS API if a quota error or network timeout occurs, preventing cascading system lockups and ensuring rapid fallback to local TTS.

Memory Caching (CACHED_LOCATION): I introduced temporal caching for GPS coordinates to minimize redundant hardware polling, significantly reducing system latency during continuous conversation loops.

Connection Pooling (requests.Session()): I instantiated a persistent HTTP session to reuse TCP connections across the application, which eliminates the overhead of repeated TLS handshakes for high-frequency cloud API calls.

Semantic Category Mapping (category_map): I designed a bilingual (English/Turkish) dictionary structure to deterministically translate fuzzy spoken intents (e.g., "markete") into strict hierarchical tags (e.g., "commercial.supermarket") required by the Geoapify routing engine.

Why this is important

Designing robust configuration and state management is what transforms a prototype into a production-ready application. By implementing circuit breakers, connection pooling, and credential security, the system guarantees high availability, safe open-source distribution, and low latency—critical metrics for real-time voice interfaces running on edge devices.

Key Takeaways

  • Environment variables are mandatory for securing secrets and adapting applications across different deployment environments.
  • Stateful circuit breakers are essential to prevent edge devices from hanging when third-party cloud services throttle or degrade.
  • Connection pooling (HTTP Keep-Alive) is a crucial micro-optimization for latency-sensitive REST API integrations.
  • Bilingual data mapping structures allow natural language models to interface predictably with strict, rigid geographic APIs.
PYTHON
def get_mic_index():
    with suppress_stderr():
        p = pyaudio.PyAudio()
        target_index = None
        for i in range(p.get_device_count()):
            dev = p.get_device_info_by_index(i)
            if "voicehat" in dev['name'].lower():
                target_index = i
                if "plug" in dev['name'].lower():
                    break
        p.terminate()
    return target_index

print("Probing I2S hardware...")
GLOBAL_MIC_IDX = get_mic_index()

# --- 3. CORE FUNCTIONS ---
def speak(text):
    global ELEVENLABS_DISABLED
    
    if not text:
        return
        
    print(f"Speaking: {text}")
    clean_text = text.replace('"', '').replace("'", "")
    
    # Only try ElevenLabs if we have a key AND the breaker hasn't tripped
    if ELEVENLABS_API_KEY and not ELEVENLABS_DISABLED:
        url = "https://api.elevenlabs.io/v1/text-to-speech/liF98vtVyeMPMLuaIUXC/stream"
        
        headers = {
            "Accept": "audio/mpeg",
            "Content-Type": "application/json",
            "xi-api-key": ELEVENLABS_API_KEY
        }
        
        payload = {
            "text": clean_text,
            "model_id": "eleven_turbo_v2_5", 
            "voice_settings": {
                "stability": 0.5,
                "similarity_boost": 0.75
            }
        }
        
        try:
            response = http_session.post(url, json=payload, headers=headers, stream=True, timeout=10)
            
            if response.status_code == 200:
                process = subprocess.Popen(
                    ['mpg123', '-q', '-f', '6552', '-o', 'alsa', '-a', 'plughw:2,0', '-'],
                    stdin=subprocess.PIPE
                )
                for chunk in response.iter_content(chunk_size=4096):
                    if chunk:
                        process.stdin.write(chunk)
                process.stdin.close()
                process.wait(timeout=15) 
                return
            elif response.status_code == 401 and "quota_exceeded" in response.text:
                print("\n[SYSTEM] ElevenLabs Quota Exceeded. Tripping circuit breaker to ensure fast responses.")
                ELEVENLABS_DISABLED = True # Breaker trips! Future calls skip ElevenLabs instantly.
            else:
                print(f"[WARNING] ElevenLabs API Error: {response.status_code}")
                print("Switching to backup Google voice...")
                
        except subprocess.TimeoutExpired:
            print("[WARNING] Hardware audio playback timed out (ALSA lockup). Resetting...")
            subprocess.run('killall -9 mpg123', shell=True)
            return
        except Exception as e:
            print(f"Network streaming error: {e}")
            print("Switching to backup Google voice...")
            
    # --- FALLBACK TO FREE GOOGLE TTS ---
    try:
        tts_lang = 'tr' if any(c in clean_text for c in ['İ', 'ü', 'ş', 'ğ', 'ç', 'ö']) else 'en'
        tts = gTTS(text=clean_text, lang=tts_lang, tld='us')
        tts.save("response.mp3")
        subprocess.run('mpg123 -q -f 6552 -o alsa -a plughw:2,0 response.mp3 >/dev/null 2>&1', shell=True, timeout=15)
    except subprocess.TimeoutExpired:
        subprocess.run('killall -9 mpg123', shell=True)
    except Exception as e:
        print(f"gTTS error: {e}")

Purpose

Handles dynamic hardware probing for the I2S microphone and orchestrates a highly resilient, streaming Text-to-Speech (TTS) pipeline to minimize latency.

Problem

In embedded Linux environments, ALSA audio device indices are unstable and can change randomly upon reboot, causing hardcoded applications to fail silently. Furthermore, traditional cloud TTS implementations download an entire audio file before playing it, creating unnatural delays in conversation. Finally, if the hardware audio decoder freezes or a premium API quota is reached, the entire application thread can lock up, requiring a hard manual reset of the edge device.

Solution

Dynamic Device Probing (get_mic_index): I utilized PyAudio to iterate through all system audio interfaces at runtime, searching specifically for the physical "voicehat" string. This guarantees the correct hardware is bound regardless of the OS boot order.

Audio Streaming via Popen: I bypassed local file saving by piping ElevenLabs API response chunks directly into the standard input of the mpg123 decoder via subprocess.Popen, significantly reducing the time-to-first-audio.

Hardware Interface Targeting: I explicitly routed playback to plughw:2,0 using ALSA command-line arguments to bypass default OS mixing layers and send the stream directly to the correct DAC.

Process Timeout & Cleanup: I engineered strict timeout parameters (process.wait(timeout=15)) and automated recovery routines (killall -9 mpg123) to aggressively terminate zombie audio threads and free locked hardware.

Automated Language Detection & Fallback: I built a character filter to dynamically detect Turkish-specific vowels (e.g., 'ş', 'ğ'), automatically switching the backup Google TTS engine (gTTS) to the correct local language profile when the primary stream fails.

Why this is important

In autonomous edge computing, hardware systems must be entirely self-healing. By implementing dynamic I2S probing and aggressive process termination, the application survives reboots and driver lockups without user intervention. Additionally, transitioning from file-based I/O to memory-piped streaming ensures the voice assistant achieves the near-instantaneous response times expected in modern AI architectures.

Key Takeaways

  • Never hardcode Linux audio device indices; dynamic hardware probing is mandatory for resilient edge deployments.
  • Streaming chunked data directly to a subprocess standard input drastically cuts memory bloat and I/O latency.
  • Any hardware-interfacing process must have a definitive timeout and a kill signal to prevent permanent systemic lockups.
  • Multi-tier fallbacks (Premium Streaming → Free Cloud API) combined with circuit breakers ensure 100% uptime for voice interfaces.
PYTHON
def _transcribe_groq(audio):
    wav_bytes = audio.get_wav_data(convert_rate=16000, convert_width=2)
    transcription = llm_client.audio.transcriptions.create(
        file=("speech.wav", wav_bytes),
        model="whisper-large-v3",
        language="en",
        response_format="text",
        temperature=0.0,
        prompt="Voice assistant commands: directions, navigate, hospital, restaurant, pharmacy, emergency, TOKI.",
    )
    return transcription.text.strip() if hasattr(transcription, "text") else transcription.strip()

def _transcribe_vosk(audio):
    raw_data = audio.get_raw_data(convert_rate=16000, convert_width=2)
    rec = KaldiRecognizer(vosk_model, 16000)
    rec.AcceptWaveform(raw_data)
    return json.loads(rec.FinalResult()).get("text", "")

def listen(prompt=None):
    if prompt:
        speak(prompt)
        
    with suppress_stderr():
        # Dynamically probe the hardware every time we listen 
        # so we always have the correct index, even if it reconnected!
        current_mic_idx = get_mic_index() 
        source = sr.Microphone(device_index=current_mic_idx)
        
        with source:
            print("\nListening...")
            
            # FIX 1: Raised from 300 to 500 to ignore background static.
            # (If it stops hearing you when you actually speak, lower this to 400).
            r.energy_threshold = 800 
            
            r.dynamic_energy_threshold = True
            r.pause_threshold = 1.2
            r.non_speaking_duration = 0.8

            try:
                audio = r.listen(source, timeout=5, phrase_time_limit=12)
            except sr.WaitTimeoutError:
                return ""

    text = ""
    try:
        text = _transcribe_groq(audio)
    except Exception as e:
        print(f"Groq STT failed ({e}), falling back to local Vosk...")

    if not text:
        try:
            text = _transcribe_vosk(audio)
        except Exception as e:
            print(f"STT Error: {e}")
            speak("Sorry, I did not understand. Please say it again.")
            return ""

    
    # Whisper often generates these exact phrases when fed pure silence.
    hallucinations = [
        "thank you for watching",
        "thanks for watching",
        "thank you.",
        "amara.org",
        "subtitles by",
        "the first step is to get the patient to the hospital"
    ]
    
    clean_text = text.lower().strip()
    
    # If the text perfectly matches a known hallucination, ignore it.
    if any(h in clean_text for h in hallucinations) or len(clean_text) < 2:
        return ""

    print("User said:", text)
    return clean_text

Purpose

Orchestrates a highly resilient, hybrid Speech-to-Text (STT) pipeline with integrated hardware noise rejection and AI hallucination filtering.

Problem

Physical microphones on edge devices capture persistent background static that can trigger false recording cycles. Furthermore, generative cloud STT models (like Whisper) are known to hallucinate phantom phrases (e.g., "thanks for watching") when fed pure silence. Finally, relying solely on cloud transcription leaves the voice assistant completely non-functional during local network outages.

Solution

Hybrid STT Architecture: I engineered a two-tier transcription pipeline. The system attempts high-speed cloud transcription via Groq's Whisper API first, seamlessly catching network exceptions and routing the audio payload to a local, offline Vosk model (KaldiRecognizer) if the cloud is unreachable.

Audio Pre-Processing: I explicitly downsampled microphone byte streams to a 16kHz, 16-bit PCM format (convert_rate=16000, convert_width=2), strictly adhering to the input tensor requirements of both the Whisper and Kaldi neural networks.

Prompt Biasing & Zero Temperature: During the Groq API call, I utilized a context prompt containing expected vocabulary (navigate, hospital, etc.) and set temperature=0.0 to force deterministic, highly accurate command recognition over creative variance.

Hardware Noise Rejection: I manually tuned the hardware input layer by raising the energy_threshold to 800 and configuring dynamic thresholding. This aggressively ignores ambient room static, preventing the microphone array from locking into a perpetual recording state.

The Hallucination Filter: I implemented a programmatic string-matching filter to intercept and discard known Whisper AI anomalies (e.g., "subtitles by", "amara.org") generated by silent audio frames, guaranteeing zero false-positive triggers in the main command loop.

Why this is important

For a voice interface to feel natural and robust in real-world, unpredictable environments, it must flawlessly distinguish between human intent and hardware noise. By combining strict AI prompting, acoustic threshold tuning, and multi-layered offline fallbacks, the system guarantees consistent uptime and prevents ghost commands, which is absolute critical for autonomous edge computing.

Key Takeaways

  • Cloud AI architectures require local safety nets; trapping network exceptions and routing to offline models prevents fatal application crashes.
  • Audio waveforms must be strictly formatted (16kHz, 16-bit PCM) before passing to STT neural networks to prevent silent transcription failures.
  • Generative AI models will predictably hallucinate on silent or noisy inputs; strict programmatic filters are mandatory to maintain system integrity.
  • Tuning physical hardware energy thresholds is just as critical as the software logic when processing real-time microphone data.
PYTHON
def turkish_mode():
    speak("Turkish mode activated. I will understand your English, but I will answer only in Turkish. Say 'exit Turkish' to stop.")
    
    while True:
        # We reuse your highly stable, default listen() function! 
        # This automatically uses the hallucination filters you already built.
        user_input = listen()
        
        if not user_input:
            continue
            
        if any(word in user_input for word in ["exit turkish", "stop turkish", "çıkış"]):
            speak("Exiting Turkish mode. Returning to English.")
            break
            
        # The prompt forces Llama to act normally, but output strictly Turkish
        prompt = f"""You are TOKI, a helpful voice assistant. 
        The user will speak to you. You must answer their query naturally and concisely, but you must ONLY speak in Turkish. 
        Do not provide translations of what they said, just answer them directly in Turkish.
        
        User: {user_input}
        Assistant:"""
        
        turkish_response = query_llm(prompt)
        
        # We use your standard speak() function since ElevenLabs 
        # turbo_v2_5 auto-detects and speaks Turkish perfectly.
        speak(turkish_response)

def stream_elevenlabs(text):
    if not ELEVENLABS_API_KEY:
        print("ElevenLabs API key missing.")
        return

    # Using the /stream endpoint instead of the standard TTS endpoint
    url = "https://api.elevenlabs.io/v1/text-to-speech/liF98vtVyeMPMLuaIUXC/stream"
    
    headers = {
        "Accept": "audio/mpeg",
        "Content-Type": "application/json",
        "xi-api-key": ELEVENLABS_API_KEY
    }
    
    payload = {
        "text": text,
        # eleven_turbo_v2_5 natively supports both Spanish and Turkish
        "model_id": "eleven_turbo_v2_5", 
        "voice_settings": {
            "stability": 0.5,
            "similarity_boost": 0.75
        }
    }
    
    try:
        response = http_session.post(url, json=payload, headers=headers, stream=True, timeout=10)
        
        if response.status_code == 200:
            process = subprocess.Popen(
                ['mpg123', '-q', '-f', '16384', '-o', 'alsa', '-a', 'plughw:2,0', '-'],
                stdin=subprocess.PIPE
            )
            for chunk in response.iter_content(chunk_size=4096):
                if chunk:
                    process.stdin.write(chunk)
            process.stdin.close()
            process.wait()
        else:
            print(f"ElevenLabs Stream Error: {response.status_code}")
            
    except Exception as e:
        print(f"Streaming hardware error: {e}")

Purpose

Implements a dedicated, state-locked bilingual conversational loop and a low-latency HTTP streaming pipeline for real-time Text-to-Speech playback.

Problem

Traditional voice assistants struggle with seamless bilingual context switching, often requiring complex routing models to understand English but reply in another language. Furthermore, waiting for standard cloud TTS endpoints to synthesize and download a complete audio file introduces unacceptable latency, breaking the illusion of a real-time, natural conversation.

Solution

Context-Locked LLM Prompting (turkish_mode): I engineered a dedicated conversational loop that wraps the user's input in a strict system prompt. This forces the underlying Llama model to process the English input contextually but strictly synthesize its output in Turkish without breaking character or providing literal translations.

Modular Architecture Reuse: I integrated my existing listen() and speak() functions directly into the loop, allowing the new language mode to automatically inherit all the hardware-level stability and hallucination filtering I had already built, without writing redundant code.

Stream API Endpoint (/stream): I targeted ElevenLabs' streaming endpoint rather than the standard generation endpoint, allowing the application to receive audio packets sequentially before the full sentence finishes generating in the cloud.

Subprocess Piping (subprocess.Popen): I established a direct memory pipe into the mpg123 hardware decoder's standard input. By iterating over the network response chunks (response.iter_content(chunk_size=4096)), I feed audio to the ALSA driver the millisecond the bytes arrive over the network.

Why this is important

True conversational AI requires both natural language flexibility and minimal latency. By forcing the LLM to handle the translation logic contextually and piping the resulting audio chunks directly into hardware memory rather than saving them to a disk, the system achieves near-human response times and seamless bilingual support without overloading the edge device's CPU.

Key Takeaways

  • Strict prompt engineering can effectively replace complex NLP routing models for bilingual interactions.
  • Streaming TTS via HTTP chunks and subprocess pipes is critical for reducing "Time to First Byte" (TTFB) latency in voice assistants.
  • Building modular core functions (like a robust listen pipeline) allows new features to be added quickly while maintaining production-level hardware stability.
PYTHON
def get_location():
    global CACHED_LOCATION, LAST_LOCATION_TIME
    
    # SPEED OPTIMIZATION: If we scanned your location in the last 5 minutes (300 seconds), 
    # reuse it instantly instead of freezing the Raspberry Pi with another network scan!
    if CACHED_LOCATION and (time.time() - LAST_LOCATION_TIME) < 300:
        print("[Speed Boost] Using cached GPS coordinates.")
        return CACHED_LOCATION

    default_coords = (40.1828, 29.0665)
    
    if not GOOGLE_GEOLOCATION_API_KEY:
        print("\n[WARNING] Google API Key is missing! Falling back to default location.")
        return default_coords

    print("Scanning nearby Wi-Fi networks...")
    mac_addresses = []

    try:
        # Reduced timeout to prevent long hangs if the Pi's Wi-Fi adapter is busy
        scan_result = subprocess.run(
            ['nmcli', '-t', '-f', 'BSSID,SIGNAL', 'dev', 'wifi', 'list'],
            capture_output=True, text=True, timeout=5 
        )
        for line in scan_result.stdout.strip().split('\n'):
            if not line:
                continue
            parts = line.rsplit(':', 1)
            if len(parts) == 2:
                clean_bssid = parts[0].strip().replace('\\', '').replace('-', ':').replace(' ', ':').upper()
                if re.match(r'^([0-9A-F]{2}[:-]){5}([0-9A-F]{2})$', clean_bssid):
                    mac_addresses.append({
                        "macAddress": clean_bssid,
                        "signalStrength": int(int(parts[1].strip()) / 2 - 100)
                    })
    except Exception as e:
        print(f"Local Wi-Fi Hardware Scan Failure: {e}")

    if len(mac_addresses) < 2:
        print("Insufficient Wi-Fi beacons. Bypassing to IP Geolocation...")
        try:
            response = http_session.get("http://ip-api.com/json/", timeout=5)
            data = response.json()
            if data.get('status') == 'success':
                CACHED_LOCATION = (data['lat'], data['lon'])
                LAST_LOCATION_TIME = time.time()
                return CACHED_LOCATION
            return default_coords
        except Exception:
            return default_coords

    try:
        url = f"https://www.googleapis.com/geolocation/v1/geolocate?key={GOOGLE_GEOLOCATION_API_KEY}"
        response = http_session.post(url, json={"considerIp": "true", "wifiAccessPoints": mac_addresses}, timeout=5)
        if response.status_code == 200:
            loc = response.json()['location']
            CACHED_LOCATION = (loc['lat'], loc['lng'])
            LAST_LOCATION_TIME = time.time()
            return CACHED_LOCATION
        return default_coords
    except Exception as e:
        print(f"Geolocation API failed: {e}")
        return default_coords

def get_route_directions(start_lat, start_lon, end_lat, end_lon, is_turkish=False):
    # This automatically swaps the API language based on your mode!
    lang_param = "tr" if is_turkish else "en"
    url = f"https://api.geoapify.com/v1/routing?waypoints={start_lat},{start_lon}|{end_lat},{end_lon}&mode=drive&details=instruction_details&lang={lang_param}&apiKey={GEOAPIFY_API_KEY}"
    
    response = http_session.get(url)
    if response.status_code != 200:
        return None
    try:
        return [step['instruction']['text'] for leg in response.json()['features'][0]['properties']['legs'] for step in leg['steps']]
    except (KeyError, IndexError) as e:
        print("Error parsing routing steps:", e)
        return None

def _haversine_meters(lat1, lon1, lat2, lon2):
    R = 6371000
    phi1, phi2 = math.radians(lat1), math.radians(lat2)
    dphi, dlambda = math.radians(lat2 - lat1), math.radians(lon2 - lon1)
    a = math.sin(dphi / 2) ** 2 + math.cos(phi1) * math.cos(phi2) * math.sin(dlambda / 2) ** 2
    return 2 * R * math.asin(math.sqrt(a))

Purpose

Calculates high-precision geographic coordinates via hardware MAC address scanning, fetches dynamic turn-by-turn routing, and calculates spherical planetary distance using native mathematical models.

Problem

Standard IP-based geolocation is highly inaccurate, often placing edge devices miles away from their actual physical location, which renders local navigation commands useless. Furthermore, querying external APIs for every single voice interaction creates severe network bottlenecks. Finally, simple flat-plane distance calculations fail to account for the Earth's curvature when determining the true physical proximity between the user and a destination.

Solution

Hardware Wi-Fi Scraping (nmcli): Instead of relying on crude IP estimation, I executed native Linux subprocesses to scrape the BSSID (MAC addresses) and signal strengths of local Wi-Fi routers. I pass this telemetry to Google's Geolocation API to achieve near-GPS accuracy indoors.

Temporal Caching (CACHED_LOCATION): I engineered a 300-second (5-minute) memory cache to store recent coordinate data. This prevents the system from triggering redundant, thread-blocking hardware scans when the assistant is queried multiple times in rapid succession.

Multi-Tier Fallbacks: I designed a robust failover cascade. If the Wi-Fi hardware scan fails (finding fewer than 2 beacons), the system gracefully degrades to IP-based location (ip-api.com), and if the network drops completely, it defaults to a hardcoded baseline coordinate.

Bilingual Routing (lang_param): I dynamically toggle the language parameter of the Geoapify routing API based on the active conversation mode, seamlessly pulling localized driving instructions without needing to run secondary translation models.

Spherical Trigonometry (_haversine_meters): I implemented the Haversine formula using Python's native math library to accurately compute the great-circle distance between two GPS coordinates entirely offline, ensuring precise proximity measurements.

Why this is important

For autonomous hardware operating in the real world, spatial context is everything. Standard IP location is simply too imprecise for micro-navigation. By aggressively scraping hardware-level Wi-Fi data on the device and calculating true spherical distances locally, the system achieves the critical location awareness required for a production-level voice assistant to provide accurate geographic routing.

Key Takeaways

  • Hardware-level network scraping (BSSID/MAC scanning) is vastly superior to IP-based routing for precise edge-device geolocation.
  • Temporal caching mechanisms are mandatory to prevent hardware throttling and UI freezes in continuous-listening voice applications.
  • Multi-layered fallbacks (Wi-Fi Scanning → IP Geolocation → Hardcoded Baseline) ensure the application never crashes due to a sensor or network failure.
  • Calculating physical distance natively using the Haversine formula reduces reliance on continuous external API calls and speeds up data processing.
PYTHON
def get_places(category, latitude, longitude, radius=5000, limit=10):
    url = f"https://api.geoapify.com/v2/places?categories={quote(category)}&filter=circle:{longitude},{latitude},{radius}&bias=proximity:{longitude},{latitude}&limit={limit}&apiKey={GEOAPIFY_API_KEY}"
    response = http_session.get(url)

    if response.status_code != 200:
        return None

    features = response.json().get('features', [])
    def dist_key(feature):
        try:
            plon, plat = feature['geometry']['coordinates']
            return _haversine_meters(latitude, longitude, plat, plon)
        except Exception:
            return float('inf')

    features.sort(key=dist_key)
    return features

def get_place_by_name(name, latitude, longitude, radius=25000, limit=5):
    url = f"https://api.geoapify.com/v1/geocode/autocomplete?text={quote(name)}&filter=circle:{longitude},{latitude},{radius}&bias=proximity:{longitude},{latitude}&limit={limit}&apiKey={GEOAPIFY_API_KEY}"
    
    response = http_session.get(url)

    if response.status_code != 200:
        return None

    features = response.json().get('features', [])
    
    # Sort the results by distance so you always get the closest one
    def dist_key(feature):
        try:
            plon, plat = feature['geometry']['coordinates']
            return _haversine_meters(latitude, longitude, plat, plon)
        except Exception:
            return float('inf')

    features.sort(key=dist_key)
    return features

def send_emergency_alert(location_link):
    print(f"\n[ACTION] Transmitting emergency coordinates: {location_link}")
    # Example using a generic webhook:
    # try:
    #     webhook_url = "https://your-webhook.url/endpoint"
    #     requests.post(webhook_url, json={"alert": "Emergency", "location": location_link})
    # except Exception as e:
    #     print(f"Failed to send alert: {e}")

def clean_location_name(destination):
    prompt = f"""You are a geocoding data processor mapping English speech to Turkish maps. 
    The user wants to navigate to: '{destination}'.
    If the input is in English, translate it to the precise Turkish location name.
    CRITICAL RULE: For Turkish government ministries or state institutions located in local provinces, use the Provincial Directorate title (e.g., output 'Sanayi ve Teknoloji İl Müdürlüğü' instead of 'Bakanlığı').
    If it is a brand or already Turkish, keep it as is.
    You must reply with ONLY valid JSON containing a single key "location". Do not add any conversational text."""
    
    try:
        response = llm_client.chat.completions.create(
            messages=[{"role": "user", "content": prompt}],
            model="llama-3.1-8b-instant",
            response_format={"type": "json_object"}, 
            temperature=0.0 
        )
        data = json.loads(response.choices[0].message.content)
        return data.get("location", destination).strip()
    except Exception as e:
        print(f"[JSON Extraction Error]: {e}")
        return destination

Purpose

Handles proximity-based geographic searches, typo-tolerant location lookups, and utilizes a large language model (LLM) as a strict data-sanitization pipeline to bridge the gap between messy human speech and rigid map APIs.

Problem

Map APIs are highly sensitive to formatting. If a user asks a voice assistant for the "Ministry of Industry" in English, a Turkish map API will fail to find it because it is cataloged as "Sanayi ve Teknoloji İl Müdürlüğü". Furthermore, geographic APIs often return search results in arbitrary orders, meaning the first result in the array might not actually be the closest physical location to the user. Finally, spoken names often contain minor mispronunciations or typos that break strict search endpoints.

Solution

Proximity-Forced Sorting (dist_key): I intercepted the raw JSON array returned by the Geoapify API and injected a custom Python sorting lambda. By running each result's coordinates through my native _haversine_meters calculation, I mathematically guarantee that the features array is strictly sorted by real-world physical proximity before the assistant reads it.

Fuzzy Searching (/autocomplete): Instead of using a rigid /search endpoint, I targeted the /autocomplete endpoint for named locations. This acts as a fuzzy-matching layer, making the system highly tolerant to minor typos generated by the Speech-to-Text engine.

LLM Data Normalization (clean_location_name): I engineered a specialized prompt for the Llama 3.1 model, forcing it to act purely as a data transformation pipeline rather than a conversational bot. It intercepts English geographic intents and translates them into precise, localized bureaucratic entities.

Strict JSON Enforcement (response_format={"type": "json_object"}): To prevent the LLM from appending useless conversational text (e.g., "Sure, here is the translation..."), I constrained the API payload to only accept valid JSON, ensuring programmatic extraction (json.loads()) never crashes.

Why this is important

In autonomous navigation systems, data integrity is everything. Real human speech is chaotic, while geographic databases require absolute precision. By using an LLM to aggressively format and translate intent before querying the map, and mathematically sorting the results after querying the map, the application achieves enterprise-grade routing accuracy and reliability.

Key Takeaways

  • Client-side Haversine sorting is mandatory; never trust an external map API to optimize results perfectly for physical proximity by default.
  • LLMs are incredibly powerful when constrained; using temperature=0.0 and JSON-object enforcement turns a generative AI into a highly reliable data-sanitization tool.
  • Fuzzy matching endpoints (/autocomplete) provide significantly better user experiences for voice assistants than strict lookup endpoints, as they naturally absorb STT inaccuracies.