Theses and Dissertations (Electrical, Electronic and Computer Engineering)
Permanent URI for this collectionhttp://hdl.handle.net/2263/31930
Browse
Recent Submissions
Now showing 1 - 20 of 492
Item A spin on compressive sensing imaging : reticle-based single-pixel imaging systemVan der Merwe, Marco Marselle (University of Pretoria, 2024-09)Various aspects of a single-pixel imaging (SPI) system are analyzed using a spinning reticle often referred to as a spinning coded aperture. The cost of imaging outside of the visual spectrum using FPA detectors becomes very expensive and an affordable alternative can have major implications in the commercial and military markets leading to more affordable imaging systems. The first documented use of a spinning disk imaging system came from the development of infrared (IR) missile seeker heads during the Second World War. These systems modulated the incoming light in amplitude of frequency to track a target. The use of spinning frequency modulation (FM) reticles as imaging devices was analyzed to determine the capability of such a system in general imaging scenarios. It was found that the frequency versus angle and radius reticle, termed FM-AR reticle, could only image simple scenes consisting of only a few targets. Using the simpler frequency versus radius reticle and modulating only a single column of pixels for each reticle rotation proved to be a viable option for imaging a wider range of scenes. However, the major drawback was that the imaging time to produce a single image was a factor of the number of columns in the image and the use of compressive sensing techniques was investigated as imaging of the entire scene in one reticle rotation was desired. compressive sensing (CS) is a signal acquisition technique to recover a sparse vector from only a few linear measurements. CS assumes a sparse vector being sampled and the use of sparsifying dictionaries is required for sampling non-sparse vectors such as images. Optimizing the sensing matrix to improve the recovery quality of the sparse vector uses coherence or cumulative coherence to determine a theoretical bound on the signal sparsity to guarantee successful recovery. SPI imaging systems employing CS techniques typically use a digital-micromirror-device (DMD) to modulate the incoming light and random binary patterns are often used. Using a reticle instead of a DMD has the potential to reduce the cost of such SPI systems as DMDs operating outside of the visual spectrum are also very expensive. The two types of reticle patterns considered related to where the patterns are constructed on the reticle. The edge-coded reticle has the reticle pattern constructed radially onto the surface and shifted towards the edge of the disk. The full-coded reticle used the entire surface for the reticle pattern. The edge-coded reticle was found to be a better alternative to the full-coded reticle as the edge-coded reticle allows controlling the CS compression ratio independently of the image dimension. Furthermore, the edge-coded reticle produces an image with a more consistent resolution compared to the full-coded reticle which consists of very small inner pixel dimensions and larger outer pixel dimensions. The genetic algorithm (GA) algorithm is a population-based metaheuristic method based on the Darwinian theory of survival of the fittest. The GA also allows the optimization of binary patterns that correspond to the opaque and transparent pixels on the reticle. Recovering sparse vectors with sensing matrices optimized by the GA for coherence and cumulative coherence generally showed improved reconstruction quality using sparse recovery algorithms. Op- timized sensing matrices consisting of floating point values showed improved recovery quality using OMP but the binary sensing matrices showed higher recovery quality using basis pursuit recovery algorithms. It was observed that the restrictions posed on the binary sensing matrix using an edge-coded reticle did not limit the resulting sensing matrix recovery ability and even improved the capability compared to the non-restricted binary sensing matrices. Training a sensing matrix on a separate image set, referred to as GA-TRAIN, using a sparse recovery algorithm to determine an average fitness score for each candidate in the GA’s population proved to be the better sensing matrix optimization technique. Using GA-TRAIN, improvements of up to 53.1% were observed compared to the GA optimizing coherence or cumulative coherence. Recovering real images using an edge-coded reticle only showed improvements comparing the first to final sensing matrix using the GA with the dB1 Daubechies wavelet dictionary in its second decomposition level (2dB1). This was only observed on the ImageNet dataset for optimizing coherence and cumulative coherence. The other dictionaries considered in this work, namely 1dB4, 5dB1 wavelet dictionary and the discrete cosine transform (DCT) dictionary, showed no noteworthy improvements in reconstruction quality when optimizing coherence and cumulative coherence on the MNIST and ImageNet datasets. The GA-TRAIN method was again used to train sensing matrices on real images and improved recovery quality was observed for all dictionaries and recovery algorithms except for the 2dB1 dictionary using GA-TRAIN. The coherence and cumulative coherence optimization showed slightly improved recovery quality on the ImageNet dataset compared to GA-TRAIN optimization method for the 2dB1 dictionary. It was generally noted that wavelet dictionaries with decomposition levels of one and two proved superior in sparse recovery compared to higher-level decomposition wavelets and the DCT in certain scenarios. The weight of a binary sensing matrix directly influences the signal-to-noise ratio (SNR) of the sampled vector and introducing noise into the sampling vector saw the coherence and cumulative coherence optimized sensing matrices struggle more than the GA-TRAIN optimized sensing matrix to accurately reconstruct the signals.Item A queueing-theoretic and dynamic-optimisation framework for end-to-end TCP congestion controlNgwenya, Dumisa Wellington (University of Pretoria, 2026-05-15)This thesis investigates end-to-end TCP congestion control using queueing-theoretic and optimal-control principles. It addresses the challenge of achieving high throughput while limiting queueing delay, round-trip-time inflation, oscillatory sending behaviour, and fairness degradation. The work links TCP congestion control to Little’s Law, Kleinrock’s Power metric, Stidham’s reward–cost formulation, and sensitivity-based throughput–delay optimality conditions. Building on these foundations, the thesis develops sender-side congestion-control mechanisms that regulate in-flight data and round-trip time while maintaining consistency between congestion-window evolution, delivery rate, and network delay. Two approaches are proposed: a queueing-theoretic stationary-delay regulator based on a predictor–corrector congestion-window update, and a nonlinear optimal-control formulation solved using Runge–Kutta integration and gradient-based optimisation. The methods are evaluated through ns-3 simulations and compared with representative TCP algorithms, including CUBIC and BBR. The thesis contributes a principled modelling framework and practical control algorithms for efficient, delay-aware, and analytically grounded TCP congestion control.Item Evaluating a human-machine student intervention framework in higher education from legacy dataCombrink , Herkulaas Morkel van Eyssen (University of Pretoria, 2026)This study aims to answer the question: “Can education be abstracted in a framework and conceptually studied for a student intervention process?” To do so, it designs and tests a student intervention framework that uses a systems approach and complexity theory to learn from different student contexts and recommend and apply systemic interventions for students. The framework is based on the assumption that education is a complex system that involves multiple interacting elements, such as students, teachers, curriculum, policies, and resources. The specific interventions and their impact on student success are not assessed in this work, as they depend on the expertise in the education domain. Rather, this work focuses on the framework that can be used to apply and evaluate different contexts and interventions. The study applies complexity theory and systemic intervention theory to understand the lens and the methods for studying the system. It also explains the education system in South Africa as the context of the study. One of the main challenges of the study is data sharing and data handling. To address this challenge, the study generates synthetic data from openly available tabular data and evaluates its conditional interdependence using different machine learning classification tasks. Then, it applies the same methods to a real-world education dataset from the University of the Free State. The study contextualizes the student intervention framework as a multi-armed bandit (MAB) stateless reinforcement learning problem and tests its performance and viability using probabilistic models. The results show that the probabilistic models yield the best results with the minimum required fine-tuning, and that they scale well to the real-world dataset. The results also indicate that the framework is viable for student intervention recommendations within the context of contextual bandits. The study concludes with a discussion of the implications and limitations of the framework, and suggests areas for future research.Item Development of solutions to facial nerve stimulation in cochlear implant recipientsOlivier, Milene (University of Pretoria, 2025)This dissertation investigates a non-invasive intervention for facial nerve stimulation, a common complication that restricts the performance of cochlear implant recipients and often leads to discomfort, reduced stimulation levels and diminished speech perception. Building on person-specific model-based predictions published in earlier work, this study evaluates the use of apical reference stimulation as an alternative grounding configuration intended to raise facial nerve excitation thresholds, expand the dynamic range and preserve or improve hearing performance. A series of clinical case studies were conducted to assess whether apical reference stimulation can reduce the prevalence of facial nerve stimulation without compromising threshold levels, upper stimulation levels or speech perception outcomes. Because many cochlear implant recipients lived far from the research facility, the study incorporated a remote programming protocol supported by a mobile follow-up application that enabled remote speech perception testing. Experimental results demonstrated that apical reference stimulation reduced or eliminated facial nerve stimulation in challenging cases and produced favourable threshold levels and dynamic ranges compared to standard stimulation strategies. The results also showed that remote programming sessions and remote speech perception assessments were comparable to conventional in-person procedures. The findings support apical reference stimulation as a promising non-invasive intervention and highlight the value of remote measures in expanding participation in cochlear implant research.Item A flexible printed non-enzymatic glucose sensor based on CeO2-doped ZnO nanostructures for plant applicationsMpala, Kimberly (University of Pretoria, 2025)In this study, a non-enzymatic glucose sensor for plants was developed and tested. Firstly, a three electrode sensor design was chosen and comprised of silver reference electrode and contacts, and carbon auxiliary and working electrodes. Experiments with varying concentrations of glucose were done to determine which printing method between inkjet printing and screen printing produces the better sensor. This was achieved by modifying the working electrode with commercial zinc oxide nanoparticle ink and glucose oxidase in order to ensure as best as possible the only variable was the printing technique. The printing technique selected was screen printing, as it generated more current for the same concentration of glucose compared to the inkjet printed electrode. Afterwards, research was done as to an appropriate dopant to improve sensitivity. This was selected to be cerium(IV) oxide due to its reported enzymatic-like properties. Zinc oxide nanoparticles with 1.7% doping and 10% doping concentrations were then synthesised. These were then used to modify the working electrode and evaluate glucose sensitivity. The 1.7% doping generated the highest current, obtained a normalised sensitivity of 190.6 μA·mM−1 ·cm−2, and was used in a plant based experiment to evaluate if it could differentiate a difference in glucose production when a plant was subjected to salt stress. The sensor was inserted into the plant, and plant sap was mixed with the appropriate electrolyte and used with the standard addition test. The results currently indicate promise with the sensor recording different current readings for a stressed and unstressed plant. These results were validated qualitatively by a glucose assay test on the same plant sap.Item Computer vision-based automated cow lameness estimation using machine learning and synthetic dataMuller, Alwyn Daniël (University of Pretoria, 2026-02)Lameness in cattle is a commonly encountered condition stemming from pain in one or more hoofs or limbs, which not only affects the animal’s movement, but also their productivity as well. Lameness can cause discomfort for cattle, reducing their quality of life, and can also lead to reduced milk production or reproductive issues. Farms with cattle exhibiting lameness can experience increased veterinary costs and economic loss. Therefore, it is important to monitor the lameness levels of cattle so that they can be treated as soon as possible. This dissertation discusses lameness in dairy cows, including various scoring systems and methods used to estimate the lameness scores of cows. The dissertation then delves into new computer vision-based algorithms that were developed to extract features such as the back arch, head, and hoof movement from cows in video recordings. These algorithms make use of You Only Look Once object detectors to detect different parts of the cow’s body to track the relevant features over time, which were then used in machine learning algorithms including neural networks and support vector machines to estimate the lameness score of the cows. The dissertation demonstrates that the lack of training data featuring lame cows is a problem and proposes a lameness simulator created in Unreal Engine 5 to simulate cows with different lameness scores as they move in a straight line past a stationary camera. The simulations validate various machine learning algorithms explained in the dissertation, and the trained machine learning models achieved an acceptable balanced accuracy for the estimated lameness scores. For the two-class lameness estimation, the best balanced accuracy score achieved was 92.65%. For the three-class lameness estimation, a balanced accuracy score 71.9% was achieved when using features from the proposed cow lameness simulator. For the four-class lameness estimation, the best balanced accuracy achieved when using features from the proposed cow lameness simulator was 65.09%. The estimation accuracy significantly drops with an increase in the number of possible lameness levels, due to the highly overlapping nature of the features between the classes. In all cases, the results were on par or better than those in published papers, although direct comparison is difficult due to the different possible scoring systems.Item Optimal dispatch strategy to minimize the cost of unserved energy during load shedding for an academic institutionNei, Mohau (University of Pretoria, 2025-11-28)Load shedding poses a significant challenge to electricity consumers as it disrupts operations, leading to lost utility and inconvenience. In response, academic institutions often rely on backup power solutions such as diesel generators to maintain continuity of activities during such outages. The Cost of Unserved Energy (CoUE) is a key metric used to quantify the economic impact of these disruptions, including the direct and indirect costs associated with mitigation strategies such as backup power systems. This study proposes an energy management solution to minimize the cost of unserved energy during load shedding for an academic institution. The aim is to determine the optimal dispatch strategy for power generation sources under varying load shedding scenarios, subject to capacity constraints while achieving minimum backup generation costs and maximum productivity. The system consists of diesel generators as the backup power source and photovoltaic systems as the renewable energy source. A case study of an academic institution in South Africa is investigated under three scenarios: (1) load shedding with no backup generation; (2) load shedding with full backup generation; and (3) load shedding with partial backup generation. The problem is formulated as a nonlinear programming (NLP) problem and solved using the ‘fmincon’ optimization solver in MATLAB. Simulation results show that, under the full backup generation, the demand is fully met but at increased cost due to reliance on diesel; when there is no backup, the operational costs are minimized but with high productivity losses; the partial backup generation offers a balance between generation cost and productivity loss. The proposed model results in more CoUE reduction under the partial backup scenario but with more unserved energy. The analysis also revealed that when PV contribution is sufficiently high, the optimizer avoids selecting any unmet demand, which led to lower CoUE during multiple outage periods. Finally, the findings showed that buildings with larger roof areas are at an advantage of mitigating the impacts of load shedding because they offer the optimizer greater freedom to allocate more PV compared to the smaller buildings.Item Onshore wind farm battery energy storage systems optimisationGwabavu, Mandisi (University of Pretoria, 2025-01)This study enhances the integration of Battery Energy Storage Systems (BESS) into onshore wind farms by improving efficiency, reliability, stability, and sustainability, thereby addressing significant challenges in renewable energy systems. Utilising advanced methodologies such as intelligent optimisation, hybrid forecasting models, and complex algorithms, the research provides innovative solutions for enhancing grid integration and energy management. The study aligns with the United Nations Sustainable Development Goals (SDGs), specifically SDGs 7 (Affordable and Clean Energy), 13 (Climate Action), and 8 (Decent Work and Economic Growth), advocates for economic prosperity, environmental stewardship, and social equity through sustainable renewable energy technologies. This study investigates the application of intelligent optimisation for integrating BESS with onshore wind farms to augment energy storage capacity and ensure grid reliability. The proposed model incorporates a neural network (NN)-based inverse model, Bayesian optimisation, Gaussian Process Regression (GPR), and Reinforcement Particle Swarm Optimisation (RPSO) to enhance wind energy production while providing accurate estimates of system performance. The study further introduces hybrid forecasting models that combine Long Short-Term Memory (LSTM) networks, Complementary Ensemble Empirical Mode Decomposition (CEEMD), and hybrid optimisation (ACO-GA-PSO) to attain accurate long-term predictions of wind energy variability. Lastly, the study reviews and applies a comprehensive techno-economic framework spanning Levelized Cost of Energy (LCOE), Levelized Cost of Storage (LCOS), Levelized Avoided Cost of Energy (LACE), and analysis together with layout and storage optimisation models to determine economically efficient BESS configurations for onshore wind. These frameworks undergo thorough validation through simulations, empirical case studies, and data analysis, with case studies of the 138 MW Gouda wind farm and the 69MW Jeffreys Bay wind farm in South Africa. The research also includes identifying resilient techno-economic models that align technical efficiency, cost-effectiveness, and long-term sustainability for BESS integration. The research significantly advances global renewable energy development by tackling challenges such as wind energy variability, battery energy storage system stability, and grid reliability. It achieves a 15-20% increase in energy efficiency and a 10% reduction in energy management errors, demonstrating practical applicability. The research supports sustainability objectives by reducing carbon footprints and greenhouse gas emissions, promoting environmental preservation, and generating economic and social benefits, including job creation and sustainable development. These advancements enhance operational efficiency, decision-making, and energy storage management, establishing onshore wind farms as fundamental components of a robust energy future. The study's innovative frameworks, validated through peer-reviewed publications and real-world applications, provide a scalable, transformative approach to renewable energy management. This research enhances global initiatives for a sustainable, low-carbon, and equitable energy framework by incorporating intelligent systems into renewable energy infrastructure, facilitating the expedited adoption of renewable energy worldwide.Item Low-cost High-performance Impedance Spectroscopy via Hardware–Software Co-design towards a CMOS ArchitectureDe Beer, Dirk (University of Pretoria, 2025-11-01)Electrochemical impedance spectroscopy (EIS) is a powerful analytical technique with critical applications in medical diagnostics, environmental monitoring, and materials science. However, its widespread adoption as a point-of-need tool has been hindered by a fundamental trade-off: high-performance instruments are typically large and expensive, whereas low-cost, portable alternatives often compromise on measurement accuracy and bandwidth. This thesis addresses this challenge by presenting a holistic approach to the design, implementation, and validation of a low-cost, high-performance impedance spectroscopy system designed to make this powerful technique accessible for ubiquitous diagnostics. This study fills a critical gap between high-end academic proofs-of-concept and limited-performance commercial solutions by co-designing a high-fidelity analog front-end, an efficient digital backend, and robust signal processing algorithms, all of which are designed to bridge the cost-performance divide.Item Preservation of vowel cue information in cochlear implant electrical stimuliPretorius, Marco (University of Pretoria, 2025-11-22)Cochlear implants (CIs) are beneficial for many people, but a large percentage of CI users are categorised as poor performing. While this poor performance is often a result of several person-specific factors that are fixed, including the health of cochlear neurons, the health of the cochlea, the electrode array insertion depth, and the proximity of the electrodes to the modiolus, choosing an optimal set of CI speech processor (clinical) parameters may reduce the negative effects of such factors and thus improve the performance of these CI users. However, it is not well known how to customise these parameters for individual users because researchers do not yet have a comprehensive understanding of how the CI distorts information. One of the reasons for this uncertainty is that the large number of interrelating factors makes it difficult for researchers to generalise the effects of clinical parameters on specific groups of CI users. This is because they need to have a homogeneous group of CI users, which is often difficult to obtain, that can participate in their listening experiments. Researchers often resort to acoustic models to eliminate (or model) person-specific factors of CI users by using normal hearing (NH) listeners who listen to speech processed through a CI processor and a model of what the distorted speech, available to the CI user, would sound like. This is useful for researchers to infer how much information is preserved by the processor. However, the estimated information from studies that use speech recognition tests may be biased due to the ability of listeners, in general, to deduce which speech stimulus material may have been presented using top-down linguistic skills. It is therefore unknown what the unbiased information is in the electrical stimulus signal at the output of a CI processor. The present study uses a method known as Continuous Feature Information Transmission Analysis (C-FITA) to estimate information about acoustic cues preserved in the electrical stimulus signal. The present study investigated the preservation of F1 and F2, since these are well known to be distorted by the processor. C-FITA was used in the present study to investigate the effect of the number of channels (m) and processing strategy on F1 and F2 information preserved in the electrical stimulus signal. The observations in the present study may lead to better guidelines for audiologists to assign clinical parameters to CI users and may serve as a guideline for CI designers to optimise the amount of F1 and F2 information preserved in the electrical stimulus signal for a given set of clinical parameters.Item Discrete radiance field representations for 3D object detection via sparse-dense convolutional neural networksVan Eeden, Janco (University of Pretoria, 2025-12)Indoor 3D object detection is a crucial computer vision task that enhances scene understanding and interpretation. Improving the efficiency and reliability of perception systems is essential for applications in robot perception, augmented reality, and virtual reality. Although 2D object detection methods have made significant progress, 3D environments present unique challenges that require specialized approaches. Despite the advancements in the 3D-domain research, existing methods rely on accurate depth information, which is often obtained through supplementary depth sensors or depth-estimation algorithms. RGB-only 3D object detection is inherently ill-posed due to ambiguous depth cues, occlusion, illumination, and camera motion. Advancements in radiance field representations, such as neural radiance fields (NeRFs) and 3D Gaussian splatting (3DGS), have shown remarkable success in novel view synthesis (NVS); their potential for 3D object detection remains underexplored. Recent methods rely on custom radiance field optimization pipelines that require significant memory overhead and compromise the modularity and generalizability of the approach. In this research, a CNN-based 3D object detection pipeline is developed that operates directly on discrete radiance field representations and thereby aims to address the generalizability and efficiency limitations of existing methods. The proposed two-stage detection system first reconstructs scene geometry using implicit or explicit radiance field representations and then performs class-agnostic 3D object detection in indoor environments. NVS approaches based on multilayer perceptrons (MLPs), 3D Gaussians, and sparse voxel-grids are extensively evaluated across multiple criteria to determine the most effective representation for 3D object detection. Two variants of the detector are developed - dense variants for superior feature extraction and sparse variants optimizing the balance between computational efficiency and detection performance. The embedded RGB-D features are extracted from the radiance field representations via efficient dense and sparse voxel-grids. The non-maximum suppression (NMS) algorithm is optimized iteratively to significantly reduce inference speed. The proposed method is validated on three challenging indoor multi-view RGB datasets (ScanNet, Hypersim, and ARKitScenes) and evaluated against state-of-the-art (SOTA) RGB-only, NeRF-based, and point-based object detection approaches. Using a smaller subset of 90 scenes from the multi-view RGB ScanNet dataset, the designed detector achieves a recall (Rec) and average precision (AP) scores of 90.0 and 45.5, respectively, at a 25% intersection-over-union (IoU) threshold. The detector’s AP performance surpasses that of existing NeRF-based detectors by approximately 17%-26% at the same IoU overlap, while performing 13% worse than existing RGB-D point-based detectors. The sparse variants, based on 3DGS and sparse voxels, provide the optimal balance between computational efficiency and detection accuracy, achieving an average inference speed of 3.4 seconds per scene on a low-end graphics processing unit (GPU) with 6GB of memory. Cross-dataset generalization experiments revealed robustness across different capture scenarios and RGB camera types, indicating that the approach captures fundamental geometric relationships. The study also established the structural similarity index (SSIM) as a reliable predictor of 3D detection performance, revealing a strong correlation between radiance field quality and detection accuracy. The modular feature extraction framework facilitates seamless adaptation to emerging radiance field representations, while the two-stage design preserves the photorealism of NVS.Item Bayesian online homography estimation and its use in multiple object trackingClaasen, Paul Johannes (University of Pretoria, 2025-06-30)A novel Bayesian homography estimation framework is proposed, which explicitly relates the homography of one video frame to the next through an affine transformation while explicitly modelling keypoint uncertainty. The literature has previously used differential homography between subsequent frames, but not in a Bayesian setting. In cases where Bayesian methods have been applied, camera motion is not adequately modelled, and keypoints are treated as deterministic. The proposed method, Bayesian Homography Inference from Tracked Keypoints (BHITK), employs a two-stage Kalman filter and significantly improves existing methods. Existing keypoint detection methods may be easily augmented with BHITK. It enables less sophisticated and less computationally expensive methods to outperform the state-of-the-art approaches in most homography evaluation metrics. Furthermore, the homography annotations of the WorldCup and TS-WorldCup datasets have been refined using a custom homography annotation tool that has been released for public use. The refined datasets are consolidated and released as the consolidated and refined WorldCup (CARWC) dataset. In addition, a novel multiple object tracking (MOT) algorithm, IMM Joint Homography State Estimation (IMM-JHSE), is proposed. IMM-JHSE uses an initial homography estimate as the only additional 3D information, whereas other 3D MOT methods use regular 3D measurements. By jointly modelling the homography matrix and its dynamics as part of track state vectors, IMM-JHSE removes the explicit influence of camera motion compensation techniques on predicted track position states, which was prevalent in previous approaches. Expanding upon this, static and dynamic camera motion models are combined using an interacting multiple model (IMM) filter. A simple bounding box motion model is used to predict bounding box positions to incorporate image plane information. In addition to applying an IMM to camera motion, a non-standard IMM approach is applied where bounding-box-based BIoU scores are mixed with ground-plane-based Mahalanobis distances in an IMM-like fashion to perform association only, making IMM-JHSE robust to motion away from the ground plane. Finally, IMM-JHSE makes use of dynamic process and measurement noise estimation techniques. IMM-JHSE improves upon related techniques, including UCMCTrack, OC-SORT, C-BIoU and ByteTrack on the DanceTrack and KITTI-car datasets, increasing HOTA by 2.64 and 2.11, respectively, while offering competitive performance on the MOT17, MOT20 and KITTI-pedestrian datasets. Using publicly available detections, IMM-JHSE outperforms almost all other 2D MOT methods and is outperformed only by 3D MOT methods---some of which are offline---on the KITTI-car dataset. Compared to tracking-by-attention methods, IMM-JHSE shows remarkably similar performance on the DanceTrack dataset, achieving 66.24 HOTA. In comparison, a variant of MeMOTR achieves 66.70 HOTA. IMM-JHSE outperforms tracking-by-attention methods on the MOT17 dataset, achieving 64.90 HOTA, where MOTIP achieves 59.2 HOTA.Item Automated cow body condition scoring using multiple 3D cameras and convolutional neural networksSummerfield, Gary Ian (University of Pretoria, 2025-08-11)Body condition scoring is an objective scoring method used to evaluate the health of a cow by determining the amount of subcutaneous fat in its body. Automated body condition scoring is becoming vital to large commercial dairy farms as it helps farmers score their cows more often and more consistently compared to manual scoring. A common approach to automated body condition scoring is to utilise a CNN-based model trained with data from a depth camera. The approach presented in this research study makes use of three depth cameras placed at different positions near the rear of a cow to train three independent CNNs. Ensemble modelling was then used to combine the estimations of the three individual CNN models. The results show that utilising the data from three depth cameras to train three separate models merged through ensemble modelling yields significantly improved automated body condition scoring accuracy compared to a single depth camera and CNN model approach.Item Enhancing voice pitch perception in cochlear implantsMaina, Sandra-Chie (University of Pretoria, 2025)Cochlear implants (CIs) are among the most successful sensory prostheses in the world, partially restoring hearing and, to many users, speech perception. However, they suffer in more complex tasks such as pitch perception. In particular, poor voice pitch perception affects CI users’ ability to perceive information necessary for contextual cues in speech, such as emotion, speaker gender, and speech intonation. This dissertation proposes three speech processing strategies designed to enhance pitch perception, and more critically, voice pitch perception. Although many pitch-enhancing schemes have improved pitch perception for CI users, their evaluation is conducted with harmonic tones or sung vowels, with added speech recognition tests to determine whether speech intelligibility is maintained. This study, instead, focuses on the perception of pitch in the context of speech to determine whether clinical standard strategies, such as ACE, are acceptable in presenting necessary voice pitch cues for non-lexical speech comprehension, as well as whether any of the proposed strategies can improve the presentation of these voice pitch cues without disrupting intelligibility or sound quality. The proposed strategies are adapted from the F0Mod strategy and are designed not only to enhance pitch but also to determine important characteristics of pitch presentation for the design of future processing strategies. The results of the perceptual experiments showed, as expected, that ACE performs well in speech recognition tasks, in quiet, providing pleasant and clear sound quality with minimal interference. However, ACE performs poorly in voice pitch perception tasks. Speech intonation recognition and voice pitch ranking capabilities are poor with this speech processing strategy, especially in noise. Therefore, ACE does not provide the necessary cues needed for full speech comprehension and must be improved. In contrast, the F0Mod strategy, particularly Milczynski et al. (2009)’s adaptation, provides good voice pitch perception, allowing for speech intonation recognition capabilities comparable to those of normal hearing listeners. The strategy is both robust to noise and does not degrade speech intelligibility. However, listeners significantly preferred the sound quality of ACE over that of Milczynski’s F0Mod strategy. In contrast, a proposed strategy, the Delay-based F0Mod strategy, was equivalent to Milczynski’s F0Mod strategy in all tasks, but was not significantly less preferred in terms of sound quality to ACE. Additionally, this strategy shows that amplitude modulation done out of phase across channels to present temporal pitch cues does not significantly degrade pitch perception, and may actually improve sound perception over Milczynski’s F0Mod strategy. In addition, the proposed Location-based F0Mod strategy showed some improvement in pitch perception over ACE, particularly at a lower fundamental frequency (F0) range. Although it is fragile in noisy environments and at higher F0 ranges, it may provide stronger temporal cues than either Milczynski’s F0Mod strategy or the Delay-based F0Mod strategy, either by aligning the cues to its corresponding tonotopic place or by only presenting the peaks of the amplitude-modulated signals, thus increasing the depth of the modulation. The strength of the presented temporal cues can be seen through the perceptual bias that was found with this strategy, where listeners perceived the pitch of the presented speech intonation stimuli as lower than with the other pitch-enhancing schemes. Additionally, speech recognition and sound quality were high with this strategy. The Location-based F0Mod strategy has shown the potential of improving pitch perception through presentation of temporal pitch cues to a selected few channels. However, it must be improved to be reliable and robust to interference. Additionally, its design can be used in future work to determine the effect of the location of the presentation of strong temporal cues by selecting different channels to modulate. Finally, while high sensitivity to pitch differences was found with the proposed Rate-based F0Mod strategy at higher frequencies, pitch perception was inconsistent with the presented stimulus and may have been limited by the proposed fundamental limit of temporal pitch. However, it may also have been an unfamiliar pitch percept that was difficult for participants to categorise, which may be improved with training. However, participants found this strategy to be particularly unintelligible and unpleasant, thereby suggesting the superiority of pitch presentation through amplitude modulation. Overall, this study has demonstrated the ineffectiveness of the ACE strategy in providing suprasegmental speech cues and the effectiveness of an amplitude modulation approach to present the necessary pitch cues. Additionally, this study has introduced three voice pitch-enhancing strategies, of which two are promising, that also inform the weight of the principles on presenting temporal pitch cues in electrical hearing and may suggest other benefits when amplitude modulation is not presented in-phase, across channels. The successful implementation of these voice pitch-enhancing strategies is, however, constrained by the successful implementation of a noise robust and real-time F0 estimator.Item Non-linear control of a fuel gas blending system with added consumer dynamicsSibiya, Mpumelelo Doctor (University of Pretoria, 2025-11-30)This dissertation contributes to existing literature on fuel gas control by providing a feasible control solution with improved economic performance for an existing fuel gas control benchmark problem. Improved economic performance is achieved by implementing a non-linear model predictive controller (NMPC) that uses state estimates provided by a moving horizon estimator (MHE) for the fuel gas composition and flame speed index (FSI) to provide continuous inputs for the controller. A new benchmark scenario is also developed which includes the effects of fuel gas consumer dynamics on the fuel gas blending system.Item Bayesian source of interest extractionHanekom, Natalie (University of Pretoria, 2025)This dissertation presents a fully Bayesian multivariate source separation model and algorithm. Full posterior model parameter and latent variable distributions were inferred using variational Bayes. The algorithm can be seen as variational Bayesian independent vector analysis for simultaneously separating multiple mixtures, such as those obtained by converting a convolutive mixture into a time-frequency domain. Key properties of the model are Gaussian mixture model source models and an explicit noise model. The algorithm was developed for guided audio source separation in a time-frequency domain. The explicit use of prior information makes Bayesian inference a natural choice in modern guided source separation algorithms, which use available information to guide algorithms to specific solutions and to achieve maximum performance for a given problem. Sources of interest were extracted from mixtures by utilising speaker-dependent Gaussian mixture model priors. The speaker-dependent models were learned from short enrolment utterances of the sources of interest, which are easily obtainable using, e.g., a mobile phone. Interfering speakers and other noise sources were jointly modelled by the noise component in the model.The algorithm was evaluated in realistic conditions with conversational speech, reverberation and diffuse babble noise from several background speakers. Speaker identification was complicated by speakers of the same gender, including related speakers (mother and daughter). The algorithm was compared to state-of-the-art source of interest extraction algorithms and achieved the best performance, frequently obtaining 10 dB better interference and noise suppression than the second-best-performing algorithm.Item A wideband high gain circularly polarized antenna with magneto-electric dipoles and a sequential phase feedMattheus, Elmien (University of Pretoria, 2025)Wideband, circularly polarized (CP) antennas are key components in communication systems due to their ability to reduce signal loss of transmitted and received electromagnetic waves, leading to enhanced data transfer over various distances. Over the years, different wideband CP antennas such as magneto-electric (ME) dipole antennas, crossed dipole antennas, spiral antennas and slot antennas have been developed. From literature, it is evident that ME dipole antennas are great CP antenna candidates to realise reliable communication systems due to their excellent radiation performance of wide impedance bandwidth (IBW), wide axial ratio bandwidth (ARBW) or high gain while maintaining a simple structure that can be manufactured with ease. However, a limitation exists to realise CP ME dipole antennas that simultaneously obtain wide usable bandwidth and high, stable gain without implementing complex structures, dual-feeding networks, large arrays, parasitic elements and/or additional cavities. Following a comprehensive investigation into wideband, CP antennas, the design process consisted of an antenna synthesis approach to design a printed, CP ME dipole radiator by combining elements of a previously designed CP ME dipole antenna and a printed, LP ME dipole antenna, integrating them into a new antenna structure and applying different modifications to the antenna structure to form a single CP ME dipole radiator. Simulations were executed to determine the radiation performance of the initial single CP ME dipole radiator. Furthermore, a parametric study was used to determine which dimensions affected the radiation performance and optimisation was conducted using these results to find the optimal dimensions of the single CP ME dipole radiator that led to good radiation performance of wide bandwidth and high gain. An sequential phase (SP)-feed was designed and optimised to obtain wide bandwidth with low magnitude and phase differences at each adjacent port. Afterwards, the SP-feed was combined with four of the CP ME dipole radiators to form a CP ME dipole antenna array. Simulations were executed to determine the radiation performance of the initial CP ME dipole antenna array. Another parametric study was used to determine which dimensions affected the radiation performance and the results were used to perform optimisation of the dimensions to find the final CP ME dipole antenna array. To confirm ease of manufacturing, a sensitivity analysis was conducted to determine the effect of the critical build parameter on the radiation performance. A prototype of the final CP ME dipole antenna array was built, assembled and measured in the compact range at the University of Pretoria. The final CP ME dipole antenna array showcased an IBW of 74.6%, a 3 dB ARBW of 68%, and a gain variation of 10.1 ± 2.1 dBic over the usable bandwidth with stable radiation patterns, well-formed main beams and low cross-polarization levels. A final size of 1.33λ 0 × 1.33λ 0 × 0.24λ 0 was achieved with a simple geometry and no additional cavities or parasitic elements. The designed CP ME dipole antenna array has good gain, wide bandwidth and outperforms other single-fed CP antennas available in literature i.t.o structure simplicity, wide radiation performance and improved gain.Item FPEVO : fused point-edge visual odometry for low-structured and low-textured scenesBrown, Dylan (University of Pretoria, 2025-11-25)Simultaneous localisation and mapping (SLAM) is the process by which an agent, such as a robot, creates a map of the environment it traverses while simultaneously determining its position relative to the generated map. Various solutions have been proposed to solve the SLAM problem, with visual SLAM methods emerging as a prominent field of research. By using visual information, visual SLAM approaches provide feature-rich representations of the environment while primarily relying on inputs from cameras, which has enabled wide accessibility and adoption. At the heart of visual SLAM lies the visual odometry component. Visual odometry is the process by which the pose of an agent is estimated using the provided visual information. Visual odometry methods provide locally accurate pose and map estimations, while the incorporation thereof in a full SLAM system aims to make the jointly estimated poses and map globally consistent. A primary limitation of existing visual odometry approaches is their inability to achieve satisfactory performance in both high- and low-textured, and well- and low-structured regions. Existing systems only cater to a subset of the aforementioned region types. To perform accurate vision-only pose estimation in both low- and high-textured and low- and well-structured regions, a robust RGB-D visual odometry method is proposed that fuses point and edge features. By combining the descriptiveness of point features with the structure provided by edge data, a method that is robust to both low-textured and low-structured scenes is developed. This is achieved using a multi-stage pipeline. Edge features are first detected and grouped based on the Gestalt principles of similarity and proximity. Edge groups are then associated between the current and previous frames. Edge pixels are matched between the associated edge groups using the structural constraints imposed by the edges. These matches are then used to estimate the motion of the agent. The developed visual odometry method is called FPEVO. Compared to state-of-the-art alternatives, FPEVO reduces the root mean square absolute trajectory error, and translational and rotational relative pose errors, by up to 71%, 81%, and 86%, respectively. It was found that the proposed method is not only more accurate than current approaches, but also more consistent, especially in low-structured and/or low-textured environments. Although the proposed method uses an RGB-D sensor, its architecture was designed to be extendable to other sensor types, such as monocular and stereo cameras.Item Modelling and characterisation of signal fading behaviour in hybrid powerline-wireless communication channelsMokise, Kealeboga (University of Pretoria, 2025)New-generation radio networks are often accompanied by increased usage of spectral and temporal resources. As a result, interaction between the propagating radio signal and the propagation environment becomes increasingly complex. This inherent complexity of signal propagation has prompted the development of channel models that capture the stochastic nature of the channel that arise from physical phenomena such as shadowing, clustering of multipath propagation components, power variation between the line-of-sight (LOS) and non-line-of-sight (NLOS) multipath propagation components and the simultaneous fading effect shadowing and multipath fading, commonly referred to as composite fading. Composite fading arises when the instantaneous variations of the received signal strength caused by multipath fading are superimposed on the long-term variations caused by shadowing. Composite fading provides a realistic account for the signal strength variation behaviour of the received signal, and therefore, it is crucial to characterise such fading behaviour to analyse the performance of radio communication systems. Various composite fading models have been proposed to characterise the simultaneous impact of multipath and shadowing in various radio communications scenarios. Composite fading models can be broadly classified into LOS shadowing, which describes shadowing of the dominant signal component, and multiplicative shadowing, which describes propagation conditions whereby both the LOS component and scattered signal components of the received signal are jointly shadowed. In the current literature, composite fading, including the other aforementioned physical phenomena of communication channels, has been intensively studied and validated in wireless communication (WLC) scenarios. The powerline communication (PLC) channel is a known medium of radio signal propagation, and it exhibits propagation characteristics which are apparent in wireless channels, such as multipath propagation and signal propagation path-loss. These common propagation characteristics shared by the PLC and WLC channels facilitate the propagation of radio signals between the two channels with minimal signal coupling requirements, which means a hybrid PLC-WLC communication system can be established by integrating the PLC and WLC technologies to leverage the different spectral resources and diversity between powerline and wireless channels. In the current literature, previous works on PLC-WLC channel characterisation primarily focus on the frequency-selective behaviour of such channels, and PLC-WLC propagation environments are typically investigated with both the transmitter and receiver remaining stationary during channel measurements. Therefore, there is a significant lack of time-selective behaviour characterisation and characterisation of hybrid PLC-WLC propagation scenarios involving relative motion between the PLC and WLC devices. Moreover, there is currently no channel model that describes the unified powerline-wireless propagation environment. Having identified these knowledge gaps and other limiting factors of hybrid PLC-WLC communication systems, this thesis focuses on three main fronts to address these gaps and limitations. The first contribution of this thesis deals with the multipath clustering problem in PLC and hybrid PLCWLC communication channels with the objective of characterising multipath clustering behaviour in indoor stationary PLC and hybrid PLC-WLC communications channels using a model-based clustering approach. The multipath clustering problem formulation is considered in wideband communication scenarios, where multipath components arrive in clusters which share similar parameters such as the arrival time delay. Using a sequence-based wideband channel sounding method, channel measurements with high-temporal resolution are obtained. An eigenvalue estimator is used to determine the number of multipath components and the space-alternating generalised expectation-maximization (SAGE) algorithm is used to extract magnitude-delay parameters of multipath components. A method for estimating a feasible range of clusters is proposed and applied in both the distance-based and model based clustering approaches. The model-based framework employs a range of finite-mixture models(FMMs) to identify clusters, while distance-based approaches use methods such as kMeans. A maximum likelihood (ML) approach is used for ftting the FMMs to the extracted multipath components, which proved to be efficient in estimating parameters for both closed-form and update-expression solutions of the likelihood functions. The corrected Akaike’s information criterion (AICc) metric was used to determine the best-ft FMM to the measurements, while the Davies-Bouldin (DB) and Calinski-Harabasz (CH) validation indexes are used to compare the model-based to the distance-based clustering methods. The obtained results indicate that both PLC and hybrid PLC-WLC communication channels present clusters of multipath components. Moreover, the results show that the model-based clustering obtains the best performance in terms of within-cluster compactness and between-cluster separation in the delay domain compared to distance-based methods. The second contribution of this thesis deals with the investigation and characterisation of fading behaviour in hybrid PLC-WLC channels with the objective of characterising the short-term (multipath) and long-term (shadowing) fading behaviour observed in indoor hybrid PLC-WLC communications channels for narrowband communications involving relative motion between the PLC and WLC devices. Novel insights into the problem formulation of composite fading behaviour in the context of PLC-WLC communication systems, considering relative motion between the PLC and WLC devices, are provided. Two new long-term fading models, namely, Gamma-Rayleigh (GR) and inverse Gamma Rayleigh (IGR), are proposed. The utility and validity of the proposed GR and IGR fading models are demonstrated through a proposed procedure of approximating the popular and well-established long-term fading models, such as the lognormal and inverse Gaussian shadowing models, using method-of-moments (MoM) estimators approach, which demonstrated excellent approximation of the probability density function (PDF) statistics. An extensive measurement campaign is carried out considering various hybrid PLC-WLC propagation scenarios classified by the powerline branching characteristics of the PLC portion of the hybrid PLC-WLC channels. Then, the Generalised Lee method is used to separate short-term and long-term fading components from the received signal envelope. The estimated long-term fading signal components are fitted to the lognormal, gamma, inverse gamma, inverse Gaussian, and the proposed GR and IGR long-term fading models. The estimated short-term fading signal components are fitted to the Rayleigh, Rician, κ-µ and Nakagami candidate short-term fading models. The obtained results indicate that the amount of short-term and long-term fading increases as more branches or discontinuities (non-idealities) are added to the PLC portion of the hybrid PLC-WLC channel. This results in increased instantaneous signal strength fluctuations and increased random mean power fluctuations in the received signal, and consequently increased composite fading behaviour of the hybrid PLC-WLC channel. The third contribution of this thesis deals with the investigation and characterisation composite fading behaviour, that is, the combined effect of short-term (multipath) and long-term (shadowing) fading, in hybrid PLC-WLC channels with the objective of proposing a new and general composite fading models to characterise composite fading behaviour observed in indoor low-voltage PLC-WLC communications scenarios involving relative motion between the PLC and WLC devices. Two new general composite fading models are proposed, namely, the κ-µ / Gamma-Rayleigh (κ-µ / GR) model and the κ-µ / inverse Gamma-Rayleigh (κ-µ / IGR) model. The κ-µ / GR is a LOS shadowing model which accounts for composite fading scenarios where shadow fading is assumed to only affect the dominant signal components, while the κ-µ / IGR is a multiplicative shadowing model which accounts for scenarios where shadowing is assumed to affect both the dominant and scattered signal components. For both models, analytical expressions are derived for the probability density function (PDF), cumulative distribution function (CDF), amount-of-fading (AF), general moments and the moment generating function (MGF) of the models, along with performance metrics such as outage probability (OP), average symbol error probability (ASEP) and average channel capacity. The validity of the derived expressions is confirmed through Monte Carlo simulations. An extensive measurement campaign is carried out considering various hybrid PLC-WLC propagation scenarios classified by the powerline branching characteristics of the PLC portion of the hybrid PLC-WLC channels, and the empirical data is fitted to the proposed models and other existing composite fading models using a non-linear fitting process. The obtained results indicate that increased branching, discontinuities and terminations in the PLC portion of a hybrid PLC-WLC channel result in enhanced multipath and shadowing effects, thereby intensifying the composite fading behaviour observed in such PLC-WLC systems. Moreover, the obtained results indicate that the appropriateness of a LOS or a multiplicative composite fading model depends on the strength of dominant signal components and the power disparity between dominant and scattered signal components. Consequently, the proposed LOS and multiplicative shadowing models proved to be generalised and necessary for accurately modelling time-selective fading across various PLC-WLC signal propagation scenarios.Item Model predictive static programming control applied to mineral processing plantsNoome, Zander Meindert (University of Pretoria, 2023-05)In a mineral processing plant, the separation of valuable material from ore has multiple stages. Usually, the ore is crushed or ground into smaller parts through multiple crushers or grinding mills. This is called the communition process. This process is typically the first stage for extracting valuable material and is important for further down-stream processes. The output of the communintion stage is usually regulated to achieve a stable throughput and a specific ore particle size. After the ore is crushed and ground to a specified size, the valuable material in the ore needs to be separated from the undesired materials. The properties of the desired material influence the method used for separation. These methods include froth flotation, gravitational separation, magnetic separation and electrostatic separation. The separation process can include multiple process streams to get a high grade of the desired minerals out of the ore. In froth flotation, the main objective is to extract the desired material from the ore to obtain a large mineral recovery. Because the flotation process relies on the flotation of particles, particle size is extremely important. The use of control systems in mineral processing plants has been adopted to improve throughput, optimize power usage, ensure safe process operation and to running at a stable operating condition. The control of these plants makes use of different advanced process control strategies which include but are not limited to cascaded control, where multiple layers of control systems are applied, and model predictive control. These different control strategies can range from regulatory control to supervisory control. Because of the large number of inputs to these plants, efficient controllers are necessary to obtain desired results. The use of Nonlinear Model Predictive Control (NMPC) is an attractive option for most mineral processing plants because of the constraint management capabilities of the controller. Unfortunately, the NMPC method has a large computational load which requires sufficient resources to make it a viable option. Another model predictive control method known as Model Predictive Static Programming (MPSP) has shown promise to improve the computational time of a standard NMPC controller. The MPSP control philosophy generates a static optimization problem which is less computationally difficult to solve compared to the dynamic optimization problem that is generated through NMPC. In this dissertation, the control of a single-stage grinding mill circuit and a four-cell flotation circuit with an MPSP controller to reduce the computational load is proposed. The computational efficiency and the output performance of MPSP controllers are compared to NMPC controllers as a motivation for the use thereof. The comparison is done by simulating two mineral processing stages, namely the communition phase and the separation phase. The simulations considered different configurations for both the MPSP and NMPC controllers. The comparison of the controllers in the simulations shows that the MPSP controller obtained similar or improved plant results while also having a reduced computational time compared to the NMPC controller. The MPSP controller also displays scalability improvements compared to the NMPC controllers which can be beneficial for supervisory control of large-scale processing plants.
