2026
Journal Articles
A C-RAN-Based Mitigation of Reactive Jamming Attacks in LoRa Networks with Jamming Cancellation
LoRa modulation has gained considerable attention in wireless Internet of Things applications, from sensor networks to smart systems, due to its long-range and low-power capabilities, and it is currently one of the major technologies used in low-power wide-area networks. However, like many wireless technologies, LoRa-based networks remain vulnerable to jamming attacks, which can lead to denial-of-service. This paper proposes a technique to mitigate jamming attacks by estimating and canceling the jamming signal in scenarios where a gateway is jammed by a nearby adversary. Additionally, the benefits of a centralized radio access network architecture are explored for networks with multiple gateways and jammers, enabling post-cancellation demodulation at a central unit to enhance the legitimate signal against residual jamming signals. The simulation results show that this approach achieves an overall frame success rate (FSR) of more than 90% in areas covered by multiple gateways, with the centralized processing providing up to 1.5× gain over the standard distributed anti-jamming approach at low signal-to-noise ratio (SNR). Increasing the code rate from 4/5 to 4/7 further improves the FSR, up to 3× at low SNR and 1.5× at high SNR, at the cost of an increase in frame air-time. Finally, jammers located within 20 m of a gateway significantly degrade the achieved FSR, keeping it below 55% even with a code rate of 4/7 and the centralized processing, suggesting that gateway placement should prioritize restricted-access locations.
IEEE Open Journal of the Communications Society. 2026. DOI : 10.1109/OJCOMS.2026.3715389.Conference Papers
HDPC Codes with LDPC Matrices: Construction Based on Social Golfer Problem
In contrast to conventional low-density parity-check (LDPC) code construction, we propose to construct a sparse decoding matrix with auxiliary variable nodes (AVNs) and then derive the corresponding high-density parity-check (HDPC) code for encoding. To leverage the sparsity gain of AVNs while controlling AVN-induced harmful substructures, we adopt a parity-check row block (PCRB) structure and show that arranging PCRBs to obtain a 4-cycle-free decoding matrix forms a class of social golfer problems (SGPs). Focusing on an explicit affine-plane-based SGP solution, we construct a family of AVN-aided affine plane (AAP) codes. For the same blocklength and minimum distance, AAP codes achieve equal or greater dimension than Reed-Muller codes and exhibit low error floors under iterative decoding, which extend finite-geometry LDPC constructions to additional code families with 4-cycle-free decoding matrices.
2026. 2026 IEEE International Symposium on Information Theory, Guangzhou, China, 2026-06-28 – 2026-07-03. DOI : 10.1109/isit62367.2026.11653623.Dynamic Dual-Window Decoding for SC-LDPC Codes with Wave Enhancement
Windowed decoding (WD) based on the belief propagation (BP) algorithm, which slides along a code chain with fixed computational resources, is an efficient decoding scheme for spatially-coupled low-density parity-check (SC-LDPC) codes, particularly for long-length codes. However, due to the lack of decoding waves, WD suffers from a performance degradation compared to full BP decoding. In this paper, we propose a novel decoding schedule for WD, termed dual-WD, which enhances the decoding waves by dynamically merging and splitting two decoding windows that slide in opposite directions. Moreover, a sliding-step selection method based on partial syndrome checks is introduced to reduce decoding complexity. Simulation results demonstrate that the proposed dual-WD outperforms the conventional WD schemes under identical computational complexity constraints, thereby narrowing the performance gap between WD and full BP decoding.
2026. 2026 IEEE International Symposium on Information Theory, Guangzhou, China, 2026-06-28 – 2026-07-03. DOI : 10.1109/isit62367.2026.11654081.CIM-FLEX: An Integer-Only Flexible Periphery for Distributed Compute-In-Memory Architectures
Distributed compute-in-memory (CIM) architectures are emerging as a path towards energy-efficient and highly parallel edge-class deep neural network (DNN) inference. However, the lack of a flexible, unified digital periphery limits the ability of distributed CIM systems to support integer-quantized DNN inference schemes, activation functions, and scalable tile-to-tile dataflow. We present CIM-FLEX, an integer-only periphery architecture that enables each CIM tile to autonomously perform partial-sum rescaling, asymmetric activation processing, and lightweight programmable activation functions. CIM-FLEX also supports mixed-precision integer multiply-and-accumulate (MAC) by decomposing higher-precision operations into uniform-precision partial MACs, without modifying existing CIM primitives. CIM-FLEX enables flexible tile-to-tile dataflow across distributed CIM tiles, exposing a tradeoff between parallel output generation and energy efficiency, where deeper accumulation chains reduce concurrent periphery utilization and improve overall efficiency. Numerical studies and TFLite-based LUT activation experiments validate a high signal-to-noise ratio (∼30 dB) for programmable activation functions, and a 14 nm CMOS implementation shows a 0.012 mm 2 area for the CIM-FLEX.
2026. ISLPED ’26: ACM/IEEE International Symposium on Low Power Electronics and Design, Evanston, IL, USA, 2026-08-05 – 2026-08-07. p. 1 – 7. DOI : 10.1145/3816440.3818582.A 2.35mm2 118.3Gbps 5.14pJ/Bit Dual-Mode LDPC Decoder for 5G/6G
A dual-mode high-throughput LDPC decoder for 5G/6G is presented. The 2.35mm2 decoder in 16 nm FinFET delivers a peak throughput of 118.3 Gbps with an energy efficiency of 5.14pJ/ bit, supporting up to 10-core parallel 5 G-LDPC decoding as well as spatially-coupled LDPC codes for 6G with code lengths up to 5x of the longest 5G-NR standard codes. Compared with the state-of-the-art, this work achieves 12.4x in peak throughput, 1.7 x in area efficiency, and 2.3 x in energy efficiency. In 6G mode, it further provides up to 1 dB coding gain advantage over 5G-LDPC codes.
2026. 2026 IEEE/JSAP Symposium on VLSI Technology and Circuits (VLSI Technology and Circuits), Honolulu, HI, USA, 2026-06-14 – 2026-06-18. p. 1 – 3. DOI : 10.1109/vlsitechnologyandcir65830.2026.11577322.Top-Metal-Only RFIC Retargeting for Fast Specs-To-Silicon Iteration Enabled by AI-Assisted Inverse Design
AI and machine learning (ML) are increasingly used to accelerate Radio-Frequency Integrated Circuit (RFIC) design, where specs-to-layout inverse-design automation can shorten design cycles and lower the expertise barrier. In practice, RFIC turnaround and design migration are constrained not only by iterative EM-driven passive design, but also by layout-to-silicon latency dominated by long fabrication cycles and the cost of full-mask preparation. To address these challenges, this paper presents a fast specs-to-silicon iteration framework for RFIC retargeting by reusing all lower-metal layers that define active devices and tuning elements (e.g., capacitors and resistors), while redesigning EM passives implemented in the reconfigurable top-metal stack above a reference mid-metal layer. We employ an AI-assisted template-seeded approach that enables fast inverse design of multi-metal-layer EM passives. As a proof of concept, two FR3 power amplifiers (PAs) at 13 and 20 GHz are designed in the GlobalFoundries 22 nm FDX+ using top-metal-only routing and pixelated matching networks. PA1 measures an OP1 dB of 18.75–21.74 dBm, a PAEOP1dB of 19.84%–32.65%, a Psat of 19.89–22.13dBm, and a PAEsat of 19.89%–33.66% across 12.5–17 GHz. PA2 measures an OP1 dB of 16.00–17.68 dBm, a PAEOP 1 dB of 18.70%–23.34%, a Psat of 17.70–18.55 dBm, and a PAE sat of 20.45%–24.10% across 19–24 GHz.
2026. 2026 IEEE Radio Frequency Integrated Circuits Symposium, Boston, MA, USA, 2026-06-07 – 2026-06-09. p. 115 – 118. DOI : 10.1109/rfic70222.2026.11602335.Proof of Concept of QCSP Frames in Earth-to-LEO Satellite Transmission
In this paper, we report on a multi-site experiment involving the transmission of a new type of frame for IoT, called Quasi-Cyclic Short Packet (QCSP). QCSP frames are characterized by the combination of a q-ary modulation using Cyclic Code Shift Keying (CCSK) and a q-ary non-binary Error Correction Code (ECC). A QCSP frame does not contain a preamble, making it energy-optimal since no resources are wasted for signaling. However, this absence poses a significant challenge for detection and synchronization, particularly in the case of a LEO satellite, which is subject to strong and fast-varying Doppler shifts. Our experiment demonstrates that such transmission can be successfully recovered.
2026. 2026 IEEE International Conference on Communications Workshops (ICC Workshops), Glasgow, United Kingdom, 2026-05-24 – 2026-05-28. p. 1 – 6. DOI : 10.1109/iccworkshops63917.2026.11586359.A Full-Custom Time-Domain Unary Sorter for Soft-Information Decoding in 65 nm CMOS
We present a full-custom time-domain (TD) unary sorter for soft-information decoding with 5-bit quantized inputs. The design exploits a time-domain unary representation combined with a bitonic sorting network to efficiently order large numbers of log-likelihood ratios (LLRs) while outputting only their indices. Implemented in a 65 nm CMOS process, the proposed sorter achieves substantially lower latency compared to conventional digital and unary designs, while also reducing area and power relative to digital implementations, making it well-suited for efficient ultra-low-latency soft-decision decoding. Furthermore, the time-domain unary representation allows trading a small amount of sorting accuracy for even lower latency, which is acceptable in many soft-decision decoding applications such as ordered reliability bits guessing random additive noise decoding (ORBGRAND) and Chase-type error-correcting decoders.
2026. 2026 IEEE International Symposium on Circuits and Systems (ISCAS), Shanghai, China, 2026-05-24 – 2026-05-27. p. 1417 – 1421. DOI : 10.1109/ISCAS66217.2026.11562817.Theses
Enabling Reliable and Scalable LoRa with Collision Resolution and Centralized Radio Access
Numerous Internet of Things (IoT) applications depend on energy-efficient, low-maintenance devices that transmit short messages over long distances. These devices often operate from batteries or harvested energy and must remain connected in locations where dense infrastructure deployment is impractical. Low-power wide-area networks (LPWANs) address these requirements by trading high throughput and low latency for long range, low cost, and low energy consumption. Among LPWAN technologies, LoRa is widely used because its chirp spread spectrum modulation enables robust communication at very low signal-to-noise ratio (SNR), while simple uplink operation keeps end-device complexity and cost low. Despite LoRa’s ability to operate at very low SNR, uplink frames near the coverage limit may still be received close to the gateway sensitivity threshold. In this operating regime, noise and interference increase the risk of missed detections and decoding failures, while clock and carrier-frequency inaccuracies in low-cost radios degrade the synchronization and, ultimately, the demodulation performance. Because LoRaWAN uses largely uncoordinated uplink access, the probability of packet collisions increases with the number of active devices. Reception failures and collisions trigger retransmissions, which increase end-device energy consumption and channel load. As IoT adoption continues to accelerate, these effects increasingly threaten the reliability, scalability, and overall performance of LoRaWAN and comparable IoT networks. Given the large installed base of LoRa end nodes, improving reliability and scalability through device modifications or channel-access changes is impractical. In this thesis, we investigate how receiver-side signal processing at gateways and centralized processing across multiple receivers can improve LoRa uplink coverage, reliability, and network scalability while preserving compatibility with existing end nodes. We first establish an interoperable LoRa physical-layer model and software-defined radio implementation that reproduce the low-level conventions required for communication with commercial transceivers and provide a reproducible baseline for LoRa receiver research. On this basis, we address missed detections and synchronization errors through a preamble detection method that is robust to rapid variations in noise and interference power, together with a synchronization algorithm that explicitly accounts for sampling frequency offset. We then treat packet collisions as a structured multi-user reception problem, showing that overlapping LoRa frames can be reliably recovered through successive interference cancellation when combined with iterative soft-decision decoding (SDD). To exploit receiver diversity beyond independent gateway processing and enable a centralized allocation of baseband processing resources, we develop a cloud radio access network (C-RAN) architecture for LPWANs. We implement and validate the proposed software-defined radio (SDR)-based architecture through a measurement campaign that also produced a reusable dataset of multi-receiver IQ samples. Within this architecture, we further derive a non-coherent likelihood combining method that exploits spatial diversity without requiring strict phase coherence among distributed receivers. Overall, the thesis demonstrates that gateway-side receiver enhancements and centralized processing provide a practical path toward more reliable and scalable
Lausanne, EPFL, 2026.2025
Journal Articles
Turbo Product Code-aided Spatiotemporal 2-d Coded Mimo for Urllc
The ultra-reliable low-latency communication (URLLC) scenario beyond the fifth-generation mobile networks requires latency to shift from ms-level to ms-level while maintaining low packet error rates. Short block-length codes can reduce latency but suffer from performance degradation. Asymptotic analyses indicate that increasing spatial degrees of freedom in multiple-input multiple-output (MIMO) systems help to achieve target error rates with a fixed block length and signal-to-noise ratio. To approach the upper bound of the average maximal achievable information rate in massive MIMO systems, spatiotemporal 2-D channel coding has been proposed. However, a considerable gap remains between the performance of the existing coding scheme and the theoretical capacity bound. To bridge this gap, this article introduces turbo product codes aided spatiotemporal 2-D coded MIMO systems (TPCAS-MIMO). Multi-level design on the baseband MIMO transceiver, involving encoding, interleaving, modulation, detection, and decoding schemes, is explored to boost the performance of TPCAS-MIMO with low latency for URLLC. In comparison to prior arts, this work achieves a 5.5 dB SNR gain at a frame error rate of 10(-5) using (64, 42)(2) polar turbo product codes in a 64 x 128 MIMO system, narrowing to a mere 2.5 dB gap from the theoretical capacity bound.
IEEE COMMUNICATIONS MAGAZINE. 2025. DOI : 10.1109/MCOM.001.2400618.GenPolar: Generative AI-Aided Complexity Reduction for Polar SCL Decoding
The CRC-aided successive cancellation list (CA-SCL) decoding algorithm for polar codes has gained widespread adoption thanks to its outstanding performance. However, with the evolution of 6G technologies, the high complexity of CA-SCL decoding poses a challenge in meeting growing performance requirements. Consequently, it is crucial to devise strategies that reduce this complexity without compromising error rates. Current efforts to mitigate the complexity mainly depend on harnessing special nodes associated with the code construction sequences, such as Fast-SCL decoding. However, these strategies suffer from redundant complexity due to ill-suited construction sequences and unnecessary sorting operations within special nodes. Addressing this issue, this paper proposes a hardware-friendly and GenAI-aided complexity reduction approach for Fast-SCL decoding, named GenPolar. This approach involves two-step optimization techniques: 1) Transformer encoder models for generating polar construction sequences, and 2) a sorting entropy based method for sorting reduction. These two-step techniques result in reduced complexity with negligible performance loss. For polar codes of length-1024 with code rates of 0.25, 0.50, and 0.75, Gen-Polar achieves latency reductions of 20.6%, 29.8%, and 40.6%, respectively. Even benchmarking against the reduced-complexity version of Fast-SCL decoding, the relative gains are 14.0%, 17.8%, and 22.3%, respectively. It should be noted that the immediate application is not limited to Fast-SCL decoding but also extends to other node-based SCL decoding algorithms like SSCL-SPC and SR-SCL.
IEEE Journal on Emerging and Selected Topics in Circuits and Systems. 2025. DOI : 10.1109/jetcas.2025.3561330.Toward Universal Belief Propagation Decoding for Short Binary Block Codes
Belief propagation (BP) decoding has been recognized for its capacity-approaching performance and high throughput when decoding long low-density parity-check (LDPC) codes. However, the application of BP decoding for short codes is hindered by dense parity-check matrices (PCMs) and prevalent short cycles in the Tanner graph. In this paper, we introduce a general method to extract an optimized sparse PCM for short binary block codes, which removes length-four cycles and enhances the connectivity of short cycles to enable BP decoding with improved performance. Notably, for short binary codes with lengths up to 64, our BP decoding performance approaches the maximum likelihood bound and surpasses the best-reported BP results with reduced computational complexity. Compared with other universal decoding algorithms, BP decoding using our extracted sparse PCMs is competitive in terms of both error-rate performance and computational complexity. These promising results suggest that our method to improve BP decoding for short codes is a step toward a practical universal BP decoder for next-generation communication systems.
IEEE Journal on Selected Areas in Communications. 2025. Vol. 43, num. 4, p. 1135 – 1152. DOI : 10.1109/JSAC.2025.3536505.Edge-Spreading Raptor-Like LDPC Codes for 6G Wireless Systems
IEEE Transactions on Communications. 2025. p. 1 – 1. DOI : 10.1109/tcomm.2025.3576907.Device-Free Floor-Scale Human Detection With Indoor LTE Antennas
This paper introduces a novel method for device-free human detection by leveraging existing wireless communication signals from 4G-long-term evolution (4G-LTE) systems. By utilizing the pervasive 4G-LTE signals, our approach enhances the efficiency and coverage of human presence detection compared to WiFi signal based approaches. A previously overlooked, but crucial human presence scenario involving subtle human activities is successfully addressed and detected. Effective human presence detection relies heavily on precise feature extraction from channel estimates and careful feature selection. Through a detailed analysis and comparison of features discussed in previous work, along with the introduction of new features, we develop a machine learning-based approach to identify the most effective features for detecting human presence. Our machine learning model, trained with these selected features, is tested across different buildings and various scenarios using a commercial 4G-LTE network. The results demonstrate that our selected features significantly enhance detection accuracy and robustness, outperforming features introduced in previous literature across diverse environments.
IEEE Open Journal of the Communications Society. 2025. DOI : 10.1109/OJCOMS.2025.3560153.HDD-RAM: A 40-nm 0.35V 25MHz Half-Select Disturb-Free Memory with Data-Aware 10T SRAM
Ultra-low-voltage SRAM is an indispensable component that is increasingly adopted in energy-efficient computing systems. However, it comes at the cost of increased sensitivity to soft errors. To address this issue, bit-interleaving SRAM is widely used to mitigate soft errors. But it suffers from halfselect disturbance. Previous works address such disturbance by using a dedicated write port or enhanced write assist scheme. However, these works may decrease write margin, induce high cell-level write latency, or incur architecture-level time/timing overhead. In this paper, we develop a high-speed bit-interleaving half-select disturb-free memory with data-aware 10T SRAM. First, we present an isolated and decoupled topology with dedicated write control to improve stability. Second, we present a data-aware write path with enhanced write-ability that effectively reduces the write access time. A 40-nm 4-Kb test chip has been fabricated to validate the optimizations above. Measurement results show that our half-select disturb-free test chip achieves a peak operating frequency of 25 MHz and an energy consumption of 0.168 fJ/bit with a supply voltage of 0.35 V. Compared with the state-of-the-art designs, it has achieved a speed up of 2.72× and an energy saving of 93.8%.
Integrated Circuits and Systems. 2025. p. 1 – 12. DOI : 10.23919/ics.2025.3565481.TRIP: Tree Search-Based Iterative Detection and Decoding for Polar-Coded MIMO Systems
With the rapid development of wireless communications, future applications call for an increase in reliability performance. Thanks to the soft information exchange between detection and decoding, iterative detection and decoding (IDD) exhibits superior performance over separate detection and decoding (SDD). In this letter, we introduce TRIP, a tree search-based IDD receiver designed for polar-coded multiple-input multiple-output (MIMO) systems. To enhance the error rate performance of TRIP and reduce its complexity, we utilize the L metric selection evolutionary multiobjective optimization algorithm (SMS-EMOA). This algorithm refines the extrinsic information and the layer width of detectors and decoders, achieving an improved balance between performance and complexity. Simulation results indicate that TRIP outperforms its SDD mode by 1.38 dB at a frame error rate (FER) of 10-3, and gains an additional 0.68 dB by fine-tuning the extrinsic information. The complexity of the detector and the decoder is reduced by 56.3% and 37.8% without compromising their performance gains, respectively.
IEEE Wireless Communications Letters. 2025. DOI : 10.1109/LWC.2024.3506655.Conference Papers
A Universal LDPC Matrix for Short-Length Codes and Long-Length Codes
6G demands an evolution of the current 5G New Radio (NR) channel codes. For eMBB+ scenarios, block lengths defined in 5G become inefficient for extremely high data rates, and a larger transport block size is thus desirable. At the same time for URLLC+ scenarios, short code block lengths are preferred for low latency, but then achieving a target frame error rate below 10−5poses a significant challenge for short-length low-density parity-check (LDPC) codes. In this work, we propose a universal LDPC (Uni-LDPC) matrix that supports both short-length and long-length codes. By using a concept of flexible coupling, the proposed Uni-LDPC code integrates the benefits of block codes and coupled codes. This unified design enables a single LDPC matrix to perform effectively across both eMBB+ and URLLC+ scenarios. Simulation results demonstrate that Uni-LDPC codes outperform 5G-NR LDPC codes over a wide range of code rates at a short block length (a few hundred information bits) with a high number of decoding iterations. Meanwhile, thanks to coupling, Uni-LDPC codes can also surpass 5G-NR LDPC codes at long code lengths (tens of thousands of information bits) with a low number of iterations.
2025. 2025 IEEE Globecom Workshops (GC Wkshps), Taipei, Taiwan, 2025-12-08 – 2025-12-12. p. 653 – 658. DOI : 10.1109/gcwkshps68340.2025.11590957.Decoding Product Codes with Belief Propagation
Product codes are considered a promising coding scheme for future optical and wireless communications. However, existing soft-input soft-output (SISO) iterative decoding for product codes involves sorting, list management, and soft-output generation, which includes hardware-unfriendly operations and dataflow. In this paper, we apply belief propagation (BP) decoding to product codes, which leverages the sparse factor graphs of the component codes. As an inherently SISO approach, the proposed product BP decoder can directly propagate soft messages across iterations with a hardware-friendly parallel structure, which outperforms the traditional Chase-Pyndiah decoding with a list size of 64 for various code rates.
2025. 2025 IEEE Workshop on Signal Processing Systems (SiPS), Hong Kong, 2025-11-01 – 2025-11-04. p. 1 – 5. DOI : 10.1109/sips66314.2025.11261227.A 5.9 μJ/Inference Reconfigurable Accelerator for Hybrid SNN-ANN Architectures in Event-based Image Recognition
Event-based images generated by event-based cameras record a stream of events with spatiotemporal information, which is different from static images. Although spiking neural networks (SNNs), which are event-based models, can be used to process event-based images naturally, they suffer from low efficiency in the training, compared to artificial neural networks (ANNs). Hybrid SNN-ANN architectures are a promising alternative to achieve both high performance and low latency. In this paper, we propose an efficient edge accelerator for hybrid SNNANN architectures. A fused reconfigurable PE array is proposed to support both SNN layers with time-parallel dataflow and ANN layers. An offline output rearrangement scheme is further proposed to improve the sparsity utilization of the ANN layers with low performance degradation and extra cost. The proposed edge accelerator is implemented with 28 nm CMOS technology, running at 400 MHz. The accelerator achieves the accuracy of 94.44% on DVS gesture dataset, and a throughput of 6,711inf/s with 5.9μ J/inf energy consuming.
2025. 2025 IEEE Biomedical Circuits and Systems Conference, Abu Dhabi, United Arab Emirates, 2025-10-16 – 2025-10-18. p. 150 – 154. DOI : 10.1109/biocas67066.2025.00042.Impact of Reactive Jamming Attacks on LoRaWAN: a Theoretical and Experimental Study
This paper investigates the impact of reactive jamming on LoRaWAN networks, focusing on showing that LoRaWAN communications can be effectively disrupted with minimal jammer exposure time. The susceptibility of LoRa to jamming is assessed through a theoretical study of how the frame success rate is impacted by only a few jamming symbols. Different jamming approaches are studied, among which repeated-symbol jamming appears to be the most disruptive, with sufficient jamming power. A key contribution of this work is the proposal of a software-defined radio (SDR)-based jamming approach implemented on GNU Radio that generates a controlled number of random symbols, independent of the standard LoRa frame structure. This approach enables precise control over jammer exposure time and provides flexibility in studying the effect of jamming symbols on network performance. The theoretical analysis is validated through experimental results, where the implemented jammer is used to assess the impact of jamming under various configurations. Our findings demonstrate that LoRa-based networks can be disrupted with a minimal number of symbols, emphasizing the need for future research on stealthy communication techniques to counter such jamming attacks.
2025. 2025 IEEE 36th International Symposium on Personal, Indoor and Mobile Radio Communications (PIMRC), Istanbul, Turkey, 2025-09-01 – 2025-09-04. p. 1 – 6. DOI : 10.1109/pimrc62392.2025.11275285.Efficient Ordered Statistics Decoder with Unfolded Gaussian Elimination
Ordered statistics decoding (OSD) is a universal decoding algorithm for short-length linear block codes that approaches maximum likelihood decoding when run at sufficiently high order. However, the high computational complexity and latency of Gaussian elimination make it challenging for OSD implementations to meet stringent latency constraints. In this paper, we design a hardware decoder for OSD with order one, featuring an unfolded Gaussian elimination architecture with configurable parallelism. The proposed architecture can efficiently process multiple matrix columns and guarantees a low decoding latency. Based on 65 nm CMOS technology, the synthesis results show a worst-case throughput of 958 Mbps for a (128, 105) polar code at a frequency of 575 MHz.
2025. 2025 13th International Symposium on Topics in Coding, Los Angeles, CA, USA, 2025-08-18 – 2025-08-22. DOI : 10.1109/istc65386.2025.11154634.Benchmark of EEG-based seizure detection algorithms with SzCORE
EEG monitoring needs reliable automated seizure detection solutions to aid in the diagnosis and treatment of epilepsy. Many reviews have attempted to catalog and summarize key studies to provide a clear overview of current advances in algorithm development. However, a recurring issue in these reviews is that the field suffers from a lack of standardization in evaluation methodologies and performance metrics, making direct comparisons between algorithms extremely difficult. This work addresses this challenge by providing a fair & transparent comparison review of state-of-the-art seizure detection algorithms in the literature, using a standardized framework: SzCORE (Seizure Community Open-source Research Evaluation). We reviewed the existing literature on patient-independent EEG-based seizure detection algorithms trained on publicly available datasets. We re-implemented some of these algorithms and evaluated them with SzCORE. We found 19 papers that matched our selection criteria. We re-implemented three of them and found notable discrepancies between reported performances and those obtained under standardized evaluation conditions, highlighting the importance of transparent benchmarking. We observed that while algorithms tended to demonstrate high sensitivity (over 90%) in detecting seizure events, they generally exhibited low precision (10–40%), revealing a persistent issue with false-positive rates. We also found a high variability of the computed performance based on evaluation datasets, which is probably partially explained by the hourly rate of seizures. This work shows the value of a standardized evaluation methodology for EEG-based seizure detection and highlights the need for continued algorithm improvements.Clinical relevance— This study highlights the importance of standardized epileptic seizure detection algorithm evaluation. It shows that current state-of-the-art algorithms obtain a high sensitivity (∼70%) at the cost of a low precision (∼14%).
2025. 2025 47th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC), Copenhagen, Denmark, 2025-07-14 – 2025-07-18. p. 1 – 7. DOI : 10.1109/embc58623.2025.11254038.An Ultra Low Power Analog/Mixed-Signal Processor for a Smart RF Signal Classification System in the ISM Band
In this work, we propose an ultra-low power highly configurable analog/mixed-signal (AMS) processor as part of a smart RF signal detection and recognition system focusing on the ISM band. The analog feature extraction approach reduces power consumption and data rate through smart feature selection. The implemented highly configurable analog circuit enables detection and recognition of various signal types across a large variety of spectrum conditions. The proposed system achieves state-of-the-art accuracy with sub-mW power consumption.
2025. 2025 IEEE Radio Frequency Integrated Circuits Symposium (RFIC), San Francisco, CA, USA, 2025-06-15 – 2025-06-17. p. 259 – 262. DOI : 10.1109/rfic61188.2025.11082785.An SDR-Based Monostatic Wi-Fi System with Analog Self-Interference Cancellation for Sensing
Wireless sensing offers an alternative to wearables for contactless monitoring of human activity and vital signs. However, most existing systems use bistatic setups, which suffer from phase imperfections due to unsynchronized clocks. Monostatic systems overcome this issue, but are hindered by strong self-interference (SI) that requires effective cancellation. We present a monostatic Wi-Fi sensing system that uses an auxiliary transmit RF chain to achieve SI cancellation levels of 40 dB, comparable to existing solutions with custom cancellation hardware. We demonstrate that the cancellation filter weights, fine-tuned using least mean squares, can be directly repurposed for target sensing. Moreover, we achieve stable SI cancellation over 30 minutes in an office environment without fine-tuning, enabling traditional vital sign monitoring using channel estimates derived from baseband samples without the adaptation of the cancellation affecting the sensing channel – a significant limitation in prior work. Experimental results confirm the detection of small, slow-moving targets, representative for breathing chest movements, at distances up to 10 meters in non-line-of-sight conditions.
2025. 2025 IEEE International Symposium on Circuits and Systems (ISCAS), London, United Kingdom, 2025-05-25 – 2025-05-28. p. 1 – 5. DOI : 10.1109/iscas56072.2025.11043800.Belief Propagation Decoding for Short Codes on Structured Sparse Parity-Check Matrices
As successfully adopted in standard long code scenarios, belief propagation (BP) decoding has been considered a promising universal decoding candidate for next-generation wireless communications. However, when applied to short codes, BP decoding suffers from poor error correction performance due to harmful cycle structures in the Tanner graph. In this paper, we address this issue by designing a structured, sparse parity-check matrix (ssPCM) framework, composed of multiple cycle-free parity-check row blocks (PCRBs). The resulting ssPCMs feature regular row weights and perform better than the state-of-theart 4 -cycle-free row redundant PCMs across Bose-Chaudhuri-Hocquenghem (BCH) codes of length 63.
2025. 2025 IEEE International Symposium on Information Theory, Ann Arbor, MI, USA, 2025-06-22 – 2025-06-27. DOI : 10.1109/ISIT63088.2025.11195685.Enhancing Channel Decoding Through Full-Custom Hardware Design
Forward error correction (FEC) codes are essential components of both wired and wireless communication systems. High-throughput decoding is necessary to meet the demands of modern communication links operating at multiple gigabits per second. However, these advanced algorithms are often complex and costly to implement in terms of power and area, which pushes the limits of the most advanced technology nodes. Traditionally, these demanding signal processing tasks are realized using semi-custom, synchronous design principles with hardware description languages (HDLs), synthesized through electronic design automation (EDA) tools. This conservative approach often cannot achieve the highest performance or the lowest power consumption. To address this challenge, we propose to use alternative circuit and design paradigms based on full-custom and mixed-signal techniques for critical blocks. We present sorters and parallel syndrome checks (SCs) as examples of such blocks that are used frequently in FEC.
2025. 59 Asilomar Conference on Signals, Systems and Computers, Pacific Grove, United States, 2025-10-26 – 2025-10-29. p. 1750 – 1754. DOI : 10.1109/IEEECONF67917.2025.11443674.Iterative Logistic Weight Based Chase Decoder for Open Forward Error Correction
We propose an iterative logistic weight based decoder for open forward error correction (oFEC) codes. Compared to Chase-Pyndiah decoding with 93 error patterns, our decoder achieves similar performance with lower complexity.
2025. Optical Fiber Communication Conference (OFC) 2025, San Francisco, United States, 2025-03-30 – 2025-04-03. DOI : 10.1364/OFC.2025.Tu2F.6.Dataset and UAV Propagation Channel Modeling for LoRa in the 860 MHz ISM Band
LoRa is one of the most widely used low-power wide- area network technology for the Internet of Things. To achieve long-range communication with low power consumption at a low cost, LoRa uses a chirp spread spectrum modulation and transmits in the sub-GHz unlicensed industrial, scientific, and medical (ISM) frequency bands. Due to the rapid densification of IoT networks, it is crucial to obtain tailored channel models to evaluate the performance of LoRa networks. While channel models for cellular technologies have been investigated extensively, specific characteristics of LoRa transmissions operating at long range with a rather small (~ 250 kHz) bandwidth require dedicated measurement campaigns and modeling efforts. In this work, we leverage an SDR-based testbed to gather and publish a dataset of LoRa frames transmitted in a campus environment. The dataset includes IQ samples of the received frames at multiple locations and allows for the evaluation of channel variations with high time resolution. Using the gathered data, we derive empirical propagation channel models for LoRa that include receiver correlation over distance for three scenarios: unmanned aerial vehicle (UAV) line-of-sight (LoS), UAV non-LoS, and pedestrian non-LoS. Furthermore, the dataset is annotated with synchronization information, enabling the evaluation of receiver algorithms using experimental data.
2025. 59th Asilomar Conference on Signals, Systems and Computers, Pacific Grove, United States, 2025-10-26 – 2025-10-29. p. 841 – 845. DOI : 10.1109/IEEECONF67917.2025.11443521.Theses
Vital Signs and Human Activity Sensing with Wireless Systems
Wireless communication systems constantly estimate and equalize the channel to ensure reliable information transmission. The channel estimation not only facilitates effective communication, but also provides an opportunity to infer the physical environment. By analyzing changes in the channel estimate, it becomes possible to uncover corresponding physical changes in the surrounding environment. This link between the channel variations and environmental changes offers opportunities for an entirely new class of applications: wireless sensing. Wireless sensing leverages the interaction between pervasive wireless signals, such as WiFi, ultra-wideband (UWB), Bluetooth, or cellular signals, and the environment to extract meaningful information, such as motion, presence, gestures, or even vital signs, without requiring physical contact. This thesis explores various wireless sensing applications, including single- and multi-target breathing rate estimation, device positioning, device-free movement tracking, gesture recognition, and human presence detection, using omnipresent communication signal sources such as WiFi, UWB, and 4G long-term evolution (4G-LTE). While wireless sensing has made significant progress over the past decade, many challenges and unresolved issues persist. One fundamental question that appears across the board is whether the frequency domain or delay domain representation of the channel estimate is more suitable for wireless sensing applications. Additionally, the lack of synchronization among sensing devices introduces distortions in the channel estimates, leading to detection errors such as misinterpreting timing offsets and phase shifts as variations in the channel caused by the sensing target. To simplify and improve the efficiency of sensing models, it is crucial to preprocess channel estimates to extract relevant features. However, given various sensing applications, the question of which features are optimal for specific sensing tasks remains largely unaddressed. In this thesis, we explore these fundamental yet underexamined questions. First, we compare the frequency domain channel frequency response (CFR) and the delay domain channel impulse response (CIR), analyzing their suitability for different sensing tasks. We also propose calibration methods tailored to different channel estimate representations and communication technologies, as well as dimension reduction techniques to enhance sensing performance by selecting relevant components or combining multiple components of channel estimates. To address the complexity of physical propagation channels, we develop learning-based approaches to map channel estimates to sensing outcomes. Additionally, we tackle the challenges of feature extraction and selection, emphasizing the importance of leveraging prior knowledge of the physical channel to pre-select features rather than directly inputting raw channel measurements into learning models. Through these contributions, this thesis advances the theoretical understanding of wireless sensing and demonstrates innovative applications that have the potential to lead to cost-effective, convenient, and reliable wireless sensing systems. These advancements contribute to the development of technologies such as smart cities, smart buildings, and enhanced human-machine interactions.
Lausanne, EPFL, 2025.Design, Optimization and Verification of Embedded Gain Cell RAMs
One of the vital components of Very Large Scale Integration (VLSI) is embedded memories. Embedded memories are widely used in various system-on-chip (SoC) architectures, including communication systems, consumer electronics, and the automotive applications. In fact, the demand for the high-performance, low power, and small area embedded memorieshas even increased dramatically with the recent advent of Artificial Intelligence. The most commonly used embedded memory in SoC devices is the six-transistor static random-access memory (6T-SRAM). However, 6T-SRAM structures are limited in terms of power and area efficiency. Each bit requires six transistors for storage, and the inherent latch structure consumes significant leakage power. Additionally, 6T-SRAM requires specific sizing ratios to ensure reliable write operations due to its latch-based architecture. These sizing constraints make 6T-SRAM highly sensitive to process variations, further limiting its scalability. An alternative to SRAM is the embedded gain-cell (GC) RAM, which generally requires only two to four transistors per bit, offering a better area density than 6T-SRAM. GC-RAMs have no direct path to the supply, resulting in lower leakage currents. Although GC-RAM can achieve speeds comparable to SRAM, it suffers from limited data retention times. Systematic refresh operations can mitigate retention time issues in mature process nodes. However, advanced process nodes face even lower retention times, necessitating frequent refresh cycles. These frequent refresh operations increase power consumption and reduce memory availability, presenting a significant challenge for the practical application of GC-RAM arrays. Further research is crucial to enhance data retention times in advanced process nodes, enabling the full utilization of GC-RAM benefits. This thesis provides a detailed exploration of the GC-RAM memory array design process, from the bitcell level to the macro level, along with innovative assist techniques implemented to maximize data retention time. These memory arrays are optimized to achieve maximized data retention time and high density, without compromising power efficiency or frequency. In this thesis, three memory arrays are designed and fabricated by using proposed verification method across three different process nodes. The process nodes are 65 nm bulk Complementary Metal-Oxide-Semiconductor (CMOS), 28 nm Fully Depleted Silicon On Insulator (FD-SOI) and 16 nm Fin Field Effect Transistor (FINFET). In each process node, the designed arrays outperform state-of-the-art GC-RAMs in the same technology node and demonstrate substantial area and power advantage over the 6T-SRAM.
Lausanne, EPFL, 2025.Efficient Algorithms and VLSI Implementations of FEC Decoders for B5G/6G
Channel coding has become an indispensable component of modern communication systems. By introducing redundancy into the transmitted data, it enables the receiver to detect and correct errors without feedback, thereby significantly improving the efficiency of transceivers. As wireless standards continue to evolve, particularly toward 6G, channel coding faces ever greater demands on performance, flexibility, and implementation efficiency. However, the increasing complexity of advanced decoders, combined with the slowdown of Moore’s law, has made it challenging to bridge the gap between coding theory and hardware implementation. Given that polar codes and low-density parity-check (LDPC) codes have been ratified as the 5G New Radio (NR) standard codes, the design of efficient decoders for these two modern codes has become a central bottleneck for 5G-NR and beyond. In addition to support for varying block sizes, code rates, and code structures, 5G-NR imposes strict demands on polar codes and LDPC codes in terms of reliability, latency, and throughput. These challenges call for a cross-layer design perspective to tightly integrate algorithmic efficiency with architectural scalability. As a newly adopted class of codes in 5G, polar codes face challenges in low-latency and high-reliability of decoders. While node-based successive cancellation list (SCL) decoding has emerged as an effective approach to reduce the latency of SCL decoding, current node-based decoders are often constrained by their limited ability to generalize across diverse node types. This restricts their flexibility and leaves room for improvement in decoding speed. In this thesis, we propose a generalized node-based SCL decoder to minimize latency in single-frame polar decoding. Moreover, we introduce a frame-interleaving architecture to explore the throughput potential of polar decoders. For LDPC codes, the primary challenge in 5G-NR is to support a wide range of code configurations while meeting the required peak throughput. To this end, we develop a fully reconfigurable 5G-NR LDPC decoder architecture to fulfill these requirements. In summary, we present algorithmic and architectural optimizations for 5G-NR codes, along with two decoder implementations that are currently among the most efficient designs meeting the standard requirements. Meanwhile, as 5G is making inroads to commercial devices, the global research community is actively exploring candidates for 6G channel coding. In contrast to 5G-NR, where code constructions are already fixed, the transition to 6G offers an opportunity to incorporate code design itself into the innovation process, prompting a revisit of coding schemes, decoding algorithms, and hardware architectures. In this thesis, we propose a new spatially-coupled LDPC code family called edge-spreading Raptor-like (ESRL) LDPC codes as a candidate for 6G next-generation mobile broadband. While preserving key features of 5G-NR standard codes, the proposed ESRL codes have advantages in error-rate performance, throughput, and hardware complexity compared to 5G-NR LDPC codes. To effectively realize the theoretical advantages in practice, we further present a fully reconfigurable high-throughput LDPC decoder implementation for ESRL codes. This ASIC can support a wide range of code rates and code lengths (up to five times longer than 5G) and achieve a high peak throughput of more than 100 Gbps, making it a promising solution for 6G wireless systems.
Lausanne, EPFL, 2025.Monostatic Wi-Fi Sensing: Self-Interference Cancellation and Channel Splicing for Robust Vital-Sign Monitoring
Continuous vital-sign monitoring enables early detection of abnormal physiological patterns, which is crucial for preventive healthcare. However, conventional monitoring devices face critical limitations: wearables have compliance issues, cameras raise privacy concerns, and clinical devices disrupt patient rest. This thesis addresses these constraints by developing robust Wi-Fi sensing infrastructure that provides a stable signal acquisition foundation upon which future diagnostic applications can be developed. Wi-Fi systems continuously estimate the wireless channel for communication. This channel estimate inherently captures motion-induced variations through human reflections, enabling detection of physiological motions like respiration. However, detecting such movements is challenging due to environmental noise and interference like walking. Furthermore, commodity Wi-Fi devices limit access to raw in-phase and quadrature (IQ) samples, which restricts channel estimation to just a fraction of the Wi-Fi frame, the training field. Moreover, standard Wi-Fi operates in bistatic configurations with spatially separated transmitter (TX) and receiver (RX), causing phase errors from unsynchronized clocks. Additionally, limited bandwidth yields coarse range resolution of several meters, which is insufficient to separate vital-sign reflections from nearby interference. These limitations motivate monostatic architectures, where co-located TX and RX chains share a common reference clock, eliminating synchronization errors while enabling easier raw IQ access. Such architectures support in-band full-duplex operation, enabling simultaneous transmission and reception. However, strong self-interference (SI) from TX-RX coupling overpowers weak target reflections. While SI cancellation (SIC) can mitigate SI, conventional analog SIC (AnSIC) systems that generate cancellation signals directly in the analog domain (A-AnSIC) require frequent tuning that inadvertently cancels physiological movements. This thesis enables stable Wi-Fi-based vital sign monitoring through architectural and signal processing innovations. First, we introduce novel digital SI cancellation (DiSIC) techniques leveraging neural networks, which effectively capture RF non-linearities to enable high-quality SIC with lower computational complexity compared to polynomial models. Then, we develop an initial monostatic sensing system, which demonstrates accurate multi-target breathing rate estimation. To provide analog SIC, we develop an AnSIC system that digitally generates the cancellation signal (D-AnSIC), achieving stable 40 dB cancellation without continuous tuning. To further remove the SI, we introduce our SI distillation and removal (SIDAR) method, which combines phase calibration with low-rank decomposition to isolate SI from weak target reflections. SIDAR enables reliable 10-meter non-line-of-sight monitoring even without D-AnSIC, though combining both yields optimal performance. Finally, we develop a channel splicing method that coherently merges multiple 20 MHz channels for 125 MHz effective bandwidth. The enhanced resolution enables precise chest-displacement estimation and achieves median respiratory rate errors below 0.5 bpm even with interference from nearby walking. These contributions address fundamental limitations in current Wi-Fi sensing systems, SI cancellation, and signal processing for future healthcare monitoring systems.
Lausanne, EPFL, 2025.Patents
Dynamic gain cell with reduced leakage
A 3T gain cell includes a write transistor, a storage transistor and a read transistor. The write transistor has a first diffusion connected to a write bit line (WBL), a gate connected to a write word line (WWL) and a second diffusion. The storage transistor has a gate connected to said second diffusion of said write transistor, a first diffusion connected to a read bit line (RBL), and a second diffusion. The read transistor has a gate connected to a read word line (RWL) a first diffusion connected to said second diffusion of said storage transistor, and a second diffusion connected to a reference voltage.
US2025285672.
2025.2024
Journal Articles
Optimizing Polar Codes for Reduced Latency Successive Cancellation List Decoder
This paper presents the application of genetic algorithms (GenAlg) to optimize polar codes for reducing the decoding latency of fast cyclic redundancy check aided successive cancellation list decoders. The theoretical decoding time steps (TS) in code construction are assessed under the constraint of a maximum special node size. The GenAlg employed iteratively utilizes mutation and crossover to reduce decoding TS by generating more large special nodes while maintaining acceptable performance. By avoiding redundant simulations, the GenAlg iterations are reduced by 5 orders of magnitude. Synthesis results indicate that the proposed scheme can reduce latency by up to 20.2% compared to 5G codes.
IEEE Communications Letters. 2024. DOI : 10.1109/LCOMM.2024.3521080.Belief Propagation Decoding for Short-Length Codes Based on Sparse Tanner Graph
In this letter, we present a novel approach to improve belief propagation (BP) decoding performance by sparsifying the Tanner graph. Our proposed method first constructs a low-weight parity-check (LWPC) matrix with increased rows based on the original parity-check matrix. Then, all 4-cycles in the LWPC matrix are eliminated by adding auxiliary variable nodes and corresponding parity-check rows, resulting in the generalized LWPC (G-LWPC) matrix. Compared to the original matrix, the sparsity is greatly improved without altering the encoding constraints. Simulation results show that BP decoding based on the G-LWPC matrix is particularly effective for short code lengths. For a (128, 20) polar code with 24-bit cyclic redundancy check, our proposed G-LWPC matrix reduces the density from 17.4% to 1.3%, and with 10 iterations, BP decoding can achieve the performance of successive cancellation list decoding with a list size of 32.
IEEE COMMUNICATIONS LETTERS. 2024. Vol. 28, num. 5, p. 969 – 973. DOI : 10.1109/LCOMM.2024.3366293.A High Dynamic Range Envelop Detector for Heterodyne Receiver Architecture
Energy detection (ED) has been of interest for various applications ranging from very low frequency biomedical signal read-out to high frequency wireless communication. The analog implementation of the ED requires circuits with a non-linear transfer characteristics which can be obtained through squaring or rectification. Comparison of these two methods is presented and implementation limitations are detailed. Interesting characteristics of rectifiers, derived from ideal transfer characteristics and from non-ideal circuit implementations, encourage rectifier’s use in the ED rather than the squarer. A high speed and high precision current mode full-wave rectifier was implemented in GF 22FDX technology. Our circuit achieves a 60 dB dynamic range and 100 MHz bandwidth with an ultra-low power consumption of 4.3 mu W from 0.8V voltage supply.
IEEE Transactions on Circuits and Systems II: Express Briefs. 2024. Vol. 71, num. 4, p. 1929 – 1933. DOI : 10.1109/TCSII.2023.3329038.A Generalized Adjusted Min-Sum Decoder for 5G LDPC Codes: Algorithm and Implementation
5G New Radio (NR) has stringent demands on both performance and complexity for the design of low-density parity-check (LDPC) decoding algorithms and corresponding VLSI implementations. Furthermore, decoders must fully support the wide range of all 5G NR blocklengths and code rates, which is a significant challenge. In this paper, we present a high-performance and low-complexity LDPC decoder, tailor-made to fulfill the 5G requirements. First, to close the gap between belief propagation (BP) decoding and its approximations in hardware, we propose an extension of adjusted min-sum decoding, called generalized adjusted min-sum (GA-MS) decoding. This decoding algorithm flexibly truncates the incoming messages at the check node level and carefully approximates the non-linear functions of BP decoding to balance the error-rate and hardware complexity. Numerical results demonstrate that the proposed fixed-point GA-MS has only a minor gap of 0.1 dB compared to floating-point BP under various scenarios of 5G standard specifications. Secondly, we present a fully reconfigurable 5G NR LDPC decoder implementation based on GA-MS decoding. Given that memory occupies a substantial portion of the decoder area, we adopt multiple data compression and approximation techniques to reduce 42.2% of the memory overhead. The corresponding 28nm FD-SOI ASIC decoder has a core area of 1.823 mm(2) and operates at 895 MHz. It is compatible with all 5G NR LDPC codes and achieves a peak throughput of 24.42 Gbps and a maximum area efficiency of 13.40 Gbps/mm(2) at 4 decoding iterations.
Ieee Transactions On Circuits And Systems I-Regular Papers. 2024. DOI : 10.1109/TCSI.2024.3368056.Multi-Ported GC-eDRAM Bitcell with Dynamic Port Configuration and Refresh Mechanism
Embedded memories occupy an increasingly dominant part of the area and power budgets of modern systems-on-chips (SoCs). Multi-ported embedded memories, commonly used by media SoCs and graphical processing units, occupy even more area and consume higher power due to larger memory bitcells. Gain-cell eDRAM is a high-density alternative for multi-ported operation with a small silicon footprint. However, conventional gain-cell memories have limited data availability, as they require periodic refresh operations to maintain their data. In this paper, we propose a novel multi-ported gain-cell design, which provides up-to N read ports and M independent write ports (NRMW). In addition, the proposed design features a configurable mode of operation, supporting a hidden refresh mechanism for improved memory availability, as well as a novel opportunistic refresh port approach. An 8kbit memory macro was implemented using a four-transistor bitcell with four ports (2R2W) in a 28 nm FD-SOI technology, offering up-to a 3x reduction in bitcell area compared to other dual-ported SRAM memory options, while also providing 100% memory availability, as opposed to conventional dynamic memories, which are hindered by limited availability.
Journal Of Low Power Electronics And Applications. 2024. Vol. 14, num. 1, p. 2. DOI : 10.3390/jlpea14010002.A Node-Based Polar List Decoder With Frame Interleaving and Ensemble Decoding Support
Node-based successive cancellation list (SCL) decoding has received considerable attention in wireless communications for its significant reduction in decoding latency, particularly with 5G New Radio (NR) polar codes. However, the existing node-based SCL decoders are constrained by sequential processing, leading to complicated and data-dependent computational units that introduce unavoidable stalls, reducing hardware efficiency. In this paper, we present a frame-interleaving hardware architecture for a generalized node-based SCL decoder. By efficiently reusing otherwise idle computational units, two independent frames can be decoded simultaneously, resulting in a significant throughput gain. Based on this new architecture, we further exploit graph ensembles to diversify the decoding space, thus enhancing the error-correcting performance with a limited list size. Two dynamic strategies are proposed to eliminate the residual stalls in the decoding schedule, which eventually results in nearly 2× throughput compared to the state-of-the-art baseline node-based SCL decoder. To impart the decoder rate flexibility, we develop a novel online instruction generator to identify the generalized nodes and produce instructions on-the-fly. The corresponding 28nm FD-SOI ASIC SCL decoder with a list size of 8 has a core area of 1.28 mm2 and operates at 692 MHz. It is compatible with all 5G NR polar codes and achieves a throughput of 3.34 Gbps and an area efficiency of 2.62 Gbps/mm2 for uplink (1024, 512) codes, which is 1.41× and 1.69× better than the state-of-the-art node-based SCL decoders.
IEEE Transactions on Circuits and Systems I: Regular Papers. 2024. Vol. 71, num. 12, p. 5457 – 5470. DOI : 10.1109/TCSI.2024.3443598.A 16-kB 65-nm GC-eDRAM Macro With Internal Bias Voltage Generation Providing Over 100-μs Retention Time
Gain-cell embedded dynamic random access memory (GC-eDRAM) has emerged as a suitable choice for embedded memory implementation due to its high density, low leakage current, and voltage scaling compatibility. This work presents a 16-kB 3T-1C GC-eDRAM macro, featuring an innovative internal reference voltage generation mechanism and an on-chip dc-dc converter for internal boosted supply generation. The memory architecture is partitioned to efficiently accommodate the reference generation and implement a variation-tolerant sensing scheme. The on-chip dc-dc converter is employed for internally generating a boosted voltage that enhances charge retention to increase the data retention time (DRT). The memory macro was implemented in a 65-nm CMOS technology and fabricated as part of a research test chip. Measurements across a spectrum of boosted voltages and different temperature points, show a significant improvement in DRT compared with similar GC-eDRAM designs, without compromising area, performance, or power dissipation.
IEEE Journal of Solid-State Circuits. 2024. DOI : 10.1109/JSSC.2024.3489793.High-Throughput and Flexible Belief Propagation List Decoder for Polar Codes
Due to its high parallelism, belief propagation (BP)decoding is amenable to high-throughput applications and thusrepresents a promising solution for the ultra-high peak datarate required by future communication systems. To bridge theperformance gap compared to the widely used successive cancel-lation list (SCL) decoding algorithm, BP list (BPL) decoding forpolar codes extends candidate codeword exploration via multiplepermuted factor graphs (PFGs) to improve the error-correctingperformance of BP decoding. However, it is a significant challengeto design a unified and flexible BPL hardware architecture thatsupports various PFGs and code configurations. In this paper,we present the first VLSI implementation of a BPL decoder forpolar codes that overcomes this implementation challenge with ahardware-friendly algorithm for on-the-fly flexible permutations.First, we introduce a sequential generation (SG) algorithm toobtain a near-optimal PFG set. Additionally, we demonstratethat any permutation can be decomposed into a combination ofmultiple fixed routings, and design a low-complexity permutationnetwork to generate graphs in an on-the-fly fashion. Our BPLdecoder has a low decoding latency by executing decoding andpermutation generation in parallel and supports arbitrary listsizes without area overhead. Experimental results based on 28nmFD-SOI technology show that for length-1024 polar codes witha code rate of one-half, our BPL decoder with 32 PFGs exhibitssimilar error-correcting performance to SCL with a list size of 4and achieves an average throughput of 25.63 Gbps and an areaefficiency of 29.46 Gbps/mm2, which is 1.82xand 4.33xfasterthan the state-of-the-art BP flip and SCL decoders, respectively
Ieee Transactions On Signal Processing. 2024. Vol. 72, p. 1158 – 1174. DOI : 10.1109/TSP.2024.3361073.Conference Papers
Centralized RAN for LPWAN: Architecture and Proof-of-Concept Prototype Implementation
Centralized radio access networks (C-RAN) are widely recognized as an enabling technology for the evolution of cellular systems, including 5G and beyond. However, the concept is not limited to 3GPP cellular systems and licensed frequency bands. In recent years, unlicensed low-power wide-area networks (LPWANs) have become popular for a variety of Internet of Things (IoT) applications. Those LPWANs can also leverage the concepts that popularized the C-RAN for cellular systems, such as coordinated multi points (CoMP), cell-free MIMO, and resources virtualization. In this paper, we first identify the requirements of popular unlicensed LPWAN technologies and detail the components of a C-RAN architecture that meets those requirements. We then present a proof-of-concept prototype implementation of the proposed architecture that supports the LoRa modulation, one of the most popular LPWAN technologies, and evaluate the latency and processing load of the system.
2024. 2024 IEEE Conference on Network Function Virtualization and Software Defined Networks, Natal, Brazil, 2024-11-05 – 2024-11-07. DOI : 10.1109/NFV-SDN61811.2024.10807469.Towards 6G: Configurable High-Throughput Decoder Implementation for SC-LDPC Codes
Spatially-coupled low-density parity-check (SC-LDPC) codes are considered an important candidate for the upcoming 6G standardization, owing to their superior error-correction capability and practical decoding complexity. However, current high-throughput SC-LDPC decoders rely on unrolled architectures tailored to a certain code and thus lack the necessary rate compatibility for wireless communications. In this paper, we present a fully reconfigurable decoder architecture for SC-LDPC codes. This hardware architecture can decode virtually any SC-LDPC code that fits in the allocated memories. Based on simplified routing networks, our SC-LDPC decoder can deliver extensive code length and code rate support and achieve high throughput to satisfy the requirements of 6G.
2024. 2024 58th Asilomar Conference on Signals, Systems, and Computers, Pacific Grove, CA, USA, 2024-10-27 – 2024-10-30. p. 980 – 984. DOI : 10.1109/ieeeconf60004.2024.10942947.LoRa Fine Synchronization with Two-Pass Time and Frequency Offset Estimation
LoRa is currently one of the most widely used low-power wide-area network (LPWAN) technologies. The physical layer leverages a chirp spread spectrum modulation to achieve long-range communication with low power consumption. Synchronization at long distances is a challenging task as the spread signal can lie multiple orders of magnitude below the thermal noise floor. Multiple research works have proposed synchronization algorithms for LoRa under different hardware impairments. However, the impact of sampling frequency offset (SFO) has mostly either been ignored or tracked only during the data phase, but it often harms synchronization. In this work, we extend existing synchronization algorithms for LoRa to estimate and compensate SFO already in the preamble and show that this early compensation has a critical impact on the estimation of other impairments such as carrier frequency offset and sampling time offset. Therefore it is critical to recover long-range signals.
2024. 2024 58th Asilomar Conference on Signals, Systems, and Computers, Pacific Grove, CA, USA, 2024-10-27 – 2024-10-30. p. 1802 – 1806. DOI : 10.1109/ieeeconf60004.2024.10942846.Learning-based Hand Gesture Classification using Channel Impulse Response with UWB
The channel impulse response (CIR) of the wireless propagation channel is influenced by the surrounding environment and thus can be used to retrieve environmental information such as location, the presence of objects, and the speed of objects. In this work, we detect hand gestures based on the complex-valued CIR from an ultra-wideband (UWB) transmission link between one transmitter and two receivers. Thanks to the high path delay resolution due to the wide (500 MHz) bandwidth, we can focus on the channel that is influenced by hand gestures in the sensing area. Using different machine learning and deep learning methods, we learn the features from sequences of CIR snapshots that are correlated to hand gestures and recognize them.
2024. 35 IEEE International Symposium on Personal, Indoor and Mobile Radio Communications, Valencia, Spain, 2024-09-02 – 2024-09-05. DOI : 10.1109/PIMRC59610.2024.10817379.A GRANDAB Decoder with 8.48 Gbps Worst-Case Throughput in 65nm CMOS
We present an error correction code (ECC) decoder based on the Guessing Random Additive Noise Decoding (GRAND) with ABandonment (GRANDAB) algorithm for applications that require constant throughput and fixed decoding latency in the Gbps range. This high constant throughput distinguishes our work from other GRAND decoders with similar average throughput, but with orders of magnitude longer worstcase latency. By leveraging the regularity, simplicity, and inherent parallelism of GRANDAB, our design maintains competitive area- and energy-efficiency. In 65 nm CMOS, we demonstrate our appoach with the BCH-(127,106) code. The ASIC implementation results show that, with a core area of 6.15mm2, our decoder maintains a fixed (best-and worst-case) throughput of 8.48 Gbps with an energy consumption of 17.5pJ/bit at 1.25 V and 3.71 Gbps with 5.7pJ/bit at 0.75 V.
2024. 50th IEEE European Solid-State Electronics Research Conference, Bruges, Belgium, 2024-09-09 – 2024-09-12. p. 685 – 688. DOI : 10.1109/ESSERC62670.2024.10719587.Monostatic Multi-Target Wi-Fi-Based Breathing Rate Sensing Using Openwifi
Continuous monitoring of human activity and vital signs has become increasingly important in healthcare and consumer applications. While wearables and others sensors are already in widespread use, they often suffer from issues such as discomfort and the need for physical contact. To address these limitations, wireless sensing using radio signals has emerged as a promising alternative. However, most previous work has focused on bistatic setups utilizing commodity Wi-Fi devices. In this work, we use the open-source openwifi platform for monostatic multi-antenna Wi-Fi sensing for contactless breathing rate measurements. We also present a method for simultaneous estimation of the breathing rate of two targets. Experimental results using CNC machines for emulating human breathing demonstrate the effectiveness of our method and setup for targets at a few meters distance. Even in challenging scenarios where the breathing rates are similar, the targets are close together, and the targets have the same angle of arrival, we achieve an average error of 0.61 breaths per minute. We further validate our setup and method by collecting human breathing data from two subjects, which can clearly be distinguished using our setup and method. Moreover, our results indicate that self-interference cancellation is not necessary for close-range breathing rate sensing.
2024. 25 IEEE Wireless Communications and Networking Conference, Dubai, United Arab Emirates, 2024-04-21 – 2024-04-24. DOI : 10.1109/WCNC57260.2024.10570912.A Low-Latency and High-Performance SCL Decoder with Frame-Interleaving
In this paper, we describe a frame-interleaving hardware architecture for a generalized node-based successive cancellation list (SCL) decoder. By efficiently reusing otherwise idle computational units, two independent frames can be decoded simultaneously, resulting in a significant throughput gain. Based on this new architecture, we also exploit graph ensembles to diversify the decoding, enhancing the error-correcting performance by 0.28 dB and reducing the worst-case latency for serial graph processing by over 32%. Implementation results show that the proposed SCL decoder with frame-interleaving architecture achieves a throughput of 7.15 Gbps and an area efficiency of 37.63 Gbps/mm2, which is 1.56× and 1.11× better than the state-of-the-art node-based SCL decoders.
2024. IEEE International Symposium on Circuits and Systems, Singapore, Singapore, 2024-05-19 – 2024-05-22. DOI : 10.1109/ISCAS58744.2024.10558152.Theses
Intelligent RF System for Ultra Low Power Spectrum Sensing with Machine Learning
The crowdedness of the RF spectrum constitutes problems in communication, however it also accommodates large amount of useful information. This information can be used to determine existing signals and operating contexts. However, spectrum condition varies profoundly and each condition possesses different processing challenges. Although an optimal power consumption is achievable for each given context, the literature lacks approaches for context-aware system adaptation and low-power hardware implementation. This thesis focuses on defining a generic method to extract information from the spectrum with the optimal power consumption under varying operating conditions. To demonstrate the generic method, a Wi-Fi, Bluetooth, and microwave oven signal detection system is proposed. Prior work showed that analog pre-processing is more power efficient than digital in low precision. Hence, the proposed system consists of an integrated highly configurable analog feature extractor (AFE) followed by off-chip digital ML classifiers. First, signal processing solutions are investigated and novel system level architectures are proposed. The proposed context-aware and intelligent duty-cycled (DC) operation schemes are detailed. I also discuss various power saving schemes. Second, modeling of aforementioned target signals and context scenarios are discussed. Sparse and medium-crowded spectra are considered in this work. Highly-crowded spectrum requires very high dynamic range, and is out of the scope of this work. Novel and patented low-power features are proposed. A top-down design methodology was used in this work and design discussions start with system level design aspects. The proposed base-band analog signal processor utilizes complex valued processing and consists of a variable-gain amplifier, a 4-channel complex band-pass filter (CBPF) bank, transconductors (OTA), and AFE blocks. The proposed sensing radio was designed and fabricated in GF 22FDX technology. A charge-based differential-pair non-linearity model is also developed for low-power analog circuit optimization. Measurement results showed excellent agreement with the analytic model even for deep sub-micron devices. Then, building block design details and characterization results are presented. The AFE consumes only 5µW and the OTA consumes 7µW. The CBPF bank consumption ranges from 39µW to 138µW depending on the frequency setting. In sparse context setting, the system incorporates the VGA and a full-band AFE, and consumes 216µW, only 11% of which is consumed by the AFE. Power consumption of the overall system consisting of the VGA, the CBPF bank, and all the AFEs ranges from 353µW to 452µW in medium-crowded context. Reported power consumptions are from a 0.8V power supply in always-on operation. The proposed DC scheme reduces average power to 58µW for sparse and to the range from 96µW to 121µW for medium-crowded contexts with only 11µs worst case detection latency. The average power consumption can be reduced down to the nano-Watt range by the proposed DC scheme with further power-latency trade-off. Last, simulated and measured feature databases are compared. Several classifier algorithms were investigated by a collaborative team. Accuracies of up to 100% and 99.9% were achieved for signal existence and recognition tests, respectively. Demonstrated classification performance proves the feasibility of low-power RF spectrum sensing and reaches the state-of-the-art accuracy levels.
Lausanne, EPFL, 2024.2023
Journal Articles
Belief-Selective Propagation Detection for MIMO Systems
Compared to the linear MIMO detectors, the Belief Propagation (BP) detector has shown greater capabilities in achieving near-optimal performance and better nature to iteratively cooperate with channel decoders. Aiming at real applications, recent works mainly fall into the category of reducing the complexity by simplified calculations, at the expense of performance sacrifice. However, the complexity is still unsatisfactory with exponentially increasing complexity or required exponentiation operations. Furthermore, the state-of-the-art (SOA) BP detectors persistently encounter error floor in high signal-to-noise ratio (SNR) region, which becomes even worse with calculation approximation. This work aims at a revised BP detector, named Belief-selective Propagation (BsP) detector by selectively utilizing the trusted incoming messages with sufficiently large a priori probabilities for updates. Two proposed strategies: symbol-based truncation (ST) and edge-based simplification (ES) squeeze the complexity (orders lower than the BP detector), while greatly relieving the error floor issue over a wide range of antenna and modulation combinations. For the 256-QAM 128 chi 64 uplink massive multiuser MIMO (MU-MIMO) system, the B(1,1) BsP detector achieves more than 1dB performance gain (@BER=10(-4)) with lower complexity than the state-of-the-art (SOA) BP detector. Trade-off between performance and complexity towards different application requirements can be conveniently obtained by tuning the parameters of the ST and ES strategies.
Ieee Transactions On Communications. 2023. Vol. 71, num. 12, p. 7244 – 7257. DOI : 10.1109/TCOMM.2023.3305534.Pipelined Architecture for Soft-Decision Iterative Projection Aggregation Decoding for RM Codes
The recently proposed recursive projection-aggregation (RPA) decoding algorithm for Reed-Muller codes has received significant attention as it provides near-ML decoding performance at reasonable complexity for short codes. However, its complicated structure makes it unsuitable for hardware implementation. Iterative projection-aggregation (IPA) decoding is a modified version of RPA decoding that simplifies the hardware implementation. In this work, we present a flexible hardware architecture for the IPA decoder that can be configured from fully-sequential to fully-parallel, thus making it suitable for a wide range of applications with different constraints and resource budgets. Our simulation and implementation results show that the IPA decoder has 41% lower area consumption, 44% lower latency, four times higher throughput, but currently seven times higher power consumption for a code with block length of 128 and information length of 29 compared to a state-of-the-art polar successive cancellation list (SCL) decoder with comparable decoding performance.
Ieee Transactions On Circuits And Systems I-Regular Papers. 2023. DOI : 10.1109/TCSI.2023.3306463.An Ultra-Low-Power Widely-Tunable Complex Band-Pass Filter for RF Spectrum Sensing
Power consumption is of utmost importance in portable devices as they are operated from a limited energy supply. Wireless radio is one of the most power-hungry blocks in these systems and needs to be activated opportunistically. Radio frequency (RF) spectrum sensing can be performed with a dedicated low power radio to control the main wireless radio. A low power receiver architecture is proposed in this work to monitor the 2.4 GHz Industrial, Scientific, and Medical (ISM) band for communication standards such as Wireless Local Area Network (WLAN), Zonal Intercommunication Global-standard (ZigBee), and Bluetooth-Low-Energy (BLE). The whole ISM band is down-converted to zero-intermediate-frequency (ZIF) and a tunable base-band complex band-pass filter (CBPF) is used to scan the spectrum. A widely tunable first order ultralow-power transconductor-capacitor (Gm-C) CBPF architecture was designed and fabricated in GF 22FDX technology, which occupies 0.0049mm(2) area. A frequency shift of +/- 60MHz was achieved with 5 – 40MHz bandwidth range. Power consumption ranges from 6.7 mu W to 99.2 mu W with the best and the worst figure-of-merit of 0.034 fJ/pole and 0.082 fJ/pole that is the energy consumption per pole normalized by the spurious free dynamic range.
Ieee Transactions On Circuits And Systems I-Regular Papers. 2023. DOI : 10.1109/TCSI.2023.3300965.A 128-kbit GC-eDRAM With Negative Boosted Bootstrap Driver for 11.3x Lower-Refresh Frequency at a 2.5% Area Overhead in 28-nm FD-SOI
Gain-cell embedded DRAM (GC-eDRAM) is a high-density logic-compatible alternative to conventional static random-access memory (SRAM) and embedded DRAM (eDRAM). However, GC-eDRAM suffers from a reduced data retention time (DRT) at deeply-scaled process nodes, leading to frequent power-hungry refresh operations. In order to reduce the refresh overhead, GC-eDRAM macros utilize external assist voltages which improve the bitcell write-ability, leading to an enhanced DRT. However, the requirement for external analog supply voltages creates additional overhead and is often impractical in the design of compact systems-on-chip (SoC). This work presents an on-chip write-assist technique implemented with a negative boosted bootstrap driver which generates the required wordline boosting on-chip without external components. The proposed circuitry is integrated compactly inside the GC-eDRAM macro to provide an area-efficient low-power solution which improves the bitcell’s write-ability and reduces its refresh requirement. A 128-kbit GC-eDRAM macro utilizing the proposed boosting circuitry has been fabricated in a 28-nm FD-SOI technology, demonstrating an 11.3(X) DRT improvement at only 2.5% area overhead.
Ieee Solid-State Circuits Letters. 2023. Vol. 6, p. 13 – 16. DOI : 10.1109/LSSC.2022.3232775.Conference Papers
Increasing LoRa Sensitivity and Reliability with an IoT Cloud RAN
LoRaWAN is one of the most popular low-power wide-area network (LPWAN) technologies. Its network architecture is based on a star-of-stars topology, in which LoRa nodes communicate with gateways that forward payloads to a central server when a packet detection occurs. As the gateway deployment is often not managed by a single entity, their locations across the network can be highly inhomogeneous. To improve the coverage as well as the reliability of LoRaWAN, we propose to use a centralized radio access network (C-RAN) architecture. We introduce an implementation of such an architecture suited for LPWANs which supports popular software-defined radios through interfaces that are compatible with the GNU Radio toolkit. We also show that a centralized processing of baseband signals received by multiple gateways can offer more resilience against fading and allow the utilization of higher datarates for the uplink. Both of these effects reduce the power consumption of LoRa devices which is critical for battery-powered edge nodes.
2023. 57 Asilomar Conference on Signals, Systems and Computers, Pacific Grove, United States, 2023-10-29 – 2023-11-01. p. 197 – 201. DOI : 10.1109/IEEECONF59524.2023.10476850.Band-of-Interest-based Channel Impulse Response Fusion for Breathing Rate Estimation with UWB
The channel impulse response (CIR) obtained from the channel estimation step of various wireless systems is a widely used source of information in wireless sensing. Breathing rate is one of the important vital signs that can be retrieved from the CIR. Recently, there have been various works that extract the breathing rate from one carefully selected CIR delay bin that contains the breathing information. However, it has also been shown that the accuracy of this estimation is very sensitive to the measurement scenario, e.g., if there is any obstacle between the transceivers and the target, the position of the target, and the orientation of the target, since only one CIR delay bin does not contain a sufficient periodic component to retrieve the breathing rate. We focus on such scenarios and propose a CIR delay bin fusion method to merge several CIR bins to achieve a more accurate and reliable breathing rate estimate. We take measurements and showcase the advantages of the proposed method across scenarios.
2023. IEEE International Conference on Communications (IEEE ICC), Rome, ITALY, MAY 28-JUN 01, 2023. p. 1695 – 1700. DOI : 10.1109/ICCWORKSHOPS57953.2023.10283555.Single-anchor Uwb Localization Using Channel Impulse Response Distributions
Ultra-wideband (UWB) devices are widely used in indoor localization scenarios. Single-anchor UWB localization shows advantages because of its simple system setup compared to conventional two-way ranging (TWR) and trilateration localization methods. In this work, we focus on single-anchor UWB localization methods that learn statistical features of the channel impulse response (CIR) in different location areas using a Gaussian mixture model (GMM). We show that by learning the joint distributions of the amplitudes of different delay components, we achieve a more accurate location estimate compared to considering each delay bin independently. Moreover, we develop a similarity metric between sets of CIRs. With this set-based similarity metric, we can further improve the estimation performance, compared to treating each snapshot separately. We showcase the advantages of the proposed methods in multiple application scenarios.
2023. ICASSP 2023 – 2023 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), Rhodes Island (Greece), 2023-06-04 – 2023-06-10. DOI : 10.1109/ICASSP49357.2023.10094570.Improved Belief Propagation Decoding of Turbo Codes
Turbo codes have been successfully adopted in 4G LTE, which can approach the channel capacity with Bahl-Cocke-Jelinek-Raviv (BCJR) decoding. With the evolution from 4G LTE to 5G NR, there is a demand to design a unified channel decoder that supports both LTE Turbo codes and NR low-density parity-check (LDPC) codes. One solution is to employ belief propagation (BP) decoding on the bipartite Tanner graph for both codes. However, although MacKay pointed out that Turbo codes have a sparse parity-check matrix, the existence of 4-cycles in such a matrix severely deteriorates the performance of BP decoding. In this paper, we propose two polynomial-based methods to optimize the parity-check matrix of Turbo codes by improving the sparsity while also removing 4-cycles and even 6-cycles compared to the original matrix. Simulation results show that the improved BP decoding for Turbo codes halves the error-correction performance gap between the original BP decoding and BCJR decoding, which is a promising step towards the unified channel decoder design based on the BP algorithm.
2023. ICASSP 2023 – 2023 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), Rhodes Island (Greece), 2023-06-04 – 2023-06-10. DOI : 10.1109/ICASSP49357.2023.10097058.Low Power LDPC Decoding by Reliable Voltage Down-Scaling
Low-Density Parity-Check (LDPC) decoder is among the power hungry building blocks of wireless communication systems. Voltage scaling down to Near-Threshold (NT) voltages substantially improves energy efficiency, in theory up 10x. However, tuning the voltage and clock frequency to the optimum error free operating point is challenging. This is mainly due to exacerbated sensitivity to Process, Voltage and Temperature (PVT) variations at reduced voltages. By definition, in many telecommunication standards, a Cyclic Redundancy Check (CRC) error detection is carried out after each forward error correction operation, e.g., LDPC decoding. Given channel information, successful CRC checking opens an opportunity for “safe” voltage down-scaling and optimum frequency tuning of LDPC decoder hardware. The strategy is explored on a Zynq System-on-Chip with CRC guiding the adaptive voltage scaling with microcontroller and LDPC decoder residing in different voltage islands. Around 40% power saving was achieved with negligible degradation in throughput.
2023. 9th IEEE Nordic Circuits and Systems Conference (NorCAS), Aalborg, DENMARK, OCT 31-NOV 01, 2023. DOI : 10.1109/NorCAS58970.2023.10305442.Spreading Factor assisted LoRa Localization with Deep Reinforcement Learning
Most of the developed localization solutions rely on RSSI fingerprinting. However, in the LoRa networks, due to the spreading factor (SF) in the network setting, traditional fingerprinting may lack representativeness of the radio map, leading to inaccurate position estimates. As such, in this work, we propose a novel LoRa RSSI fingerprinting approach that takes into account the SF. The performance evaluation shows the prominence of our proposed approach since we achieved an improvement in localization accuracy by up to 6.67% compared to the state-of-the-art methods. The evaluation has been done using a fully connected deep neural network (DNN) set as the baseline. To further improve the localization accuracy, we propose a deep reinforcement learning model that captures the ever-growing complexity of LoRa networks and copes with their scalability. The obtained results show an improvement of 48.10% in the localization accuracy compared to the baseline DNN model.
2023. 97th IEEE Vehicular Technology Conference (VTC-Spring), Florence, ITALY, JUN 20-23, 2023. DOI : 10.1109/VTC2023-Spring57618.2023.10200189.2022
Journal Articles
A Maximum-Likelihood-Based Two-User Receiver for LoRa Chirp Spread-Spectrum Modulation
Long Range (LoRa) is an emerging low-power wide-area network technology offering long-range wireless connectivity to Internet of Things (IoT) devices. For energy efficiency reasons, LoRa end nodes implement a nonslotted ALOHA multiple access scheme to transmit packets to the gateway. Due to the lack of synchronization between end nodes, collisions between uplink packets have been identified as the main obstacle to the scaling of dense LoRa networks. To tackle this issue, we present in this article a LoRa receiver that is capable of decoding colliding packets from two interfering end nodes. The proposed two-user detector is derived from the maximum-likelihood principle using a detailed model of two colliding LoRa packets. As the complexity of the maximum-likelihood sequence estimation is prohibitive, complexity-reduction techniques are introduced to enable practical implementations of the receiver. An in-depth performance analysis highlights that the proposed two-user detector inherently leverages the differences in received power, time offsets, and frequency offsets between the users to separate and demodulate their respective signals. To demonstrate the practicality of the proposed detector, an interference-robust synchronization algorithm is then designed and evaluated. Simulation results indicate that a LoRa receiver combining the proposed synchronization algorithm and two-user detector is capable of detecting and demodulating two interfering users with satisfactorily error rates.
Ieee Internet Of Things Journal. 2022. Vol. 9, num. 22, p. 22993 – 23007. DOI : 10.1109/JIOT.2022.3186732.Ultra-High-Throughput EMS NB-LDPC Decoder with Full-Parallel Node Processing
This paper presents an ultra-high-throughput decoder architecture for NB-LDPC codes based on the Hybrid Extended Min-Sum algorithm. We introduce a new processing block that updates a check node and its associated variable nodes in a fully pipelined way, thus allowing the decoder to process one row of the parity check matrix per clock cycle. The work specifically focuses on a rate 5/6 code of size (N, K) = (144, 120) symbols over GF(64). The synthesis results on a 28-nm technology show that for a 0.789 M NAND-gates complexity complexity, the architecture reaches a decoding throughput of 0.9 Gbps with 30 decoding iterations. Compared to the 5G binary LDPC code of the same size and code rate, the proposed architecture offers a gain of 0.3 dB at a Frame Error Rate of 10(-3).
Journal Of Signal Processing Systems For Signal Image And Video Technology. 2022. DOI : 10.1007/s11265-022-01795-y.Design for Test With Unreliable Memories by Restoring the Beauty of Randomness
This article presents a design-for-test methodology for embedded memories. The methodology relies on a fully random fault model of post-fabrication errors, which results in a low-overhead test strategy. The methodology’s effectiveness is demonstrated on an embedded system with faulty memories.
Ieee Design & Test. 2022. Vol. 39, num. 2, p. 112 – 120. DOI : 10.1109/MDAT.2021.3081687.A Low-Complexity LoRa Synchronization Algorithm Robust to Sampling Time Offsets
LoRaWAN is nowadays one of the most popular protocols for low-power Internet of Things communications. Although its physical layer, namely LoRa, has been thoroughly studied in the literature, aspects related to the synchronization of LoRa receivers have received little attention so far. The estimation and correction of carrier frequency and sampling time offsets (STOs) are, however, crucial to attain the low sensitivity levels offered by the LoRa spread-spectrum modulation. The goal of this article is to build a low-complexity, yet efficient synchronization algorithm capable of correcting both offsets. To this end, a complete analytical model of a LoRa signal corrupted by these offsets is first derived. Using this model, we propose a new estimator for the STO. We also show that the estimations of the carrier frequency and the STOs cannot be performed independently. Therefore, to avoid a complex joint estimation of both offsets, an iterative low-complexity synchronization algorithm is proposed. To reach a packet error rate of 10(-3), performance evaluations show that the proposed receiver requires only 1 or 2 dB higher signal-to-noise ratio than a theoretical perfectly synchronized receiver, while incurring a very low computational overhead.
Ieee Internet Of Things Journal. 2022. Vol. 9, num. 5, p. 3756 – 3769. DOI : 10.1109/JIOT.2021.3101002.Dynamic SCL Decoder With Path-Flipping for 5G Polar Codes
Since polar were ratified as part of the 5G standard, low-complexity polar decoders with close-to-optimum error-rate performance have received significant attention. Compared to successive cancellation (SC) decoding, both SC list and SC flip decoding can improve the error rate performance by increasing the number of considered candidate solutions. The combination of both strategies leads to SC list flip (SCLF) decoding, which can provide a tradeoff between error rate performance and area as well as energy. In this letter, we derive a new flip metric for the SCLF decoding process, based on which we propose the dynamic SCLF (D-SCLF) decoding algorithm. Moreover, we exploit the distributed CRC defined in the 5G standard to further optimize the D-SCLF decoding. Numerical results show that for the downlink control channel, our D-SCLF decoder with a list size of only four and only three additional attempts can achieve the performance of a regular list decoder with a list size of eight, leading to an overall memory and average complexity (energy) reduction.
Ieee Wireless Communications Letters. 2022. Vol. 11, num. 2, p. 391 – 395. DOI : 10.1109/LWC.2021.3129470.64-kB 65-nm GC-eDRAM With Half-Select Support and Parallel Refresh Technique
Gain-cell-embedded DRAM (GC-eDRAM) is an attractive alternative to traditional 6T SRAM, as it offers higher density, lower leakage power, and two-ported functionality. However, its refresh requirement also results in power consumption and memory access limitations. In this letter, we present a GC-eDRAM architecture designed to overcome the refresh disadvantages using a novel technique for improving the availability of the memory. In addition, by using a read-before-write mechanism, half select is supported. The macro avoids the need for supply boosting by employing 3T-1C bitcells and also integrates a replica bit line for optimal access timing to improve performance and power consumption. A 64- kB GC-eDRAM macro was fabricated in a 65- nm process technology, providing a 40% area reduction compared to a 6T SRAM cell, while achieving a 99.99% bit yield with a 16 mu s retention time.
Ieee Solid-State Circuits Letters. 2022. Vol. 5, p. 170 – 173. DOI : 10.1109/LSSC.2022.3182531.A Sequence Repetition Node-Based Successive Cancellation List Decoder for 5G Polar Codes: Algorithm and Implementation
Due to the low-latency and high-reliability require-ments of 5G, low-complexity node-based successive cancellation list (SCL) decoding has received considerable attention for use in 5G communications systems. By identifying special constituent codes in the decoding tree and immediately decoding these, node-based SCL decoding provides a significant reduction in decoding latency compared to conventional SCL decoding. However, while there exists many types of nodes, the current node-based SCL decoders are limited by the lack of a more generalized node that can effi-ciently decode a larger number of different constituent codes to further reduce the decoding time. In this paper, we extend a recent generalized node, the sequence repetition (SR) node, to SCL decod-ing, and describe the first implementation of an SR-List decoder. By merging certain SR-List decoding operations and applying various optimizations for 5G New Radio (NR) polar codes, our optimized SR-List decoding algorithm increases the throughput by almost 2x compared to a similar state-of-the-art node-based SCL decoder. We also present our hardware implementation of the optimized SR-List decoding algorithm which supports all 5G NR polar codes. Synthesis results show that our SR-List decoder can achieve a 2.94 Gbps throughput and 6.70 Gbps/mm2 area efficiency for L = 8.
Ieee Transactions On Signal Processing. 2022. Vol. 70, p. 5592 – 5607. DOI : 10.1109/TSP.2022.3216921.Conference Papers
Increasing Cellular Network Energy Efficiency for Railway Corridors
Modern trains act as Faraday cages making it challenging to provide high cellular data capacities to passengers. A solution is the deployment of linear cells along railway tracks, forming a cellular corridor. To provide a sufficiently high data capacity, many cell sites need to be installed at regular distances. However, such cellular corridors with high power sites in short distance intervals are not sustainable due to the infrastructure power consumption. To render railway connectivity more sustainable, we propose to deploy fewer high-power radio units with intermediate low-power support repeater nodes. We show that these repeaters consume only 5% of the energy of a regular cell site and help to maintain the same data capacity in the trains. In a further step, we introduce a sleep mode for the repeater nodes that enables autonomous solar powering and even eases installation because no cables to the relays are needed.
2022. 25th Design, Automation and Test in Europe Conference and Exhibition (DATE), ELECTR NETWORK, Mar 14-23, 2022. p. 1103 – 1106. DOI : 10.23919/DATE54114.2022.9774757.Beam Selection and Tracking for Amplify-and-Forward Repeaters
Mobile network operators constantly have to upgrade their cellular network to satisfy the public’s hunger for increasing data capacity. However, regulatory limits regarding allowed electromagnetic field strength on existing cell sites often limit or prevent the installation of new cells on additional frequency bands. Further densifying the radio access infrastructure by means of new cell sites takes time because of the site acquisition, the need for permissions, and the civil construction. To this end, amplify-and-forward repeaters are a cost-efficient method to provide mobile wireless data capacity to concealed areas outdoors, inside buildings, or vehicles, where low signal levels drastically limit the achievable capacity or even prevent communication with distant base stations. Repeaters with high-gain beamforming antennas can be used to decrease the path loss and thereby offer a higher capacity than user equipment otherwise could achieve. Since those highly directional beams must be constantly adjusted, we propose methods to align and track the beam even without having access to in-band beam control mechanisms. Numerical analyses with measurement data show that such non-invasive (system data-agnostic) methods are feasible and only incur a few percent in throughput reduction.
2022. IEEE 95th Vehicular Technology Conference: (VTC-Spring), Helsinki, FINLAND, Jun 19-22, 2022. DOI : 10.1109/VTC2022-Spring54318.2022.9860548.List Ordered Statistics Decoders for Polar Codes
The de-facto standard decoding algorithm for polar codes, successive cancellation list (SCL) decoding, is a breadth-first search algorithm. By keeping a list of candidate codewords, SCL decoding improves the performance as the list size L increases. However, generating this list requires storing L copies of internal log-likelihood ratios (LLRs). On the contrary, near-maximum likelihood (near-ML) decoding does not require internal LLRs. This algorithm returns the ML solution of the coded bits in a finite search space and shows superior error-rate performance for high code rates. In this paper, we propose a novel near-ML decoding algorithm based on the breadth-first search, a list ordered statistics decoding (List-OSD) algorithm, which estimates the coded bits starting from the most reliable ones. Simulation results show that the List-OSD exhibits 0.25 dB gain compared with SCL decoding at a frame error rate of 10(-3) for (128, 105) polar code with much less memory consumption for LLRs. Moreover, we design the pipelined hardware architecture of the proposed algorithm based on SMIC 65 nm technology, delivering a 353.14 Mbps throughput and 0.7 mm(2) area when L = 32. To the best of our knowledge, this is the first work that presents ASIC implementation results for OSD-based decoders.
2022. 56th Asilomar Conference on Signals, Systems, and Computers, ELECTR NETWORK, Oct 31-Nov 02, 2022. p. 628 – 633. DOI : 10.1109/IEEECONF56349.2022.10051951.Fast Sequence Repetition Node-Based Successive Cancellation List Decoding for Polar Codes
Compared with the bit-wise successive cancellation list (SCL) decoding of polar codes, the node-based Fast SCL decoding significantly reduces the decoding latency by identifying special constituent codes and decoding these in parallel. To further reduce the latency of current Fast SCL decoders, we first propose a fast sequence repetition (SR) node-based SCL (Fast SR-SCL) decoding algorithm, which only involves one type of node in the SCL decoding tree. Furthermore, we employ the adaptive path splitting (APS) strategy to terminate the path splitting in the SR node early, without degrading the error-correcting performance. Numerical results show that for 5G uplink codes with a length of 1024 and rates of 1/4, 1/2, and 3/4, our decoder can deliver the same decoding performance while reducing the average latency by 34.5%, 38.0%, and 39.6% compared with the state-of-the-art Fast SCL decoder for a list size L = 8.
2022. IEEE International Conference on Communications (ICC), Seoul, SOUTH KOREA, May 16-20, 2022. p. 116 – 122. DOI : 10.1109/ICC45855.2022.9839072.Theses
Cellular Fronthauling for Data Capacity Increase in Underserved Spaces
The increase in wireless data traffic continues and is a product of several factors. First, new technologies and capabilities enable new use cases for which new products emerge. Then, with the growing user adoption over time, the data traffic is further increased. As a result, the actual growth numbers vary year over year, but the trend is sustained. People nowadays also expect ubiquitous availability of mobile connectivity and use cloud services or stream music and video. Moreover, an increasing number of machines are being connected too. While mobile network operators (MNOs) deploy outdoor cellular networks, a significant portion of the mobile data traffic is consumed or originates indoors or while traveling in trains. To this end, MNOs constantly upgrade their networks to satisfy the capacity demand. Additionally, new and wider frequency bands are made available by regulators. However, the mid-band frequencies between 3 GHz and 6 GHz poorly enter buildings, and outdoor cells on millimeter-wave (mmWave) frequencies (24 GHz to 100 GHz) are practically unusable indoors and inside modern railways with their metallic hull and coated windows acting as Faraday cages. The increasing network densification with local cell sites (also inside shielded structures) can provide a partial solution. However, the required fiber and roll-out process is very costly and takes time. Moreover, providing sufficient capacity for the increasing data traffic demand inside buildings and trains is difficult where an optical fiber backhaul is too costly or impossible. State-of-the-art solutions are analyzed for solving the indoor capacity challenge. However, we found that these rely on or impact the existing outdoor cellular network and thus cannot provide additional capacity to the network that is used indoors. As a solution, we propose the mmWave bridge, an amplify-and-forward out-of-band repeater concept. It is a radio access technology (RAT) transparent and cost-efficient method to fronthaul mobile cells. Moreover, it provides wireless data capacity inside buildings, vehicles, or concealed areas outdoors, where low signal levels drastically limit the achievable capacity or prevent communication with distant base stations. The cellular signal of a base station is fronthauled over newly available mmWave frequencies outdoor without interfering with the existing cellular network. Mid-band frequencies are used indoors to provide sufficient coverage beyond rooms and, at the same time, benefit from the outdoor-to-indoor attenuation reducing possible interference. The benefits of the mmWave bridge are described, and we discuss the corresponding challenges and solutions. Moreover, a hardware prototype was developed based on commercial off-the-shelf components. We have tested our prototype in three use cases to demonstrate the functionality and compatibility with commercial infrastructure and mobile terminals and present the measurement results. We can show that the entire capacity of a mobile cell can be fronthauled over distances in relation to the mmWave frequency propagation and signal power. Finally, the amplify-and-forward out-of-band repeater concept is generalized regarding the fronthaul carrier frequency. The use of beamforming antennas on RAT-transparent repeaters without access to in-band beam control is investigated, and a solution with a minimal impact of a few percent in throughput reduction is presented.
Lausanne, EPFL, 2022.Working Papers
ErgoDEC: A Fault Tolerant 28 nm LDPC Decoder Providing Stable FER Quality with Unreliable Memories
Communication systems have been associated with an inherent fault tolerance to hardware reliability issues. Therefore, many publications have studied the impact of such issues on, for example, channel (mostly LDPC) decoders. Since all of these studies are so far only based on simulations, we implement an LDPC decoder in a 28nm technology, specifically to verify the corresponding assumptions and quality metrics. Our measurements indicate a large quality spread across a population of dies, which is not visible when considering an average quality across simulation runs with random fault injection. Such a quality spread is detrimental since in high-volume manufacturing a certain minimum quality must be guaranteed for all accepted dies. The latter is only possible with an ergodic behaviour in which the quality of any faulty die is indicative for the quality of all dies in the population. To alleviate this issue, ErgoDEC provides simple architectural measure to restore an ergodic behaviour in the presence of memory reliability issues. This more stable quality can also be observed in our measurements and can be exploited to further improve the resilience against faulty bits.
2022
Student Projects
Space Localisation
2022.2021
Journal Articles
On the Advantage of Coherent LoRa Detection in the Presence of Interference
It has been shown that the coherent detection of long range (LoRa) signals only provides marginal gains of around 0.7 dB on the additive white Gaussian noise (AWGN) channel. However, ALOHA-based massive Internet-of-Things systems, including LoRa, often operate in the interference-limited regime. Therefore, in this work, we examine the performance of the LoRa modulation with coherent detection in the presence of interference from another LoRa user with the same spreading factor. We derive rigorous symbol- and frame error rate (FER) expressions as well as bounds and approximations for evaluating the error rates. The error rates predicted by these approximations are compared against error rates found by Monte Carlo simulations and shown to be very accurate. We also compare the performance of LoRa with coherent and noncoherent receivers and we show that the coherent detection of LoRa is significantly more beneficial in interference scenarios than in the presence of only AWGN. For example, we show that coherent detection leads to a 2.5-dB gain over the standard noncoherent detection for a signal-to-interference ratio (SIR) of 3 dB and up to a 10-dB gain for an SIR of 0 dB. Moreover, we show that with coherent detection it is easier to obtain useful and relevant FER values even for negative SIR values.
Ieee Internet Of Things Journal. 2021. Vol. 8, num. 14, p. 11581 – 11593. DOI : 10.1109/JIOT.2021.3058792.An Energy-Autonomous Wireless Sensor With Simultaneous Energy Harvesting and Ambient Light Sensing
Wireless sensor nodes (WSNs) are generally powered by batteries, which results in a substantial limitation to the places where the nodes can be installed, to the maximum number of deployable devices, and to the node lifetime. To meet the demand for Internet-of-Things (IoT) applications that require a large number of maintenance-free, low cost, wireless sensor nodes, this paper proposes a wireless sensor platform with a single photovoltaic transducer that performs the dual role of harvesting energy and sensing ambient light. This dual use allows even smaller and cheaper nodes that do not require any form of supporting external power, with a reduced component count. The device implements off-the-shelf components on a 2x2cm(2) printed circuit board (PCB) with a thickness of 0.45cm. It features Bluetooth Low Energy (BLE) communication and can harvest and sense indoor ambient light with a limit of detection of 200 lux.
Ieee Sensors Journal. 2021. Vol. 21, num. 12, p. 13744 – 13752. DOI : 10.1109/JSEN.2021.3068134.Adding Indoor Capacity without Fiber Backhaul: An mmWave Bridge Prototype
Today, a large portion of mobile data traffic is consumed behind the shielding walls of buildings or in the Faraday cage of trains. This renders cellular network coverage from outdoor cell sites difficult. Indoor small cells and distributed antennas along train tracks are often considered as a solution, but the cost and the need for optical fiber backhaul are often prohibitive. To alleviate this issue, we describe an out-of-band repeater that converts a sub-6 GHz cell signal from a small cell installed at a cell tower to an mmWave frequency for the fronthaul to buildings or distributed antenna sites, where the signal is downconverted to the original frequency and emitted, for example, inside a building. This concept does not require fiber deployment, provides backward compatibility to equipment already in use, and additional indoor capacity is gained while outdoor networks are offloaded. The architecture and hardware prototype implementation is described, and measurements are reported to demonstrate the functionality and compatibility with commercial infrastructure and mobile terminals.
Ieee Communications Magazine. 2021. Vol. 59, num. 4, p. 110 – 115. DOI : 10.1109/MCOM.001.2000722.Polarization-Adjusted Convolutional (PAC) Codes: Sequential Decoding vs List Decoding
In the Shannon lecture at the 2019 International Symposium on Information Theory (ISIT), Arikan proposed to employ a one-to-one convolutional transform as a pre-coding step before the polar transform. The resulting codes of this concatenation are called polarization-adjusted convolutional (PAC) codes. In this scheme, a pair of polar mapper and demapper as pre- and postprocessing devices are deployed around a memoryless channel, which provides polarized information to an outer decoder leading to improved error correction performance of the outer code. In this paper, the list decoding and sequential decoding (including Fano decoding and stack decoding) are first adapted for use to decode PAC codes. Then, to reduce the complexity of sequential decoding of PAC/polar codes, we propose (i) an adaptive heuristic metric, (ii) tree search constraints for backtracking to avoid exploration of unlikely sub-paths, and (iii) tree search strategies consistent with the pattern of error occurrence in polar codes. These contribute to the reduction of the average decoding time complexity from 50% to 80%, trading with 0.05 to 0.3 dB degradation in error correction performance within FER = 10(-3) range, respectively, relative to not applying the corresponding search strategies. Additionally, as an important ingredient in Fano decoding of PAC/polar codes, an efficient computation method for the intermediate LLRs and partial sums is provided. This method is effective in backtracking and avoids storing the intermediate information or restarting the decoding process. Eventually, all three decoding algorithms are compared in terms of performance, complexity, and resource requirements.
Ieee Transactions On Vehicular Technology. 2021. Vol. 70, num. 2, p. 1434 – 1447. DOI : 10.1109/TVT.2021.3052550.Enhancing the Reliability of Dense LoRaWAN Networks With Multi-User Receivers
LoRaWAN is a low-power wireless technology that provides long-range connectivity to battery-powered Internet of Things (IoT) devices. To minimize the energy consumption of the IoT nodes, LoRaWAN networks use for the uplink a pure non-slotted ALOHA multiple access scheme. Since the devices are not synchronized in time, collisions between uplink packets are the main source of errors when the number of nodes becomes important. To improve the reliability of dense LoRaWAN networks, we propose in this paper a successive interference cancellation LoRa receiver capable of decoding frames from two colliding users. The proposed two-user detector leverages the bit-interleaved coded modulation scheme of LoRa to improve the detection and cancellation of the strongest interfering user. We show that in the presence of two interfering users, the usage of a low coding-rate and iterative soft-detection are essential to attain error rates sufficiently close to the single-user scenario. Using network-level simulations, we subsequently evaluate the gains of our proposed two-user receiver in a realistic LoRaWAN network. To this end, we build advanced models of the studied receivers using Monte-Carlo simulations at the physical layer. For an overall packet error rate of 1%, simulation results indicate that a LoRaWAN network employing our two-user detector may serve 4.7 times more devices than a network with only a single-user receiver at the gateway.
Ieee Open Journal Of The Communications Society. 2021. Vol. 2, p. 2725 – 2738. DOI : 10.1109/OJCOMS.2021.3134091.Low-Cost Side-Channel Secure Standard 6T-SRAM-Based Memory With a 1% Area and Less Than 5% Latency and Power Overheads
Side-channel attacks constitute a concrete threat to IoT systems-on-a-chip (SoCs). Embedded memories implemented with 6T SRAM macrocells often dominate the area and power consumption of these SoCs. Regardless of the computational platform, the side-channel sensitivity of low-hierarchy cache memories can incur significant overhead to protect the memory content (i.e., data encryption, data masking, etc.). In this manuscript, we provide a silicon proof of the effectiveness of a low cost side-channel attack protection that is embedded within the memory macro to achieve a significant reduction in information leakage. The proposed solution incorporates low-cost impedance randomization units, which are integrated into the periphery of a conventional 6T SRAM macro in fine-grain memory partitions, providing possible protection against electromagnetic adversaries. Various blocks of unprotected and protected SRAM macros were designed and fabricated in a 55 nm test-chip. The protected ones had little as 1% area overhead and less than 5% performance and power penalties compared to a conventional SRAM design. To evaluate the security of the proposed solution, we applied a robust mutual information metric and an adaptation to the memory context to enhance this evaluation framework. Assessment of the protected memory demonstrated a significant information leakage reduction from 8 bits of information exposed after only 100 cycles of attack to less than similar to 1.5 bits of mutual information after 160K traces. The parametric nature of the protection mechanisms are discussed while specifying the proposed design parameters. Overall, the proposed methodology enables designs with higher security-level at a minimal cost.
Ieee Access. 2021. Vol. 9, p. 91764 – 91776. DOI : 10.1109/ACCESS.2021.3088991.Conference Papers
Intrinsically Self-powered, Battery-free, and Sensor-free Ambient Light Control System
This paper deals with an energy-autonomous wireless sensing system with an indirect measurement of the light energy for ambient lighting sensing. The system can perform power measurements without battery or external power sources and generates time-domain data. The advantage of the proposed solution consists of greater energy efficiency, a reduced number of components, higher miniaturization, and a reduction in implementation costs. The device is implemented using off-the-shelf components on a printed circuit board (PCB) with the size of 2 x 2 cm(2) and the thickness of 0.45 cm, it can harvest and detect sunlight ambient light as low as 200 lux.
2021. 20th IEEE Sensors Conference, ELECTR NETWORK, Oct 31-Nov 04, 2021. DOI : 10.1109/SENSORS47087.2021.9639712.A Novel RF Spectrum Monitoring Architecture for an Ultra-Low-Power Wi-Fi Geopositioning System
Wireless radio consumes the highest power in many systems and must be activated wisely to save power especially in battery-powered systems. Hence, gathering insight into the spectrum activity is needed to control the wireless radio. In this work, classic full-band Fast Fourier Transform (FFT) and sequential digital spectrum scanning systems are presented with their high energy consumption and latency drawbacks. A context-aware, multi-layer-duty-cycled, multi-channel, ultra-low-power analog spectrum monitoring architecture is proposed as a solution to the drawbacks of the classic systems with the emphasis on Wi-Fi signal detection for a Basic Service Set Identifier (BSSID)-based geopositioning shipment tracking application. The proposed architecture provides more than 3 order of magnitude power saving in detection compared to the classic sequential spectrum scanning while maintaining the full functionality under vast variety of operating conditions.
2021. 19th IEEE International New Circuits and Systems Conference (NEWCAS), ELECTR NETWORK, Jun 13-16, 2021. DOI : 10.1109/NEWCAS50681.2021.9462768.Improving Railway Track Coverage with mmWave Bridges: A Measurement Campaign
Bringing cellular capacity into modern trains is challenging because they act as Faraday cages. Building a radio frequency (RF) corridor along the railway tracks ensures a high signal-to-noise ratio and limits handovers. However, building such RF corridors is difficult because of the administrative burden of excessive formalities to obtain construction permissions and costly because of the sheer number of base stations. Our contribution in this paper is an unconventional solution of mmWave fronthauled low-power out-of-band repeater nodes deployed in short intervals on existing masts between high-power macro cell sites. The paper demonstrates the feasibility of the concept with an extensive measurement campaign on a commercial railway line. The benefit of using many low-power nodes with low-gain antennas compared to a baseline with only high-gain macro antennas is discussed, and the coverage improvement is evaluated. Based on the measurement results, a simple path loss model is calibrated. This model allows evaluation of the potential of the mmWave repeater architecture to increase the macro cell inter-site distance and reduce deployment costs.
2021. 1st Workshop on 5G Measurements, Modeling, and Use Cases (5G-MeMU), ELECTR NETWORK, Aug 23, 2021. p. 8 – 13. DOI : 10.1145/3472771.3472774.Dynamic Range and Complexity Optimization of Mixed-Signal Machine Learning Systems
Audio processing had been in demand throughout the electronic era. Recent advances in neural networks increased the demand on audio processing for speech recognition applications. In this work, a rigorous study on the dynamic range and system complexity optimization is presented for a mixed-signal keyword spotting system. The proposed system consists of an analog feature extractor and a neural network based keyword classifier. The results showed that with the proposed method, more than an order of magnitude power saving can be achieved in the analog feature extraction compared to the digital state-of-the-art counterpart.
2021. IEEE International Symposium on Circuits and Systems (IEEE ISCAS), Daegu, SOUTH KOREA, May 22-28, 2021. DOI : 10.1109/ISCAS51556.2021.9401331.ComplexBeat: Breathing Rate Estimation from Complex CSI
In this paper, we explore the use of channel state information (CSI) from a WiFi system to estimate the breathing rate of a person in a room. In order to extract WiFi CSI components that are sensitive to breathing, we propose to consider the delay domain channel impulse response (CIR), while most state-of-the-art methods consider its frequency domain representation. One obstacle while processing the CSI data is that its amplitude and phase are highly distorted by measurement uncertainties. We thus also propose an amplitude calibration method and a phase offset calibration method for CSI measured in orthogonal frequency-division multiplexing (OFDM) multiple-input multiple-output (MIMO) systems. Finally, we implement a complete breathing rate estimation system in order to showcase the effectiveness of our proposed calibration and CSI extraction methods.
2021. 35th IEEE Workshop on Signal Processing Systems (IEEE SiPS), Coimbra, PORTUGAL, Oct 19-21, 2021. p. 217 – 222. DOI : 10.1109/SiPS52927.2021.00046.Theses
Adaptive Body Biasing in Strong Body Factor Technologies
With the advent of intelligent sensor nodes in everyday life, low power aspects of system design become more and more important. Adaptive body biasing is a promising methodology to achieve dynamic adaptation of the tradeoff between performance and energy by shifting the threshold voltage of transistors in a digital design. This approach, in combination with a low supply voltage, provides a strong knob to the designer to rapidly shift the circuits operating point from deep sub threshold operation for slow and low leakage retention, to fast, higher performance operation to for short, but demanding tasks. This thesis concentrates on such designs using the deeply depleted channel technology in which body control is particularly effective. The first part of this thesis is dedicated to strategies and tools supporting the digital design process of circuits using adaptive body biasing. A methodology to compare a standard cell library characterised in different operating points, defined by supply voltage, process corner, temperature, and bias points is presented first. Next, we present a methodology to exhaustively and rapidly map out the supply-voltage/bias-voltage design space using a heavily pruned cell library. We extract speed and power of a simple example design across the entire design space and show a methodology to scale the characteristics of the small reference up to a more complex design. In a case study this modelling approach achieves an error of less than 1% on the total power relative to an actual characterisation of the full library at the same design point. The second part of the thesis analyses three chips implementing different biasing schemes in USJC 55nm DDC. The first two were designed by CSEM with components and measurements contributed from this thesis while the third one was entirely designed for this thesis. The first chip utilises a biasing scheme based on the first order approximation that the circuit speed is proportional to the on-current which can be driven by PMOS and NMOS transistors. This is implemented using an analog control loop, setting both PMOS and NMOS on-currents equal to a reference current provided by current DAC. The SoC characterisation is presented with the objective of identifying suitable operating modes and bias points, including a reliability and retention analysis of the SRAM. A series of ring oscillators constructed from the most common standard cells has also been integrated and provides measurement support for the first part of the thesis. The second chip extends the Calanda biasing scheme with a secondary regulation loop that is based on an FLL, designed in this thesis, in combination with a configurable standard cell based ring oscillator. The user can directly program a target frequency and the biasing system regulates the current DAC accordingly. We show that this approach effectively overcomes the drawbacks of the current based approach resulting in an effective regulation. Finally the third chip presents a novel biasing scheme that was designed in this thesis and is tailored toward simplicity. It utilises two constant voltages for the PMOS bias to switch between retention and operation. A charge pump controlled by a standard cell compatible distributed on current balance sensor regulates the NMOS bias such that the on-currents of PMOS and NMOS match. We show that this simple approach, in conjunction with a well chosen operating point, can be efficient across corners.
Lausanne, EPFL, 2021.Student Projects
Le fil d’Ariane, redécouverte du Musée des Thermes et de Cluny à Paris
Le musée de Cluny est le musée national du Moyen Âge situé au cœur de Paris. Le site est composé d’un ensemble de bâtiments de différentes époques, dont deux monuments historiques : les thermes gallo-romains du Nord et l’hôtel médiéval des abbés de Cluny, prototype du modèle des hôtels parisiens entre cour et jardin. Suite à des ajouts successifs, notamment les extensions muséales des XIXe et XXIe siècles, le musée apparaît aujourd’hui comme un ensemble hétéroclite de bâtiments, un archipel perdu au milieu de la ville. L’objectif du projet est de rendre sa cohérence à l’ensemble, tout en respectant le caractère du site et ses évolutions à travers le temps. D’abord, la visibilité des monuments et du musée dans la ville est renforcée, d’un côté par l’intégration et la transformation dans le site du square situé au nord de l’îlot, qui relie le musée aux grands boulevards adjacents et, de l’autre, par la modification de l’implantation de l’accueil, en façade nord, accessible depuis la place nouvellement créée sur ce jardin. Basée sur une trame régulière, une série d’interventions structurelles permet d’unifier l’ensemble du site. Celles-ci se déclinent en éléments de toiture, de sol ou encore en dispositifs muséaux. Leurs variations répondent aux besoins des lieux : recréation d’un parcours simple qui aide à la compréhension du site, couverture des ruines exposées, création de nouveaux espaces intérieurs et extérieurs et intégration de l’ensemble dans la ville.
2021.2020
Journal Articles
Replica Bit-Line Technique for Internal Refresh in Logic-Compatible Gain-Cell Embedded DRAM
Embedded memories, mostly implemented with static random access memory (SRAM), dominate the area and power of integrated circuits. Gain-cell embedded DRAM (GC-eDRAM) is an alternative to SRAM due to its high density, low power consumption, and two-ported functionality. However, GC-eDRAM requires periodic refresh cycles to maintain its data due to its dynamic storage mechanism. The refresh operation is typically handled at a memory controller level, resulting in an energy overhead and limited memory availability. In this paper, we propose a new approach for the realization of the refresh operation using an internal refresh mechanism, which supports an efficient row-wise refresh operation within a single clock cycle, providing 100% write access availability at a reduced refresh latency and power. An 8 kbit GC-eDRAM array with integrated internal refresh and replica bit-line was implemented, demonstrating up-to 30% reduced refresh latency and 65% reduced read energy at a low cost of 2.4% array area overhead, compared to a conventional GC-eDRAM array without internal refresh capabilities.
Microelectronics Journal. 2020. Vol. 101, p. 104781. DOI : 10.1016/j.mejo.2020.104781.Hardware Implementation of Neural Self-Interference Cancellation
In-band full-duplex systems can transmit and receive information simultaneously and on the same frequency band. However, due to the strong self-interference caused by the transmitter to its own receiver, the use of non-linear digital self-interference cancellation is essential. In this work, we describe a hardware architecture for a neural network-based non-linear self-interference (SI) canceller and we compare it with our own hardware implementation of a conventional polynomial based SI canceller. Our results show that, for the same SI cancellation performance, the neural network canceller has an 8.1x smaller area and requires 7.7x less power than the polynomial canceller. Moreover, the neural network canceller can achieve 7 dB more SI cancellation while still being 1.2x smaller than the polynomial canceller and only requiring 1.3x more power. These results show that NN-based methods applied to communications are not only useful from a performance perspective, but can also lead to order-of-magnitude implementation complexity reductions.
IEEE Journal of Emerging and Selected Topics in Circuits and Systems. 2020. Vol. 10, num. 2, p. 204 – 216. DOI : 10.1109/JETCAS.2020.2992370.Artificial Intelligence for 5G and Beyond 5G: Implementations, Algorithms, and Optimizations
This Special Issue of the IEEE Journal on Emerging and Selected Topics in Circuits and Systems (JETCAS) is dedicated to demonstrating the latest research progress on artificial intelligence for 5G and beyond 5G (B5G) with respect to implementations, algorithms, and optimizations.
IEEE Journal of Emerging and Selected Topics in Circuits and Systems. 2020. Vol. 10, num. 2, p. 145 – 148. DOI : 10.1109/JETCAS.2020.2999944.A Standalone FPGA-Based Miner for Lyra2REv2 Cryptocurrencies
Lyra2REv2 is a hashing algorithm that consists of a chain of individual hashing algorithms, and it is used as a proof-of-work function in several cryptocurrencies. The most crucial and exotic hashing algorithm in the Lyra2REv2 chain is a specific instance of the general Lyra2 algorithm. This work presents the first hardware implementation of the specific instance of Lyra2 that is used in Lyra2REv2. Several properties of the aforementioned algorithm are exploited in order to optimize the design. In addition, an FPGA-based hardware implementation of a standalone miner for Lyra2REv2 on a Xilinx Multi-Processor System on Chip is presented. The proposed Lyra2REv2 miner is shown to be significantly more energy efficient than both a GPU and a commercially available FPGA-based miner. Finally, we also explain how the simplified Lyra2 and Lyra2REv2 architectures can be modified with minimal effort to also support the recent Lyra2REv3 chained hashing algorithm.
Ieee Transactions On Circuits And Systems I-Regular Papers. 2020. Vol. 67, num. 4, p. 1194 – 1206. DOI : 10.1109/TCSI.2020.2970923.Current-Based Data-Retention-Time Characterization of Gain-Cell Embedded DRAMs Across the Design and Variations Space
The rise of data-intensive applications has resulted in an increasing demand for high-density and low-power on-chip embedded memories. Gain-cell embedded DRAM (GC-eDRAM) is a logic-compatible alternative to conventional static random access memory (SRAM) which offers higher density, lower leakage power, and two-ported operation. However, in order to maintain the stored data, GC-eDRAM requires periodic refresh cycles, which are determined according to the worst-case data retention time (DRT) across process, voltage and temperature (PVT) variations. Even though several DRT characterization methodologies have been reported in literature, they often require unfeasible run-times for accurate DRT evaluation, or they result in highly pessimistic design margins due to their inaccuracy. In this work, we propose an current-based DRT (IDRT) characterization methodology that enables accurate DRT evaluation across process variations without the need for a large number of costly electronic design automation (EDA) software licenses. The presented approach is compared with other DRT characterization methodologies for both accuracy and run-time across several gain-cell structures at different process technologies, providing less than a 4% DRT error and over 100x shorter run-time compared to a conventional DRT evaluation methodology.
Ieee Transactions On Circuits And Systems I-Regular Papers. 2020. Vol. 67, num. 4, p. 1207 – 1217. DOI : 10.1109/TCSI.2020.2971695.Gain-Cell Embedded DRAMs: Modeling and Design Space
Among the different types of dynamic random-access memories (DRAMs), gain-cell embedded DRAM (GC-eDRAM) is a compact, low-power, and CMOS-compatible alternative to conventional static random-access memory (SRAM). GC-eDRAM achieves high memory density, as it relies on a storage cell that can be implemented with as few as two transistors and that can be fabricated without additional process steps. However, since the performance of GC-eDRAMs relies on many interdependent variables, the optimization of the performance of these memories for the integration into their hosting system, as well as the design investigation of future GC-eDRAMs, proves to be highly complex tasks. In this context, modeling tools of memories are key enablers for the exploration of this large design space in a short amount of time. In this article, we present GC-eDRAM modeling tool (GEMTOO), the first modeling tool that estimates timing, memory availability, bandwidth, and area of GC-eDRAMs. The tool considers parameters related to technology, circuits, and memory architecture, and it enables the evaluation of architectural transformations as well as advanced transistor-level effects, such as the increase in the access delay due to the deterioration of the stored data. The timing is estimated with a maximum deviation of 15% from postlayout simulations in a 28-nm FD-SOI technology for different memory sizes and architectures. Moreover, the measured random cycle frequency of a GC-eDRAM fabricated in a 28-nm CMOS bulk process is estimated with a 9% deviation when considering 6-sigma random process variations of the bitcells. The proposed GEMTOO modeling tool is used to show the intricacies in design optimization of GC-eDRAMs, and based on the results, optimal design practices are derived.
Ieee Transactions On Very Large Scale Integration (Vlsi) Systems. 2020. Vol. 28, num. 3, p. 646 – 659. DOI : 10.1109/TVLSI.2019.2955933.Design and Decoding of Irregular LDPC Codes Based on Discrete Message Passing
We consider discrete message passing (MP) decoding of low-density parity check (LDPC) codes based on information-optimal symmetric look-up table (LUT). A link between discrete message labels and the associated log-likelihood ratio values (defined in terms of density evolution distributions) is established. This link gives rise to an algebraic structure on the message labels and leads to an interpretation of LUT decoding as a form of quantized belief propagation. We then exploit the algebraic structure for low-complexity LUT decoder designs. Our LUT decoding framework is the first to also apply to irregular LDPC codes by taking into account the degree distribution in a joint LUT design. We exploit the relation between LUT decoding and belief propagation to obtain stability conditions and irregular LDPC code designs optimized for LUT decoding. The resulting decoders outperform floating-point precision min-sum decoders at LUT resolutions as low as 3 bit s for regular codes and 4 bits for irregular codes.
Ieee Transactions On Communications. 2020. Vol. 68, num. 3, p. 1329 – 1343. DOI : 10.1109/TCOMM.2019.2944159.On the Error Rate of the LoRa Modulation With Interference
LoRa is a chirp spread-spectrum modulation developed for the Internet of Things (IoT). In this work, we examine the performance of LoRa in the presence of both additive white Gaussian noise and interference from another LoRa user. To this end, we extend an existing interference model, which assumes perfect alignment of the signal of interest and the interference, to the more realistic case where the interfering user is neither chip- nor phase-aligned with the signal of interest and we derive an expression for the error rate. We show that the existing aligned interference model overestimates the effect of interference on the error rate. Moreover, we prove two symmetries in the interfering signal and we derive low-complexity approximate formulas that can significantly reduce the complexity of computing the symbol and frame error rates compared to the complete expression. Finally, we provide numerical simulations to corroborate the theoretical analysis and to verify the accuracy of our proposed approximations.
Ieee Transactions On Wireless Communications. 2020. Vol. 19, num. 2, p. 1292 – 1304. DOI : 10.1109/TWC.2019.2952584.A 161-mW 56-Gb/s ADC-Based Discrete Multitone Wireline Receiver Data-Path in 14-nm FinFET
This article introduces a wireline receiver (RX) data-path employing discrete multi-tone (DMT) modulation for communicating over electrical links. The DMT RX incorporates a fully digital equalization data-path, with a synthesized and automatically placed and routed digital signal processor (DSP) following a 10-bit time-interleaved pipelined successive-approximation register analog-to-digital converter (TI-PISAR ADC). The prototype RX chip implemented in a 14-nm FinFET process demonstrates a lane data rate of 56 Gb/s dissipating 161 mW including the ADC and the DSP power. The energy efficiency of 1.2 pJ/b for the DSP and 2.9 pJ/b for the entire RX was achieved with the data-rate of 56 Gb/s for communicating over channels exhibiting up to 28-dB loss at 14 GHz with a bit-error-rate (BER) better than 2e-4.
Ieee Journal Of Solid-State Circuits. 2020. Vol. 55, num. 1, p. 38 – 48. DOI : 10.1109/JSSC.2019.2938414.A 1-Mbit Fully Logic-Compatible 3T Gain-Cell Embedded DRAM in 16-nm FinFET
Gain-cell embedded DRAM (GC-eDRAM) is a logic-compatible embedded memory alternative to SRAM, offering higher density, lower leakage power consumption, and an inherent two-ported functionality. However, increased leakage currents and process variations under technology scaling lead to a reduced data retention time (DRT), resulting in increased refresh power and reduced memory availability, currently limiting its implementation to planar 28-nm technologies and above. This letter presents the first GC-eDRAM in 16-nm FinFET technology, featuring a mixed-VT 3T gain-cell structure to minimize the storage node (SN) leakage. The implemented 1-Mbit 3T GC-eDRAM is fully logic-compatible and provides a 2x smaller bitcell size compared to a 6T SRAM with similar design rules, offering the highest density logic-compatible memory cell in 16-nm technology. Measurement results demonstrate a 77-mu s DRT under a 600-mV VDD, which is over 10x longer than previously reported GC-eDRAMs in 28-nm technologies. The memory was fully operational at temperatures spanning -40 degrees C to 125 degrees C and under a supply voltage as low as 450 mV, providing the lowest measured VDDmin and widest temperature range reported in the literature for GC-eDRAM.
Ieee Solid-State Circuits Letters. 2020. Vol. 3, p. 110 – 113. DOI : 10.1109/LSSC.2020.3006496.Conference Papers
Identification of Non-Linear RF Systems Using Backpropagation
In this work, we use deep unfolding to view cascaded non-linear RF systems as model-based neural networks. This view enables the direct use of a wide range of neural network tools and optimizers to efficiently identify such cascaded models. We demonstrate the effectiveness of this approach through the example of digital self-interference cancellation in full-duplex communications where an IQ imbalance model and a non-linear PA model are cascaded in series. For a self-interference cancellation performance of approximately 44.5 dB, the number of model parameters can be reduced by 74% and the number of operations per sample can be reduced by 79% compared to an expanded linear-in-parameters polynomial model.
2020. IEEE International Conference on Communications (IEEE ICC) / Workshop on NOMA for 5G and Beyond, ELECTR NETWORK, Jun 07-11, 2020. DOI : 10.1109/ICCWorkshops49005.2020.9145367.Lupulus: A Flexible Hardware Accelerator For Neural Networks
Neural networks have become indispensable for a wide range of applications, but they suffer from high computational- and memory-requirements, requiring optimizations from the algorithmic description of the network to the hardware implementation. Moreover, the high rate of innovation in machine learning makes it important that hardware implementations provide a high level of programmability to support current and future requirements of neural networks. In this work, we present a flexible hardware accelerator for neural networks, called Lupulus, supporting various methods for scheduling and mapping of operations onto the accelerator. Lupulus was implemented in a 28nm FD-SOI technology and demonstrates a peak performance of 380GOPS/GHz with latencies of 21.4 ms and 183.6 ms for the convolutional layers of AlexNet and VGG-16, respectively.
2020. IEEE International Conference on Acoustics, Speech, and Signal Processing, Barcelona, SPAIN, May 04-08, 2020. p. 1608 – 1612. DOI : 10.1109/ICASSP40776.2020.9054764.Training Channel Selection for Learning-based 1-bit Precoding in Massive MU-MIMO
Learning-based algorithms have gained great popularity in communications since they often outperform even carefully engineered solutions by learning from training samples. In this paper, we show that the selection of appropriate training examples can be important for the performance of such learning-based algorithms. In particular, we consider non-linear 1-bit precoding for massive multi-user MIMO systems using the C2PO algorithm. While previous works have already shown the advantages of learning critical coefficients of this algorithm, we demonstrate that straightforward selection of training samples that follow the channel model distribution does not necessarily lead to the best result. Instead, we provide a strategy to generate training data based on the specific properties of the algorithm, which significantly improves its error floor performance.
2020. IEEE International Conference on Communications (IEEE ICC) / Workshop on NOMA for 5G and Beyond, ELECTR NETWORK, Jun 07-11, 2020. DOI : 10.1109/ICCWorkshops49005.2020.9145443.Coded LoRa Frame Error Rate Analysis
In this work, we study the coded frame error rate (FER) of LoRa under additive white Gaussian noise (AWGN) and under carrier frequency offset (CFO). To this end, we use existing approximations for the bit error rate (BER) of the LoRa modulation under AWGN and we present a FER analysis that includes the channel coding, interleaving, and Gray mapping of the LoRa physical layer. We also derive the LoRa BER under carrier frequency offset and we present a corresponding FER analysis. We compare the derived frame error rate expressions to Monte Carlo simulations to verify their accuracy.
2020. IEEE International Conference on Communications (IEEE ICC) / Workshop on NOMA for 5G and Beyond, ELECTR NETWORK, Jun 07-11, 2020. DOI : 10.1109/ICC40277.2020.9148806.An Open-Source LoRa Physical Layer Prototype on GNU Radio
LoRa is the proprietary physical layer (PHY) of LoRaWAN, which is a popular Internet-of-Things (IoT) protocol enabling low-power devices to communicate over long ranges. A number of reverse engineering attempts have been published in the last few years that helped to reveal many of the LoRa PHY details. In this work, we describe our standard-compatible LoRa PHY software-defined radio (SDR) prototype based on GNU Radio. We show how this SDR prototype can be used to develop and evaluate receiver algorithms for LoRa. As an example, we describe the sampling time offset and the carrier frequency offset estimation and compensation blocks. We experimentally evaluate the error rate of LoRa, both for the uncoded and the coded cases, to illustrate that our publicly available open-source implementation is a solid basis for further research.
2020. 21st IEEE International Workshop on Signal Processing Advances in Wireless Communications (IEEE SPAWC), ELECTR NETWORK, May 26-29, 2020. DOI : 10.1109/SPAWC48557.2020.9154273.Complexity-efficient Fano Decoding of Polarization-adjusted Convolutional (PAC) Codes
Polarization-adjusted convolutional (PAC) codes are modified polar codes in which a one-to-one convolutional transformation is employed before the classical polar transform. Fano decoding of PAC codes in the Shannon lecture at ISIT2019 showed an outstanding performance at the cost of a high time-complexity, particularly at low SNR regimes. In order to reduce this complexity, an adaptive heuristic metric is proposed that improves the comparability of the variable-length paths and adjusts itself in response to the channel noise level. This metric can significantly reduce the number of nodes visited on average in tree-traversal. Additionally, a partial rewinding of the successive cancellation process is proposed to efficiently compute the intermediate LLRs and partial sums when backtracking occurs in the Fano algorithm. This method avoids storing the intermediate results of the decoding process or restarting (full rewinding) the decoding process.
2020. International Symposium on Information Theory and its Applications (ISITA), ELECTR NETWORK, Oct 24-27, 2020. p. 200 – 204.GC-eDRAM with Body-Bias Compensated Readout and Error Detection in 28nm FD-SOI
Gain-cell embedded DRAM (GC-eDRAM) is an attractive alternative to conventional SRAM due to its high-density, low-leakage, and inherent two-ported functionality. However, its dynamic storage mechanism requires power-hungry refresh cycles to maintain data. This problem is aggravated due to the impact of Process-Voltage-Temperature (PVT) variations at deeply-scaled technology nodes and low voltages. In this paper, we present a GC-eDRAM with body-bias compensated readout, which is dynamically configured to extend the data retention time (DRT) of the memory under varying operating conditions. The proposed GC-eDRAM exploits the body-biasing capabilities of FD-SOI technology to adjust the switching threshold of the sense inverter under PVT variations. An additional, unbiased, sense inverter is added to provide a dual sampling mechanism to the readout path, enabling error detection to further reduce design guard bands. An 8 kb GC-eDRAM with integrated body-bias compensated readout and error detection was implemented in 28 nm FD-SOI technology. Silicon measurements of the manufactured array demonstrate up-to 75% DRT improvement and up-to 86% energy savings under PVT and frequency variations compared to a conventional guard banded memory design.
2020. IEEE International Symposium on Circuits and Systems (ISCAS), ELECTR NETWORK, Oct 10-21, 2020. DOI : 10.1109/ISCAS45731.2020.9180997.Gain-Cell Embedded DRAMs: Modeling and Design Space
Among the different types of DRAMs, gain-cell embedded DRAM (GC-eDRAM) is a compact, low-power and CMOS-compatible alternative to conventional SRAM. GC-eDRAM achieves high memory density as it relies on a storage cell that can be implemented with as few as two transistors and that can be fabricated without additional process steps. However, since the performance of GC-eDRAMs relies on many interdependent variables, the optimization of the performance of these memories for the integration into their hosting system, as well as the design investigation of future GC-eDRAMs, prove to be highly complex tasks. In this context, modeling tools of memories are key enablers for the exploration of this large design space in a short amount of time. In this paper, we present GEMTOO, the first modeling tool that estimates timing, memory availability, bandwidth, and area of GC-eDRAMs. The tool considers parameters related to technology, circuits, and memory architecture and it enables the evaluation of architectural transformations as well as of advanced transistor-level effects, such as the increase of the access delay due to deterioration of the stored data. The timing is estimated with a maximum deviation of 15% from post-layout simulations in a 28nm FD-SOI technology for different memory sizes and architectures. Moreover, the measured random cycle frequency of a GC-eDRAM fabricated in 28nm CMOS bulk process is estimated with a 9% deviation when considering 6-sigma random process variations of the bitcells. The proposed GEMTOO modeling tool is used to show the intricacies in design optimization of GC-eDRAMs and, based on the results, optimal design practices are derived.
2020. IEEE International Symposium on Circuits and Systems (ISCAS), ELECTR NETWORK, Oct 10-21, 2020. DOI : 10.1109/ISCAS45731.2020.9180999.A Maximum-Likelihood-based Multi-User LoRa Receiver Implemented in GNU Radio
LoRa is a popular low-power wide-area network (LPWAN) technology that uses spread-spectrum to achieve long-range connectivity and resilience to noise and interference. For energy efficiency reasons, LoRa adopts a pure ALOHA access scheme, which leads to reduced network throughput due to packet collisions at the gateways. To alleviate this issue, in this paper we analyze and implement a LoRa receiver that is able to decode LoRa packets from two interfering users. Our main contribution is a two-user detector derived in a maximum-likelihood fashion using a detailed interference model. As the complexity of the maximum-likelihood sequence estimation is prohibitive, a complexity-reduction technique is introduced to enable a practical implementation of the proposed two-user detector. This detector has been implemented along with an interference-robust synchronization algorithm on the GNU Radio software-defined radio (SDR) platform. The SDR implementation shows the effectiveness of the proposed method and also allows its experimental evaluation. Measurements indicate that our detector inherently leverages the time offset between the two colliding users to separate and demodulate their signals.
2020. 54th Asilomar Conference on Signals, Systems, and Computers, ELECTR NETWORK, Nov 01-05, 2020. p. 1106 – 1111. DOI : 10.1109/IEEECONF51394.2020.9443502.On the Implementation Complexity of Digital Full-Duplex Self-Interference Cancellation
In-band full-duplex systems promise to further increase the throughput of wireless systems, by simultaneously transmitting and receiving on the same frequency band. However, concurrent transmission generates a strong self-interference signal at the receiver, which requires the use of cancellation techniques. A wide range of techniques for analog and digital self-interference cancellation have already been presented in the literature. However, their evaluation focuses on cases where the underlying physical parameters of the full-duplex system do not vary significantly. In this paper, we focus on adaptive digital cancellation, motivated by the fact that physical systems change over time. We examine some of the different cancellation methods in terms of their performance and implementation complexity, considering the cost of both cancellation and training. We then present a comparative analysis of all these methods to determine which perform better under different system performance requirements. We demonstrate that with a neural network approach, the reduction in arithmetic complexity for the same cancellation performance relative to a state-of-the-art polynomial model is several orders of magnitude.
2020. 54th Asilomar Conference on Signals, Systems, and Computers, ELECTR NETWORK, Nov 01-05, 2020. p. 969 – 973. DOI : 10.1109/IEEECONF51394.2020.9443274.Theses
Energy- and Cost-Efficient VLSI DSP Systems Design with Approximate Computing
Technology scaling has progressed to enable integrated circuits with extremely high density enabling systems of tremendous complexity with manageable power consumption. With the continuation of Moore’s law for many years, electronic chips have been able to accommodate the growing performance demand and energy-efficiency requirement of many applications. Unfortunately, in the past decade, we have also seen diminishing returns from the most advanced process nodes in terms of performance and power consumption. The issue lies in the fact that it has become increasingly difficult to guarantee a reliable operation without costly design margins due to process variations. The diminishing benefit from technology scaling is mainly due to the worst-case design paradigm which requires a conservative design margin. It has been therefore proposed to abandon the conservative and 100% error-free design paradigm while exploiting the inherent fault-tolerance of many applications through approximate computing to avoid the need for conservative margins. The contributions of this thesis focus on algorithmic and architectural techniques for the design of approximate and efficient hardware and are summarized as follows. First, we propose a design methodology that allows us to drop the requirement of 100% reliable operation and to accept dies with unreliable memory components. We examine the proposed methodology for multiple applications such as image processing kernels and a channel decoder in communication systems and we show the first measured example of an integrated circuit that delivers a stable performance despite the presence of errors in its memories. Second, we propose a systematic statistical framework to dynamically adapt the output quality of a channel decoder at run-time, as an example of an iterative algorithm, which provides a notable reduction in the energy consumption for different approximation levels. Finally, we propose algorithmic and architectural techniques at design-time, inspired by static approximate computing paradigm, to reduce the complexity of the arithmetic units of a channel decoder resulting in a significant reduction in both design area and energy consumption while still achieving the best-in-class throughput.
Lausanne, EPFL, 2020.Physical Layer Aspects of LoRa and Full-Duplex Wireless Transceivers
Wireless communications are currently faced with two main challenges. The first challenge stems from the enormous number of Internet of Things (IoT) devices that transmit very small amounts of data. The second challenge is the need for ever-increasing data rates required by users of multimedia rich services, as well as the extremely low latency required in emerging applications such as autonomous vehicles and augmented reality. In this thesis we deal with important physical layer (PHY) aspects that have not been analyzed in-depth in the existing literature, and whose study can help to address the aforementioned challenges. Low-power wide-area networks (LPWANs) comprise a big part of the IoT. For energy efficiency reasons, most of LPWAN technologies adopt uncoordinated channel access schemes which result in collisions. This issue becomes more severe as the number of devices increases, putting the scalability of LPWANs at risk as they become interference-limited. To evaluate and support LPWAN scalability, in the first part of this thesis we perform a thorough analysis of the performance of one of the most important LPWAN technologies, namely LoRa. We analyze the LoRa performance in interference scenarios, and we derive expressions, as well as very accurate low-complexity approximations, for the error rate of LoRa for both the uncoded and coded cases, and with carrier frequency offset (CFO). We also propose and analyze the coherent demodulation of LoRa under interference, as a potential receiver improvement in collision scenarios. Finally, we build a standard-compatible LoRa PHY software-defined radio (SDR) prototype based on GNU Radio, which can be used for measurements of LoRa PHY performance. The second part of this thesis focuses on full-duplex radios, which allow simultaneous transmission and reception in the same frequency band, and have been proposed as a possible solution to overcome the capacity bottleneck of high data-rate applications. However, full-duplex transceivers suffer from strong self-interference. Perfect self-interference cancellation is difficult to achieve due to the presence of strong non-linear signal components, which are introduced by hardware imperfections inherent in the transmitter and receiver chains. We propose the digital predistortion of the transmit signal to compensate for the cascade of the transceiver non-linearities and enhance self-interference cancellation. Unfortunately, a residual self-interference component always remains, particularly when operating at realistic transmit powers. To increase the usefulness of full-duplex technology, we examine communication schemes where using full-duplex transceivers can significantly improve the performance in terms of both throughput and latency, even under imperfect self-interference suppression. In particular, we examine the use of full-duplex technology in cognitive radios, and in communication links with asymmetric capacity requirements between the uplink and downlink channels.
Lausanne, EPFL, 2020.2019
Journal Articles
GC-eDRAM With Body-Bias Compensated Readout and Error Detection in 28-nm FD-SOI
Gain-cell embedded DRAM (GC-eDRAM) is an attractive alternative to conventional SRAM due to its high-density, low-leakage, and inherent two-ported functionality. However, its dynamic storage mechanism requires power-hungry refresh cycles to maintain data. This problem is aggravated due to the impact of process-voltage-temperature (PVT) variations at deeply scaled technology nodes and low voltages. In this brief, we present a gain-cell embedded DRAM (GC-eDRAM) with body-bias compensated readout, which is dynamically configured to extend the data retention time (DRT) of the memory under varying operating conditions. The proposed GC-eDRAM exploits the body-biasing capabilities of FD-SOI technology to adjust the switching threshold of the sense inverter under PVT variations. An additional, unbiased, sense inverter is added to provide a dual-sampling mechanism to the readout path, enabling error detection to further reduce design guard bands. An 8-kb GC-eDRAM with integrated body-bias compensated readout and error detection was implemented in 28-nm FD-SOI technology. Silicon measurements of the manufactured array demonstrate up to 75 DRT improvement and up to 86 energy savings under PVT and frequency variations compared to a conventional guard banded memory design.
IEEE Transactions on Circuits and Systems II: Express Briefs. 2019. Vol. 66, num. 12, p. 2042 – 2046. DOI : 10.1109/TCSII.2019.2896164.2019 International Symposium on Low Power Electronics and Design
Ieee Design & Test. 2019. Vol. 36, num. 6, p. 82 – 83. DOI : 10.1109/MDAT.2019.2941713.Improving Energy-Efficiency in Dynamic Memories Through Retention Failure Detection
A gain-cell embedded DRAM (GC-eDRAM) is an attractive logic-compatible alternative to the conventional static random access memory (SRAM) for the implementation of embedded memories, as it offers higher density, lower leakage, and two-ported operation. However, it requires periodic refresh cycles to maintain its data which deteriorates due to leakage. The refresh-rate, which is traditionally set according to the worst cell in the array under extreme operating conditions, leads to a significant refresh power consumption and decreased memory availability. In this paper, we propose to reduce the cost of GC-eDRAM refresh by employing failure detection to lower the refresh-rate. A 4T dynamic complementary dual-modular redundancy bitcell is proposed to offer per-bit error detection, resulting in a substantial decrease in the refresh-rate and over 60% power reduction compared with the SRAM. The proposed approach is also compared with the conventional SRAM and GCeDRAM implementations with integrated error correction codes, demonstrating significant area and latency reductions.
Ieee Access. 2019. Vol. 7, p. 27641 – 27649. DOI : 10.1109/ACCESS.2019.2901738.Conference Papers
Scalable Boolean Methods in a Modern Synthesis Flow
With the continuous push to improve Quality of Results (QoR) in EDA, Boolean methods in logic synthesis have been recently drawing the attention of researchers. Boolean methods achieve better QoR than algebraic methods but require higher computational cost. In this paper, we introduce the Scalable Boolean Method (SBM) framework. The SBM consists of 4 optimization engines designed to be scalable in a modern synthesis flow. The first presented engine is a generalized resubstitution framework based on computing, and implementing, the Boolean difference between two nodes. The second consists of a gradient-based AIG optimization, while the third one is based on heterogeneous elimination for kerneling. The last proposed engine is a revisiting of maximum set of permissible functions computation with BDDs. Altogether, the SBM framework enables significant synthesis results. We improve 12 of the best known area results in the EPFL synthesis competition. Embedded in a commercial EDA flow, the new Boolean methods enable -2.20% combinational area savings and -5.99% total negative slack reduction, after physical implementation, at contained runtime cost.
2019. Design, Automation & Test in Europe Conference & Exhibition (DATE), Florence, ITALY, Mar 25-29, 2019. p. 1643 – 1648. DOI : 10.23919/DATE.2019.8714776.Neural-Network Optimized 1-bit Precoding for Massive MU-MIMO
Base station (BS) architectures for massive multi-user (MU) multiple-input multiple-output (MIMO) wireless systems are equipped with hundreds of antennas to serve tens of users on the same time-frequency channel. The immense number of BS antennas incurs high system costs, power, and interconnect bandwidth. To circumvent these obstacles, sophisticated MU precoding algorithms that enable the use of 1-bit DACs have been proposed. Many of these precoders feature parameters that are, traditionally, tuned manually to optimize their performance. We propose to use deep-learning tools to automatically tune such 1-bit precoders. Specifically, we optimize the biConvex 1-bit PrecOding (C2PO) algorithm using neural networks. Compared to the original C2PO algorithm, our neural-network optimized (NNO-)C2PO achieves the same error-rate performance at 2x lower complexity. Moreover, by training NNO-C2PO for different channel models, we show that 1-bit precoding can be made robust to vastly changing propagation conditions.
2019. 20th IEEE International Workshop on Signal Processing Advances in Wireless Communications (SPAWC), Cannes, FRANCE, Jul 02-05, 2019. DOI : 10.1109/SPAWC.2019.8815519.LDPC Coded Multiuser Shaping for the Gaussian Multiple Access Channel
The joint design of input constellation and low-density parity-check (LDPC) codes to approach the symmetric capacity of the two-user Gaussian multiple access channel is studied. More specifically, multilevel coding is employed at each user to construct a high-order input constellation and the constellations of the users are jointly designed so as to maximize the multiuser shaping gain. At the receiver, each layer of the multilevel coding is jointly decoded among users, while successive cancellation is employed across layers. The LDPC code employed by each user in each layer is designed using EXIT charts to support joint decoding among users for the prescribed per-layer rate and SNR. Numerical simulations are provided to validate the proposed constellation and LDPC code designs.
2019. IEEE International Symposium on Information Theory (ISIT), Paris, FRANCE, Jul 07-12, 2019. p. 2609 – 2613. DOI : 10.1109/ISIT.2019.8849785.A 24 kb Single-Well Mixed 3T Gain-Cell eDRAM with Body-Bias in 28 nm FD-SOI for Refresh-Free DSP Applications
Logic-compatible gain-cell embedded DRAM (GC-eDRAM) is an emerging alternative to conventional SRAM for memory-dominated system-on-chip (SoC) designs due to its high-density, low-power, and two-ported operation. Although GCs have a limited data retention time (DRT) at deeply scaled technology nodes, there are many DSP applications which only require short-term data storage and can therefore avoid refresh. In this paper, we present a novel single-well mixed 3T GC implementation in 28 nm FD-SOI technology. The proposed GC is supplied with body-bias control to improve the DRT by suppressing the leakage through the write port, and extend the maximum operating frequency by forward body-biasing the read port. A 24 kbit GC-eDRAM macro implementing the proposed 3T GC was fabricated in 28 nm FD-SOI technology, resulting in the highest density logic-compatible embedded memory fabricated in any 28 nm process with over 2x higher density compared to a 6T SRAM cell, over 4x higher DRT compared to a conventional 3T GC, and 38 x 47 x lower static power compared to conventional single-ported and two-ported SRAMs.
2019. 15th IEEE Asian Solid-State Circuits Conference (A-SSCC), Macao, PEOPLES R CHINA, Nov 04-06, 2019. p. 219 – 222. DOI : 10.1109/A-SSCC47793.2019.9056985.LoRa Symbol Error Rate Under Non-Aligned Interference
In this work, we examine the performance of the LoRa chirp spread spectrum modulation in the presence of both additive white Gaussian noise and interference from another LoRa user. To this end, we extend an existing interference model to the more realistic case where the interfering user is neither chip- nor phase-aligned with the signal of interest and we derive an expression for the SER. We show that the existing interference model overestimates the effect of interference on the error rate. Moreover, we derive a low-complexity approximate formula that can significantly reduce the complexity of computing expression.
2019. 53rd Asilomar Conference on Signals, Systems, and Computers, Pacific Grove, CA, Nov 03-06, 2019. p. 1957 – 1961. DOI : 10.1109/IEEECONF44664.2019.9048665.A 4.8pJ/b 6Gb/s ADC-Based PAM-4 Wireline Receiver Data -Path with Cyclic Prefix in 14nm FinFET
This work presents an ADC-bascd receiver (RX) data-path for frame-based PAM-4 modulation with a cyclic prefix (CP). Similar to discrete multi-tone (DMT) modulation, a frame of PAM-4 symbols arc protected from the channel delay spread by the CP taps. A PAM-4 frame window including CP taps is viewed as a DMT symbol and is equalized similarly to a DMT signal equalization, based on a discrete-time Fourier transform (DFT) and frequency-domain equalizer (FDE). The RX prototype implemented in 14nm FinFET achieves 56Gb/s datarate at less than 3c-5 pre-FEC BER over a 19dB loss channel at 14GHz dissipating 270mW including the ADC and the DSP data-path excluding the inverse DFT and the BER checker.
2019. 15th IEEE Asian Solid-State Circuits Conference (A-SSCC), Macao, PEOPLES R CHINA, Nov 04-06, 2019. p. 239 – 240. DOI : 10.1109/A-SSCC47793.2019.9056940.Feedback-Aware Precoding for Millimeter Wave Massive MIMO Systems
Millimeter wave (mmWave) communication is a promising solution for coping with the ever-increasing mobile data traffic because of its large bandwidth. To enable a sufficient link margin, a large antenna array employing directional beamforming, which is enabled by the availability of channel state information at the transmitter (CSIT), is required. However, CSIT acquisition for mmWave channels introduces a huge feedback overhead due to the typically large number of transmit and receive antennas. Leveraging properties of mmWave channels, this paper proposes a precoding strategy which enables a flexible adjustment of the feedback overhead. In particular, the optimal unconstrained precoder is approximated by selecting a variable number of elements from a basis that is constructed as a function of the transmitter array response, where the number of selected basis elements can be chosen according to the feedback constraint. Simulation results show that the proposed precoding scheme can provide a near-optimal solution if a higher feedback overhead can be afforded. For a low overhead, it can still provide a good approximation of the optimal precoder.
2019. 30th IEEE Annual International Symposium on Personal, Indoor and Mobile Radio Communications (PIMRC), Istanbul, TURKEY, Sep 08-11, 2019. p. 1021 – 1027. DOI : 10.1109/PIMRC.2019.8904332.Minimum Energy Point in Constant Frequency Designs under Adaptive Supply Voltage and Body Bias Adjustment in 55 nm DDC
In this paper, we describe a systematic low-power design methodology for technologies that offer a strong body factor. Specifically, we explore both the body bias voltage and the supply voltage knobs in order to find the MEP (minimum energy point) for a constant target frequency. Our methodology accounts for process and temperature (PT) variations while charting the design space for a simple reference design. We then show how to scale the energy data of this reference design to any arbitrary design. A case study of a 32 bit RISC microprocessor achieves an energy estimation match of our significantly less complex estimation methodology within 1% of traditional signoff results.
2019. 15th Conference on Ph.D Research in Microelectronics and Electronics (PRIME), Lausanne, SWITZERLAND, Jul 15-18, 2019. p. 285 – 288. DOI : 10.1109/PRIME.2019.8787736.Data-Retention-Time Characterization of Gain-Cell eDRAMs across the Design and Variations Space
The rise of data-intensive applications has increased the demand for high-density and low-power embedded memories. Among them, the gain-cell embedded DRAM (GC-eDRAM) is a suitable alternative to the static random access memory (SRAM) due to its high memory density and low leakage current. However, as the GC-eDRAM dynamically stores data, its memory content has to be periodically refreshed according to the data retention time (DRT). Even though different DRT characterization methodologies have been reported in the literature, a practical and accurate method to quantify the DRT across Monte Carlo (MC) runs to evaluate the impact of local process variations (LPVs) has not been proposed yet. Thus, the minimum memory refresh rate is generally estimated with large design guard bands to avoid any loss of data, at the expense of a higher power consumption and less memory bandwidth. In this work, we present a current-based DRT characterization methodology that enables an accurate LPV analysis without the need of a large number of costly electronic design automation (EDA) software licenses. The presented approach is compared with other DRT characterization methodologies for both accuracy as well as practical aspects. Furthermore, the DRT of a 3-transistor (3T) gain cell (GC) designed in 28nm FD-SOI process technology is measured for different design choices, global and local variations. The analysis of the results shows that LPVs have the most degrading effect on the DRT and therefore that the proposed approach is key for either the design of GC-eDRAMs or the choice of their refresh rate to avoid the need for overly pessimistic worst-case margins.
2019. IEEE International Symposium on Circuits and Systems (IEEE ISCAS), Sapporo, JAPAN, May 26-29, 2019. DOI : 10.1109/ISCAS.2019.8702393.Power Analysis Resilient SRAM Design Implemented with a 1% Area Overhead Impedance Randomization Unit for Security Applications
Power analysis attacks are an effective tool to extract sensitive information using side-channel analysis, forming a serious threat to IoT systems-on-a-chip (SoCs). Embedded memories implemented with conventional 6T SRAM macrocells often dominate the area and power of these SoCs. In this paper, for the first time, we use silicon measurements to prove that conventional SRAM arrays leak valuable information and that their data can be extracted using power analysis attacks. In order to provide a power analysis resilient embedded memory and adhere to the area constraints of modern SoCs, we implement a low-cost impedance randomization unit, which is integrated into the periphery of a conventional 6T SRAM macro. Preliminary silicon measurements of a 55 nm test-chip implementing the proposed memory array demonstrate a significant information leakage reduction at a low-cost 1% area overhead and no speed and power penalties compared to a conventional SRAM design.
2019. IEEE 45th European Solid State Circuits Conference (ESSCIRC), Cracow, POLAND, Sep 23-26, 2019. p. 69 – 72. DOI : 10.1109/ESSCIRC.2019.8902622.A 0.5 V 2.5 mu W/MHz Microcontroller with Analog-Assisted Adaptive Body Bias PVT Compensation with 3.13 nW/kB SRAM Retention in 55 nm Deeply-Depleted Channel CMOS
Microcontroller systems operating at low supply voltage in near- or sub-threshold regime suffer both from increased effects of PVT (Process, Voltage, Temperature) variation and from a larger share of leakage on overall power due to the reduced frequency. We show how to overcome these effects for the core and memory by exploiting the strong body factor of deeply-depleted channel CMOS at 0.5 V, compensating frequency over PVT to +/- 6%, achieving 30x frequency and 20x leakage scaling in a 2.56 mu W/MHz 32 bit RISC Core with 3.13 nW/kB 2.5 mu W/MHz SRAM. Frequency-leakage configurability in core and SRAM through adaptive body bias at fixed supply voltage is implemented using a novel automatic analog-assisted I-ON-controlled approach.
2019. 40th Annual IEEE Custom Integrated Circuits Conference (CICC), Austin, TX, Apr 14-17, 2019. DOI : 10.1109/CICC.2019.8780199.3.5 GHz Coverage Assessment with a 5G Testbed
Today, cellular networks have saturated frequencies below 3 GHz. Because of increasing capacity requirements, 5th generation (5G) mobile networks target the 3.5 GHz band (3.4 to 3.8 GHz). Despite its expected wide usage, there is little empirical path loss data and mobile radio network planning experience for the 3.5 GHz band available. This paper presents the results of rural, suburban, and urban measurement campaigns using a pre-standard 5G prototype testbed operating at 3.5 GHz, with outdoor as well as outdoor-to-indoor scenarios. Based on the measurement results, path loss models are evaluated, which are essential for network planning.
2019. 89th IEEE Vehicular Technology Conference (VTC Spring), Kuala Lumpur, MALAYSIA, Apr 28-May 01, 2019. DOI : 10.1109/VTCSpring.2019.8746551.A 161mW 56Gb/s ADC-Based Discrete Multitone Wireline Receiver Data-Path in 14nm FinFET
2019. IEEE International Solid- State Circuits Conference (ISSCC), San Francisco, CA, Feb 17-21, 2019. p. 476 – 478. DOI : 10.1109/ISSCC.2019.8662505.FPGA-Based Emulation of Embedded DRAMs for Statistical Error Resilience Evaluation of Approximate Computing Systems
Embedded DRAM (eDRAM) requires frequent power-hungry refresh according to the worst-case retention time across PVT variations to avoid data loss. Abandoning the error-free paradigm, by choosing sub-critical refresh rates that gracefully degrade the eDRAM content, unlocks considerable power-saving opportunities, but requires to understand the effect of stochastic memory errors at the system/application level. We propose an FPGA-based platform featuring faulty eDRAM emulation based on advanced retention time models and silicon measurements for statistical error resilience evaluation of applications in a complete embedded system. We analyze the statistical QoS for various benchmarks under different sub-critical refresh rates and retention time distributions.
2019. 56th ACM/EDAC/IEEE Design Automation Conference (DAC), Las Vegas, NV, Jun 02-06, 2019. DOI : 10.1145/3316781.3317830.A Lyra2 FPGA Core for Lyra2REv2-Based Cryptocurrencies
Lyra2REv2 is a hashing algorithm that consists of a chain of individual hashing algorithms and it is used as a proof-of-work function in several cryptocurrencies that aim to be ASIC-resistant. The most crucial hashing algorithm in the Lyra2REv2 chain is a specific instance of the general Lyra2 algorithm. In this work we present the first FPGA implementation of the aforementioned instance of Lyra2 and we explain how several properties of the algorithm can be exploited in order to optimize the design.
2019. IEEE International Symposium on Circuits and Systems (IEEE ISCAS), Sapporo, JAPAN, May 26-29, 2019. DOI : 10.1109/ISCAS.2019.8702498.Lora Digital Receiver Analysis And Implementation
Low power wide area network technologies (LPWANs) are attracting attention because they fulfill the need for long range low power communication for the Internet of Things. LoRa is one of the proprietary LPWAN physical layer (PHY) technologies, which provides variable data-rate and long range by using chirp spread spectrum modulation. This paper describes the basic LoRa PHY receiver algorithms and studies their performance. The LoRa PHY is first introduced and different demodulation schemes are proposed. The effect of carrier frequency offset and sampling frequency offset are then modeled and corresponding compensation methods are proposed. Finally, a software-defined radio implementation for the LoRa transceiver is briefly presented.
2019. 44th IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Brighton, ENGLAND, May 12-17, 2019. p. 1498 – 1502. DOI : 10.1109/ICASSP.2019.8683504.Advanced Machine Learning Techniques for Self-Interference Cancellation in Full-Duplex Radios
In-band full-duplex systems allow for more efficient use of temporal and spectral resources by transmitting and receiving information at the same time and on the same frequency. However, this creates a strong self-interference signal at the receiver, making the use of self-interference cancellation critical. Recently, neural networks have been used to perform digital self-interference with lower computational complexity compared to a traditional polynomial model. In this paper, we examine the use of advanced neural networks, such as recurrent and complex-valued neural networks, and we perform an in-depth network architecture exploration. Our neural network architecture exploration reveals that complex-valued neural networks can significantly reduce both the number of floating-point operations and parameters compared to a polynomial model, whereas the real-valued networks only reduce the number of floating-point operations. For example, at a digital self-interference cancellation of 44:51dB, a complex-valued neural network requires 33:7% fewer floating-point operations and 26:9% fewer parameters compared to the polynomial model.
2019. 53rd Asilomar Conference on Signals, Systems, and Computers, Pacific Grove, CA, Nov 03-06, 2019. p. 1149 – 1153. DOI : 10.1109/IEEECONF44664.2019.9048900.JESD204B Compliant 12.5 Gb/s LVDS and SST Transmitters in 28 nm FD-SOI CMOS
JESD204B compliant low-voltage differential signaling (LVDS), and source-series-terminated (SST) transmitters in 28 nm FD-SOI CMOS technology are presented with 1.1 pJ/bit and 1.7 pJ/bit at 12.5 Gb/s, respectively. An external 6.25 GHz single-ended clock, which is terminated internally with mid-common-mode termination, is used for half-rate operation in both designs, which can achieve open eye diagrams at 12.5 Gb/s data rate. Both transmitters are measured and compared in terms of design complexity, power consumption, area, and signal integrity performance. From architecture selection to circuit design, power consumption is minimized while maintaining the maximum data rate that the JESD204B standard supports. High-speed standard cell ESD diodes are employed for the pads to achieve >1kV HBM ESD protection, while adding 200 fF parasitic capacitance.
2019. 15th Conference on Ph.D Research in Microelectronics and Electronics (PRIME), Lausanne, SWITZERLAND, Jul 15-18, 2019. p. 101 – 104. DOI : 10.1109/PRIME.2019.8787786.Theses
Low-Power Design of Digital VLSI Circuits around the Point of First Failure
As an increase of intelligent and self-powered devices is forecasted for our future everyday life, the implementation of energy-autonomous devices that can wirelessly communicate data from sensors is crucial. Even though techniques such as voltage scaling proved to effectively reduce the energy consumption of digital circuits, additional energy savings are still required for a longer battery life. One of the main limitations of essentially any low-energy technique is the potential degradation of the quality of service (QoS). Thus, a thorough understanding of how circuits behave when operated around the point of first failure (PoFF) is key for the effective application of conventional energy-efficient methods as well as for the development of future low-energy techniques. In this thesis, a variety of circuits, techniques, and tools is described to reduce the energy consumption in digital systems when operated either in the safe and conservative exact region, close to the PoFF, or even inside the inexact region. A straightforward approach to reduce the power consumed by clock distribution while safely operating in the exact region is dual-edge-triggered (DET) clocking. However, the DET approach is rarely taken, primarily due to the perceived complexity of its integration. In this thesis, a fully automated design flow is introduced for applying DET clocking to a conventional single-edge-triggered (SET) design. In addition, the first static true-single-phase-clock DET flip-flop (DET-FF) that completely avoids clock-overlap hazards of DET registers is proposed. Even though the correct timing of synchronous circuits is ensured in worst-case conditions, the critical path might not always be excited. Thus, dynamic clock adjustment (DCA) has been proposed to trim any available dynamic timing margin by changing the operating clock frequency at runtime. This thesis describes a dynamically-adjustable clock generator (DCG) capable of modifying the period of the produced clock signal on a cycle-by-cycle basis that enables the DCA technique. In addition, a timing-monitoring sequential (TMS) that detects input transitions on either one of the clock phases to enable the selection of the best timing-monitoring strategy at runtime is proposed. Energy-quality scaling techniques aimat trading lower energy consumption for a small degradation on the QoS whenever approximations can be tolerated. In this thesis, a low-power methodology for the perturbation of baseline coefficients in reconfigurable finite impulse response (FIR) filters is proposed. The baseline coefficients are optimized to reduce the switching activity of the multipliers in the FIR filter, enabling the possibility of scaling the power consumption of the filter at runtime. The area as well as the leakage power of many system-on-chips is often dominated by embedded memories. Gain-cell embedded DRAM (GC-eDRAM) is a compact, low-power and CMOS-compatible alternative to the conventional static random-access memory (SRAM) when a higher memory density is desired. However, due to GC-eDRAMs relying on many interdependent variables, the adaptation of existing memories and the design of future GCeDRAMs prove to be highly complex tasks. Thus, the first modeling tool that estimates timing, memory availability, bandwidth, and area of GC-eDRAMs for a fast exploration of their design space is proposed in this thesis.
Lausanne, EPFL, 2019.2018
Journal Articles
Design of LDPC Codes for the Unequal Power Two-User Gaussian Multiple Access Channel
In this letter, we describe an LDPC code design framework for the unequal power two-user Gaussian multiple access channel using EXIT charts. We show that the sumrate of the LDPC codes designed using our approach can get close to the maximal sum-rate of the two-user Gaussian multiple access channel. Moreover, we provide numerical simulation results that demonstrate the excellent finite-length performance of the designed LDPC codes.
Ieee Wireless Communications Letters. 2018. Vol. 7, num. 5, p. 868 – 871. DOI : 10.1109/LWC.2018.2833855.Design Techniques for High-Speed Multi-Level Viterbi Detectors and Trellis-Coded-Modulation Decoders
The implementation of a 25.6-Gb/s four-level pulse-amplitude-modulation (4-PAM) reduced-state sliding-block Viterbi detector (VD) is presented. The power consumption of the VD is 105 orilV at a supply voltage of 0.7 V, corresponding to an energy efficiency of 4.1 pJ/b. A data rate of 30.4 Gb/s is achieved with an energy efficiency of 5.3 pith at a supply voltage of 0.8 V. The VD, implemented in an experimental chip fabricated in 14-nm CMOS FINFET, exploits set-partitioning principles and embedded per-survivor decision feedback to reduce implementation complexity and power consumption. The active area of the VD with 12 slices, each operating at one-eighth of the modulation rate, is 0.507 x 0.717 mm(2). Experimental results showing system performance are obtained by using a (2(15)-1)-bit pseudo-random binary sequence. The impact of the synchronization length and survivor path memory length on the detector design and system performance are shown. A new pipelined reduced-state sequence detector algorithm is presented for high-speed implementations. A novel speculative symbol timing recovery scheme is proposed. New simulation results are obtained to compare the performance of the Reed-Solomon (RS)-encoded 4-PAM scheme with that of the concatenated RS 4-D 5-PAM trellis-coded-modulation (TCM) scheme over an ideal band-limited additive-white-Gaussian-noise channel. Drawing on the results achieved for the VD, novel design techniques for a high-speed low-complexity eight-state 4-D 5-PAM TCM decoder is proposed.
Ieee Transactions On Circuits And Systems I-Regular Papers. 2018. Vol. 65, num. 10, p. 3529 – 3542. DOI : 10.1109/TCSI.2018.2803735.A 9.52 dB NCG FEC Scheme and 162 b/Cycle Low-Complexity Product Decoder Architecture
IEEE TRANSACTIONS ON CIRCUITS AND SYSTEMS I-REGULAR PAPERS. 2018. Vol. 65, num. 4, p. 1420 – 1431. DOI : 10.1109/TCSI.2017.2745902.Fast Low-Complexity Decoders for Low-Rate Polar Codes
Polar codes are capacity-achieving error-correcting codes with an explicit construction that can be decoded with low-complexity algorithms. In this work, we show how the state-of-the-art low-complexity decoding algorithm can be improved to better accommodate low-rate codes. More constituent codes are recognized in the updated algorithm and dedicated hardware is added to efficiently decode these new constituent codes. We also alter the polar code construction to further decrease the latency and increase the throughput with little to no noticeable effect on error-correction performance. Rate-flexible decoders for polar codes of length 1024 and 2048 are implemented on FPGA. Over the previous work, they are shown to have from 22% to 28% lower latency and 26% to 34% greater throughput when decoding low-rate codes. On 65 nm ASIC CMOS technology, the proposed decoder for a (1024, 512) polar code is shown to compare favorably against the state-of-the-art ASIC decoders. With a clock frequency of 400 MHz and a supply voltage of 0.8 V, it has a latency of 0.41 μs and an area efficiency of 1.8 Gbps/mm2 for an energy efficiency of 77 pJ/info. bit. At 600 MHz with a supply of 1 V, the latency is reduced to 0.27 μs and the area efficiency increased to 2.7 Gbps/mm2 at 115 pJ/info. bit.
Journal of Signal Processing Systems. 2018. Vol. 90, num. 5, p. 675 – 685. DOI : 10.1007/s11265-016-1173-y.A 4-Transistor nMOS-Only Logic-Compatible Gain-Cell Embedded DRAM With Over 1.6-ms Retention Time at 700 mV in 28-nm FD-SOI
IEEE TRANSACTIONS ON CIRCUITS AND SYSTEMS I-REGULAR PAPERS. 2018. Vol. 65, num. 4, p. 1245 – 1256. DOI : 10.1109/TCSI.2017.2747087.An 800-MHz Mixed-V-T 4T IFGC Embedded DRAM in 28-nm CMOS Bulk Process for Approximate Storage Applications
IEEE Journal of Solid-State Circuits. 2018. Vol. 53, num. 7, p. 2136 – 2148. DOI : 10.1109/JSSC.2018.2820145.Wireless Communication and Security Issues for Cyber-Physical Systems and the Internet-of-Things
Wireless sensors and actuators connected by the Internet-of-Things (IoT) are central to the design of advanced cyber-physical systems (CPSs). In such complex, heterogeneous systems, communication links must meet stringent requirements on throughput, latency, and range, while adhering to tight energy budget and providing high levels of security. In this paper, we first summarize wireless communication principles from the perspective of the connectivity needs of IoT and CPS. Based on these principles, we then review the most relevant wireless communication standards before focusing on the key security issues and features of such systems. In particular, the gap between the security features in the communication standards used in CPSs and IoT and their actual vulnerabilities are pointed out with practical examples and recent attacks. We emphasize the need for a more in-depth study of the security issues across all the protocol layers, including both logical layer security and physical layer security.
Proceedings Of The IEEE. 2018. Vol. 106, num. 1, p. 38 – 60. DOI : 10.1109/Jproc.2017.2780172.A 588-Gb/s LDPC Decoder Based on Finite-Alphabet Message Passing
IEEE TRANSACTIONS ON VERY LARGE SCALE INTEGRATION (VLSI) SYSTEMS. 2018. Vol. 26, num. 2, p. 329 – 340. DOI : 10.1109/TVLSI.2017.2766925.Faulty Successive Cancellation Decoding of Polar Codes for the Binary Erasure Channel
IEEE TRANSACTIONS ON COMMUNICATIONS. 2018. Vol. 66, num. 6, p. 2322 – 2332. DOI : 10.1109/TCOMM.2017.2771243.Conference Papers
Fast-SSC-Flip Decoding of Polar Codes
Polar codes are widely considered as one of the most exciting recent discoveries in channel coding. For short to moderate block lengths, their error-correction performance under list decoding can outperform that of other modern error-correcting codes. However, high-speed list-based decoders with moderate complexity are challenging to implement. Successive-cancellation (SC)-flip decoding was shown to be capable of a competitive error-correction performance compared to that of list decoding with a small list size, at a fraction of the complexity, but suffers from a variable execution time and a higher worst-case latency. In this work, we show how to modify the state-of-the-art high-speed SC decoding algorithm to incorporate the SC-flip ideas. The algorithmic improvements are presented as well as average execution-time results tailored to a hardware implementation. The results show that the proposed fast-SSC-flip algorithm has a decoding speed close to an order of magnitude better than the previous works while retaining a comparable error-correction performance.
2018. DOI : 10.1109/WCNCW.2018.8369026.A Timing-Monitoring Sequential for Forward and Backward Error-Detection in 28 nm FD-SOI
The increasing impact of variability on near-threshold nanometer circuits calls for a tighter online monitoring and control of the available timing margins. Error-detection sequentials are widely used together with error-correction techniques to operate digital designs with such carefully controlled far-below-worst-case margins, ensuring their correct operation even in the presence of uncertainties and variations. However, these registers are often designed only to either detect setup timing violations or to measure the available positive timing slack for a small detection-window. In this paper we propose a timing-monitoring sequential that provides both timing-monitoring modes, which can be selected at run-time depending on the desired timing-monitoring strategy. As the detection window of the presented circuit depends on the duty-cycle of the clock, either slow paths or fast paths can be monitored and measured with wide timing windows. The performance of this timing-monitoring sequential is evaluated in a 28nm FD-SOI process with post-layout simulations which show that the circuit is able to monitor a positive timing slack as small as 140 ps or to measure a path delay as fast as 50 ps. The proposed circuit is applied to a digital multiplier that was fabricated in a test chip and measurements show that the timing-monitoring sequentials are able to measure the critical path of the multiplier with a 1% accuracy and without incurring any timing violation.
2018. IEEE International Symposium on Circuits and Systems (ISCAS), Florence, ITALY, May 27-30, 2018. DOI : 10.1109/ISCAS.2018.8351043.Non-Linear Digital Self-Interference Cancellation for In-Band Full-Duplex Radios Using Neural Networks
Full-duplex systems require very strong self-interference cancellation in order to operate correctly and a significant part of the self-interference signal is due to non-linear effects created by various transceiver impairments. As such, linear cancellation alone is usually not sufficient and sophisticated non-linear cancellation algorithms have been proposed in the literature. In this work, we investigate the use of a neural network as an alternative to the traditional non-linear cancellation method that is based on polynomial basis functions. Measurement results from a full-duplex testbed demonstrate that a small and simple feed-forward neural network canceler works exceptionally well, as it can match the performance of the polynomial non-linear canceler with significantly lower computational complexity.
2018. IEEE 19th International Workshop on Signal Processing Advances in Wireless Communications (SPAWC), Kalamata, GREECE, Jun 25-28, 2018. p. 1 – 5. DOI : 10.1109/SPAWC.2018.8445987.Majority Logic Synthesis
The majority function (xyz) evaluates to true, if at least two of its Boolean inputs evaluate to true. The majority function has frequently been studied as a central primitive in logic synthesis applications for many decades. Knuth refers to the majority function in the last volume of his seminal The Art of Computer Programming as “probably the most important ternary operation in the entire universe.” Majority logic sythesis has recently regained signficant interest in the design automation community due to nanoemerging technologies which operate based on the majority function. In addition, majority logic synthesis has successfully been employed in CMOS-based applications such as standard cell or FPGA mapping. This tutorial gives a broad introduction into the field of majority logic synthesis. It will review fundamental results and and describe recent contributions from theory, practice, and applications.
2018. 37th IEEE/ACM International Conference on Computer-Aided Design (ICCAD), San Diego, CA, Nov 05-08, 2018. DOI : 10.1145/3240765.3267501.On the Tradeoff Between Accuracy and Complexity in Blind Detection of Polar Codes
Polar codes are a recent family of error-correcting codes with a number of desirable characteristics. Their disruptive nature is illustrated by their rapid adoption in the 5 th-generation mobile-communication standard, where they are used to protect control messages. In this work, we describe a two-stage system tasked with identifying the location of control messages that consists of a detection and selection stage followed by a decoding one. The first stage spurs the need for polar-code detection algorithms with variable effort to balance complexity between the two stages. We illustrate this idea of variable effort for multiple detection algorithms aimed at the first stage. We propose three novel blind detection methods based on belief-propagation decoding inspired by early-stopping criteria. Then we show how their reliability improves with the number of decoding iterations to highlight the possible tradeoffs between accuracy and complexity. Additionally, we show similar tradeoffs for a detection method from previous work. In a setup where only one block encoded with the polar code of interest is present among many other blocks, our results notably show that, depending on the complexity budget, a variable number of undesirable blocks can be dismissed while achieving a missed-detection rate in line with the block-error rate of a complex decoding algorithm.
2018. 10th IEEE International Symposium on Turbo Codes & Iterative Information Processing (ISTC), HONG KONG, PEOPLES R CHINA, Dec 03-07, 2018. DOI : 10.1109/ISTC.2018.8625366.Design and Implementation of a Neural Network Aided Self Interference Cancellation Scheme for Full-Duplex Radios
In-band full-duplex systems are able to transmit and receive information simultaneously on the same frequency band. Due to the strong self-interference caused by the transmitter to its own receiver, the use of non-linear digital self interference cancellation is essential. In this work, we present a hardware architecture for a neural network based non-linear self-interference canceller and we compare it with our own hardware implementation of a conventional polynomial based canceller. We show that, for the same cancellation performance, the neural network canceller has a significantly higher throughput and requires fewer hardware resources.
2018. 52nd Asilomar Conference on Signals, Systems, and Computers, Pacific Grove, CA, Oct 28-Nov 01, 2018. p. 589 – 593. DOI : 10.1109/ACSSC.2018.8645295.2017
Journal Articles
An FPGA-Based 4 Mbps Secret Key Distillation Engine for Quantum Key Distribution Systems
Quantum key distribution (QKD) enables provably secure communication between two parties over an optical fiber that arguably withstands any form of attack. Besides the need for a suitable physical signalling scheme and the corresponding devices, QKD also requires a secret key distillation protocol. This protocol and the involved signal processing handle the reliable key agreement process over the fragile quantum channel, as well as the necessary post-processing of key bits to avoid leakage of secret key information to an eavesdropper. In this paper we present in detail an implementation of a key distillation engine for a QKD system based on the coherent one-way (COW) protocol. The processing of key bits by the key distillation engine includes agreement on quantum bit detections (sifting), information reconciliation with forward error correction coding, parameter estimation, and privacy amplification over an authenticated channel. We detail the system architecture combining all these processing steps, and discuss the design trade-offs for each individual system module. We also assess the performance and efficiency of our key distillation implementation in terms of throughput, error correction capabilities, and resource utilization. On a single-FPGA ( Xilinx Virtex-6 LX240T) platform, the system supports distilled key rates of up to 4 Mbps.
Journal Of Signal Processing Systems For Signal Image And Video Technology. 2017. Vol. 86, num. 1, p. 1 – 15. DOI : 10.1007/s11265-015-1086-1.Energy-Efficient Near-Threshold Parallel Computing: The PULPv2 Cluster
This article presents an ultra-low-power parallel computing platform and its system-on-chip (SoC) embodiment, targeting a wide range of emerging near-sensor processing tasks for Internet of Things (IoT) applications. The proposed SoC achieves 193 million operations per second (MOPS) per mW at 162 MOPS (32 bits), improving the first-generation Parallel Ultra-Low-Power (PULP) architecture by 6.4 and 3.2 times in performance and energy efficiency, respectively.
Ieee Micro. 2017. Vol. 37, num. 5, p. 20 – 31. DOI : 10.1109/MM.2017.3711645.Automated Integration of Dual-Edge Clocking for Low-Power Operation in Nanometer Nodes
Clocking power, including both clock distribution and registers, has long been one of the primary factors in the total power consumption of many digital systems. One straightforward approach to reduce this power consumption is to apply dual-edge-triggered (DET) clocking, as sequential elements operate at half the clock frequency while maintaining the same throughput as with conventional single-edge-triggered (SET) clocking. However, the DET approach is rarely taken in modern integrated circuits, primarily due to the perceived complexity of integrating such a clocking scheme. In this article, we first identify the most promising conditions for achieving low-power operation with DET clocking and then introduce a fully automated design flow for applying DET to a conventional SET design. The proposed design flow is demonstrated on three benchmark circuits in a 40nm CMOS technology, providing as much as a 50% reduction in clock distribution and register power consumption.
ACM Transactions on Design Automation of Electronic Systems. 2017. Vol. 22, num. 4, p. 62. DOI : 10.1145/3054744.Multipliers-Driven Perturbation of Coefficients for Low-Power Operation in Reconfigurable FIR Filters
Reconfigurable finite-impulse response (FIR) filters are one of the most widely implemented components in Internet of Things systems that require flexibility to support several target applications while consuming the minimum amount of power to comply with the strict design requirements of portable devices. Due to the significant power consumption in the multiplier components of the FIR filter, various techniques aimed at reducing the switching activity of these multipliers have been proposed in the literature. However, these techniques rarely exploit the flexibility on the algorithmic level, which can lead to additional benefits. In this paper, FIR filter multipliers are extensively characterized with power simulations, providing a methodology for the perturbation of the coefficients of baseline filters at the algorithm level to trade-off reduced power consumption for filter quality. The proposed optimization technique does not require any hardware overhead and it enables the possibility of scaling the power consumption of the filter at runtime, while ensuring the full baseline performance of any programmed filter whenever it is required. The analyzed FIR filters were fabricated in a 28nm FD-SOI test chip and measured at a near-threshold, 600mV supply voltage. For example, by carefully choosing slightly perturbed coefficients in a low-pass configuration, power savings of up to 33% are achieved when accepting a 3dB degradation on the stopband, as compared with the baseline implementation of the filter.
IEEE Transactions on Circuits and Systems I: Regular Papers. 2017. Vol. 64, num. 9, p. 2388 – 2400. DOI : 10.1109/TCSI.2017.2698138.PolarBear: A 28-nm FD-SOI ASIC for Decoding of Polar Codes
Polar codes are a recently proposed class of block codes that provably achieve the capacity of various communication channels. They received a lot of attention as they can do so with low-complexity encoding and decoding algorithms, and they have an explicit construction. Their recent inclusion in a 5G communication standard will only spur more research. However, only a couple of ASICs featuring decoders for polar codes were fabricated, and none of them implements a list-based decoding algorithm. In this paper, we present ASIC measurement results for a fabricated 28 nm CMOS chip that implements two different decoders: the first decoder is tailored toward error-correction performance and flexibility. It supports any code rate as well as three different decoding algorithms: successive cancellation (SC), SC flip and SC list (SCL). The flexible decoder can also decode both non-systematic and systematic polar codes. The second decoder targets speed and energy efficiency. We present measurement results for the first silicon-proven SCL decoder, where its coded throughput is shown to be of 306.8 Mbps with a latency of 3.34 us and an energy per bit of 418.3 pJ/bit at a clock frequency of 721 MHz for a supply of 1.3 V. The energy per bit drops down to 178.1 pJ/bit with a more modest clock frequency of 308 MHz, lower throughput of 130.9 Mbps and a reduced supply voltage of 0.9 V. For the other two operating modes, the energy per bit is shown to be of approximately 95 pJ/bit. The less flexible high-throughput unrolled decoder can achieve a coded throughput of 9.2 Gbps and a latency of 628 ns for a measured energy per bit of 1.15 pJ/bit at 451 MHz.
IEEE Journal of Emerging and Selected Topics in Circuits and Systems. 2017. Vol. 7, num. 4, p. 616 – 629. DOI : 10.1109/JETCAS.2017.2745704.Conference Papers
Blind detection of polar codes
Polar codes were recently chosen to protect the control channel information in the next-generation mobile communication standard (5G) defined by the 3GPP. As a result, receivers will have to implement blind detection of polar coded frames in order to keep complexity, latency, and power consumption tractable. As a newly proposed class of block codes, the problem of polar-code blind detection has received very little attention. In this work, we propose a low-complexity blind-detection algorithm for polar-encoded frames. We base this algorithm on a novel detection metric with update rules that leverage the a priori knowledge of the frozen-bit locations, exploiting the inherent structures that these locations impose on a polar-encoded block of data. We show that the proposed detection metric allows to clearly distinguish polar-encoded frames from other types of data by considering the cumulative distribution functions of the detection metric, and the receiver operating characteristic. The presented results are tailored to the 5G standardization effort discussions, i.e., we consider a short low-rate polar code concatenated with a CRC.
2017. IEEE International Workshop on Signal Processing Systems (SiPS), Lorient, France, October 3-5, 2017. DOI : 10.1109/SiPS.2017.8109977.Comparison of Polar Decoders with Existing Low-Density Parity-Check and Turbo Decoders
Polar codes are a recently proposed family of provably capacity-achieving error-correction codes that received a lot of attention. While their theoretical properties render them interesting, their practicality compared to other types of codes has not been thoroughly studied. Towards this end, in this paper, we perform a comparison of polar decoders against LDPC and Turbo decoders that are used in existing communications standards. More specifically, we compare both the error-correction performance and the hardware efficiency of the corresponding hardware implementations. This comparison enables us to identify applications where polar codes are superior to existing error-correction coding solutions as well as to determine the most promising research direction in terms of the hardware implementation of polar decoders.
2017. IEEE Wireless Communications and Networking Conference (WCNC), San Francisco, CA, USA, Mar. 2017. p. 1 – 6. DOI : 10.1109/WCNCW.2017.7919106.Theses
High-Speed Wireline Link Design
High-speed serial links are a crucial application of semiconductor technology and have been the enabler of the scaling of computing systems. The increasing data-rate requirements of these links have only been partially satisfied by advancements in process technologies. Implementing serial links operating at a data rate of more than 25 Gb/s thus requires constantly inventing smarter algorithms, architectures, power-management techniques, and new standards. The Viterbi algorithm (VA), an attractive solution for symbol detection in the presence of intersymbol interference (ISI) and noise, minimizes the error probability in detecting the whole symbol sequence, instead of a single symbol as in decision-feedback equalization. The bit-error-rate (BER) performance of the VA is thus better than that of symbol-by-symbol detectors because the VA does not cancel the ISI, but rather uses the information embedded therein to maximize the reliability of its decisions. However, implementing a maximum-likelihood sequence detector (MLSD) realizing the VA may be prohibitive because design specifications regarding area, latency, power consumption, and speed may not be satisfied concurrently. Suboptimal solutions, such as feed-forward and decision-feedback equalizers (DFEs), may therefore be chosen instead of MLSDs due to their lower implementation complexity. A reduced-state sequence detector (RSSD) reduces the implementation complexity of the MLSD with negligible performance degradation by exploiting set-partitioning principles and embedded per-survivor decision feedback. A sliding-block or systolic-array Viterbi detector (VD) breaks the speed bottleneck in sequence-detector implementations by parallelizing the operation of the VA. In this thesis, we implement a 56-Gb/s four-level pulse-amplitude-modulation (4-PAM) DFE to demonstrate the feasibility of a multi-level analog-to-digital-converter (ADC)-based symbol-by-symbol detector in terms of achievable area, BER, energy-efficiency, latency, and speed figures comparable to those of analog solutions. Furthermore, to improve the BER figures, we implement a 25.6-Gb/s 4-PAM reduced-state sliding-block VD to demonstrate the feasibility of a multi-level sequence detector in terms of achievable area, energy-efficiency, latency, and speed figures comparable to those of symbol-by-symbol detectors. Moreover, we develop a sliding-block VA with optimized unequal synchronization and survivor path memory lengths to reduce its implementation complexity, latency, and power consumption. An increase in speed is thereby achieved at the same implementation complexity, BER, latency, and power consumption. We then develop a novel VD with embedded per-survivor decision feedback, whose longest path contains only one adder, to increase its speed significantly. Next, we propose a concatenated-coding scheme using an outer Reed—Solomon code and a four-dimensional (4-D) 5-PAM inner trellis-coded-modulation (TCM) scheme to achieve signal-to-noise-ratio gains without bandwidth expansion. Finally, we implement a 70-Gb/s 4-D 5-PAM systolic-array TCM decoder with eight states, which includes an inverse Tomlinson—Harashima precoder, to demonstrate the feasibility of a multi-level ADC-based sequence decoder in terms of achievable speed figures comparable to those of sequence detectors.
Lausanne, EPFL, 2017.Talks
Polar codes and APSK modulation – Just good friends
Information Theory and Applications Workshop (ITA), San Diego, CA, USA, Feb. 12-17, 2017.2016
Journal Articles
Cross-Layer Energy-Efficiency Optimization of Packet Based Wireless MIMO Communication Systems
Energy in today’s short-range wireless communication is mostly spent on the analog-and digital hardware rather than on radiated power. Hence, purely information-theoretic considerations fail to achieve the lowest energy per information bit and the optimization process must carefully consider the overall transceiver. In this paper, we propose to perform cross-layer optimization, based on an energy-aware rate adaptation scheme combined with a physical layer that is able to properly adjust its processing effort to the data rate and the channel conditions to minimize the energy consumption per information bit. This energy proportional behavior is enabled by extending the classical system modes with additional configuration parameters at the various layers. Fine grained models of the power consumption of the hardware are developed to provide awareness of the physical layer capabilities to the medium access control layer. The joint application of the proposed energy-aware rate adaptation and modifications to the physical layer of an IEEE 802.11n system, improves energy-efficiency ( averaged over many noise and channel realizations) in all considered scenarios by up to 44 %.
Journal Of Signal Processing Systems For Signal Image And Video Technology. 2016. Vol. 85, num. 1, p. 129 – 142. DOI : 10.1007/s11265-015-1003-7.A Low-Voltage Radiation-Hardened 13T SRAM Bitcell for Ultralow Power Space Applications
Continuous transistor scaling, coupled with the growing demand for low-voltage, low-power applications, increases the susceptibility of VLSI circuits to soft-errors, especially when exposed to extreme environmental conditions, such as those encountered by space applications. The most vulnerable of these circuits are memory arrays that cover large areas of the silicon die and often store critical data. Radiation hardening of embedded memory blocks is commonly achieved by implementing extremely large bitcells or redundant arrays and maintaining a relatively high operating voltage; however, in addition to the resulting area overhead, this often limits the minimum operating voltage of the entire system leading to significant power consumption. In this paper, we propose the first radiation-hardened static random access memory (SRAM) bitcell targeted at low-voltage functionality, while maintaining high soft-error robustness. The proposed 13T employs a novel dual-driven separated-feedback mechanism to tolerate upsets with charge deposits as high as 500 fC at a scaled 500-mV supply voltage. A 32×32 bit memory macro was designed and fabricated in a standard 0.18-mu m CMOS process, showing full read and write functionality down to the subthreshold voltage of 300 mV. This is achieved with a cell layout that is only 2x larger than a reference 6T SRAM cell drawn with standard design rules.
Ieee Transactions On Very Large Scale Integration (Vlsi) Systems. 2016. Vol. 24, num. 8, p. 2622 – 2633. DOI : 10.1109/Tvlsi.2016.2518220.Wireless Channel Characterization in Burning Buildings Over 100-1000 MHz
A 3-D implementation of the finite-difference time-domain (FDTD) method is used to model 100-1000-MHz radio wave propagation in a generalized office building. Fire within this building is modeled as a cold plasma medium. The presence of fire is found to decrease the sector-averaged received power by up to 10 dB. The FDTD results also showing propagation through fire can introduce rotation in linearly polarized signals, increasing the power of cross-polarized components. Uncertainties in the plasma properties are modeled using nonintrusive polynomial chaos, and can introduce up to +/- 8 dB variation in the sector-averaged power.
Ieee Transactions On Antennas And Propagation. 2016. Vol. 64, num. 7, p. 3265 – 3269. DOI : 10.1109/Tap.2016.2562671.Power, Area, and Performance Optimization of Standard Cell Memory Arrays Through Controlled Placement
Embedded memory remains a major bottleneck in current integrated circuit design in terms of silicon area, power dissipation, and performance; however, static random access memories (SRAMs) are almost exclusively supplied by a small number of vendors through memory generators, targeted at rather generic design specifications. As an alternative, standard cell memories (SCMs) can be defined, synthesized, and placed and routed as an integral part of a given digital system, providing complete design flexibility, good energy efficiency, low-voltage operation, and even area efficiency for small memory blocks. Yet implementing an SCM block with a standard digital flow often fails to exploit the distinct and regular structure of such an array, leaving room for optimization. In this article, we present a design methodology for optimizing the physical implementation of SCM macros as part of the standard design flow. This methodology introduces controlled placement, leading to a structured, noncongested layout with close to 100% placement utilization, resulting in a smaller silicon footprint, reduced wire length, and lower power consumption compared to SCMs without controlled placement. This methodology is demonstrated on SCM macros of various sizes and aspect ratios in a state-of-the-art 28nm fully depleted silicon-on-insulator technology, and compared with equivalent macros designed with the noncontrolled, standard flow, as well as with foundry-supplied SRAM macros. The controlled SCMs provide an average 25% reduction in area as compared to noncontrolled implementations while achieving a smaller size than SRAM macros of up to 1Kbyte. Power and performance comparisons of controlled SCM blocks of a commonly found 256 x 32 (1 Kbyte) memory with foundry-provided SRAMs show greater than 65% and 10% reduction in read and write power, respectively, while providing faster access than their SRAM counterparts, despite being of an aspect ratio that is typically unfavorable for SCMs. In addition, the SCM blocks function correctly with a supply voltage as low as 0.3V, well below the lower limit of even the SRAM macros optimized for low-voltage operation. The controlled placement methodology is applied within a full-chip physical implementation flow of an OpenRISC-based test chip, providing more than 50% power reduction compared to equivalently sized compiled SRAMs under a benchmark application.
Acm Transactions On Design Automation Of Electronic Systems. 2016. Vol. 21, num. 4, p. 59. DOI : 10.1145/2890498.Spatial Multiplexing of QPSK Signals With a Single Radio: Antenna Design and Over-the-Air Experiments
This paper describes the implementation and performance analysis of the first fully operational beam-space multiple-input multiple-output (MIMO) antenna for the spatial multiplexing of two QPSK streams. The antenna is composed of a planar three-port radiator with two varactor diodes terminating the passive ports. Pattern reconfiguration is used to encode the MIMO information onto orthogonal virtual basis patterns in the far field. A measurement campaign was conducted to compare the performance of the beam-space MIMO system with a conventional 2 x 2 MIMO system under realistic propagation conditions. Propagation measurements were conducted for both systems and the mutual information and symbol error rates were estimated from Monte-Carlo simulations over the measured channel matrices. The results show the beam-space MIMO system and the conventional MIMO system exhibit similar finite-constellation capacity and error performance in nonline-of-sight scenarios when there is sufficient scattering in the channel. In comparison, in line-of-sight channels, the capacity performance is observed to depend on the relative polarization of the receiving antennas.
Ieee Transactions On Antennas And Propagation. 2016. Vol. 64, num. 12, p. 5131 – 5145. DOI : 10.1109/Tap.2016.2624138.Synthesis of Dual Mode Logic
In recent years, the major focus of VLSI design has shifted from high-speed to low-power consumption. While standard CMOS-based digital design provides substantial flexibility during pre-silicon design phases, the characteristics of the gates are set by fabrication variations and environmental conditions and cannot easily be changed at runtime. The recently proposed Dual Mode Logic (DML) family provides a novel approach to provide this capability by introducing two configurable operating modes, static and dynamic, that enable fine-grained control of the power-performance tradeoff of a logic path. However, the introduction of a new topology requires the development of both a design methodology and techniques for integration in a robust design automation flow. Standard synthesis tools do not support dynamic gates, and in particular, dual-characteristic gates. Therefore, until now, DML has been limited to small, custom-made blocks and components. In this paper, we present a novel approach for the integration of DML into standard electronic design automation tools, as part of the standard digital design flow. The development of this approach and the accompanying design methodology enables DML to be used in larger designs, such as state-of-the-art, high-speed and/or low-power SoCs. We demonstrate the employment of the proposed approach in order to benefit from DML properties, and reduce the power consumption, while simultaneously improving the operating frequency of a number of test designs. (C) 2016 Elsevier Ltd. All rights reserved.
Integration-The Vlsi Journal. 2016. Vol. 55, p. 246 – 253. DOI : 10.1016/j.vlsi.2016.07.004.Ultra Low Voltage Synthesizable Memories: A Trade-Off Discussion in 65 nm CMOS
In this study, design considerations for ultra low voltage (ULV) standard-cell based memories (SCM) are presented. Trade-offs for area cost, leakage power, access time, and access energy are discussed and realized using different read logic styles, latch architecture designs, and process options. Furthermore, deployment of multiple threshold voltages (Vth) options in a single standard-cell/bitcell enables additional architectural choices. Silicon measurements from five memory designs, optimized at the transistor level in conjunction with gate-level optimizations, are considered to demonstrate the different trade-off corners. Measurements show that substituting the storage element in an SCM with a D-latch using transistor stacking and channel length stretching results in lowest leakage power. Alternatively, a pass-transistor based latch as storage element reduces the area footprint at a cost of reduced access speed, which can be compensated by using a lower-Vth pass-transistor. However, relatively high speed (tens of MHz) in the near-to subthreshold (sub-Vth) region is achievable if general purpose transistors are used instead of low power transistors. A discussion is included to illustrate when to implement ULV memories using SCMs and when to choose sub-Vth SRAMs. The discussion shows that the border is between 4-6 kb, depending on the number of words and the wordlength configuration.
Ieee Transactions On Circuits And Systems I-Regular Papers. 2016. Vol. 63, num. 6, p. 806 – 817. DOI : 10.1109/Tcsi.2016.2537931.Single-Supply 3T Gain-Cell for Low-Voltage Low-Power Applications
Logic compatible gain cell (GC) embedded DRAM (eDRAM) arrays are considered an alternative to SRAM due to their small size, non-ratioed operation, low static leakage, and 2-port functionality. However, traditional GC-eDRAM implementations require boosted control signals in order to write full voltage levels to the cell to reduce the refresh rate and shorten access times. These boosted levels require either an extra power supply or on-chip charge pumps, as well as non-trivial level shifting and toleration of high voltage levels. In this paper, we present a novel, logic compatible, 3T GC-eDRAM bitcell that operates with a single supply voltage and provides superior write capability to conventional GC structures. The proposed circuit is demonstrated with a 2kb memory macro that was designed and fabricated in a mature 0.18um CMOS process, targeted at low-power, energy-efficient applications. The test array is powered with a single supply of 900mV, showing an 0.8ms worst-case retention time, a 1.3ns write-access time, and 2.4pW/bit of retention power. The proposed topology provides a bitcell area reduction of 43%, as compared to a redrawn 6T SRAM in the same technology, and an overall macro area reduction of 67% including peripherals.
Ieee Transactions On Very Large Scale Integration (Vlsi) Systems. 2016. Vol. 24, num. 1, p. 358 – 362. DOI : 10.1109/TVLSI.2015.2394459.An Efficient Tool for the Assisted Design of SAR ADCs Capacitive DACs
The optimal design of SAR ADCs requires the accurate estimate of nonlinearity and parasitic capacitance effects in the feedback charge redistribution DAC. Since both contributions depend on the specific array topology, complex calculations, custom modeling and heavy simulations in common circuit design environments are often required. This paper presents a MATLAB-based numerical environment to assist the design of the charge redistribution DACs adopted in SAR ADCs. The tool performs both parametric and statistical simulations taking into account capacitive mismatch and parasitic capacitances computing both differential and integral nonlinearity (DNL, INL). An excellent agreement is obtained with the results of circuit simulators (e.g. Cadence Spectre) featuring up to 104 shorter simulation time, allowing statistical simulations that would be otherwise impracticable. The switching energy and SNDR degradation due to static nonlinear effects are also estimated. Simulations and measurements on three designed and two fabricated prototypes confirm that the proposed tool can be used as a valid instrument to assist the design of a charge redistribution SAR ADC and to predict its static and dynamic metrics. (C) 2015 Elsevier B.V. All rights reserved.
Integration, the VLSI Journal (Elsevier). 2016. Vol. 53, num. March 2016, p. 88 – 99. DOI : 10.1016/j.vlsi.2015.12.005.Silicon-Proven, Per-Cell Retention Time Distribution Model for Gain-Cell Based eDRAMs
Gain-cell embedded DRAM (GC-eDRAM) is an interesting alternative to SRAMfor reasons such as high density, low bitcell leakage, logic compatibility, and suitability for 2-port memories. Themajor drawbacks of GC-eDRAMs are their limited data retention times (RTs) and the large spread of RT across an array, which degrade energy-efficiency due to refresh cycles. While the array refresh rate can be determined according to circuit simulation or post-manufacturing calibration, there is a lack of analytical and statistical RT models for GC-eDRAM that could unveil the limiters and circuit parameters that lead to the large observed RT spreads. In this work, we derive the first comprehensive analytical model for the statistical distribution of the per-cell retention time of 2T-bitcell GC-eDRAMs, which is found to follow a log-normal distribution. The accuracy of the proposed retention time model is verified by extensive Monte Carlo and worst case distance circuit simulations and silicon measurements of an 0.18 mu m test chip. Furthermore, a sensitivity analysis unveils the circuit parameters that have the highest impact on the RT spread. Interestingly, the variability of the threshold voltage of the write access transistor has a much higher impact on the RT spread than the variability of any other circuit parameter, including the storage node capacitor. This holds true under process scaling, for nodes as advanced as 28 nm, as shown through simulation. The insights gained from the retention time model help circuit designers achieve better GC-eDRAMs with longer RTs and sharper RT distributions. In addition, the herein presented model can be used as a basis to study the reliability/energy trade-off for GC-eDRAM usage in fault-tolerant VLSI systems.
Ieee Transactions On Circuits And Systems I-Regular Papers. 2016. Vol. 63, num. 2, p. 222 – 232. DOI : 10.1109/Tcsi.2015.2512706.Conference Papers
Digital Predistortion of Power Amplifier Non-Linearities for Full-Duplex Transceivers
Non-linearities introduced by the power amplifier stage can significantly reduce the performance of self-interference cancellation in full-duplex transceivers. Accordingly, we propose a full-duplex system architecture that predistorts the digital baseband transmit signal to account for the non-linear memory effects of the power amplifier. Implementation results for a 5 MHz OFDM signal (operating with 20 dBm average transmit power) on a full-duplex testbed show that a further 13 dB suppression can be obtained, compared to the case when no predistortion is applied. The power levels of out-of-band emissions are also significantly reduced.
2016. 17th IEEE International workshop on Signal Processing Advances in Wireless Communications, Edinburgh, Scotland, UK, July 3-6, 2016. DOI : 10.1109/SPAWC.2016.7536811.DynOR: A 32-bit Microprocessor in 28 nm FD-SOI with Cycle-By-Cycle Dynamic Clock Adjustment
This paper presents DynOR, a 32-bit 6-stage OpenRISC microprocessor with dynamic clock adjustment. To alleviate the issue of unused dynamic timing margins, the clock period of the processor is adjusted on a cycle-by-cycle level, based on the instruction types currently in flight in the pipeline. To this end, we employ a custom designed clock generation unit, capable of immediate glitch-free adjustments of the clock period over a wide range with fine granularity. Our chip measurements in 28nm FD-SOI technology show that DynOR provides an average speedup of 19% in program execution over a wide range of operating conditions, with a peak speedup for certain applications of up to 41%. Furthermore, this speedup can be traded off against energy, to reduce the chip power consumption for a typical die by up to 15%, compared to a static clocking scheme based on worst case excitation.
2016. 42nd European Solid-State Circuits Conference (ESSCIRC), Lausanne, Switzerland, September 12-15, 2016. p. 261 – 264. DOI : 10.1109/ESSCIRC.2016.7598292.High-Speed Link With Trellis-Coded Modulation and Reed Solomon Coding
A high-performance low-latency transmission system based on a concatenated code consisting of inner four-dimensional five-level pulse-amplitude-modulation (5-PAM) trellis-coded modulation and an outer Reed-Solomon (RS) code is proposed for a high-speed data link over time-dispersive channels. The implementation of high-speed sequence detection in state-of-the-art technology nodes is enabled by a novel reduced state Viterhi detector which includes only one addition in the longest path. The bit-error-rate performance of the proposed transmission scheme for a real-world channel model is evaluated by simulations. Compared with the RS-encoded 4-PAM transmission scheme in the IEEE P802.3bj standard, the proposed transmission system achieves a signal-to-noise ratio advantage of 0.7 dB at a bit error rate of 10(-6) without expanding the bandwidth.
2016. IEEE Conference on Standards for Communications and Networking (CSCN), Berlin, GERMANY, OCT 31-NOV 02, 2016. DOI : 10.1109/CSCN.2016.7785181.Energy vs. Reliability Trade-offs Exploration in Biomedical Ultra-Low Power Devices
State-of-the-art wearable devices such as embedded biomedical monitoring systems apply voltage scaling to lower as much as possible their energy consumption and achieve longer battery lifetimes. While embedded memories often rely on Error Correction Codes (ECC) for error protection, in this paper we explore how the characteristics of biomedical applications can be exploited to develop new techniques with lower power overhead. We then introduce the Dynamic eRror compEnsation And Masking (DREAM) technique, that provides partial memory protection with less area and power overheads than ECC. Different tradeoffs between the error correction ability of the techniques and their energy consumption are examined to conclude that, when properly applied, DREAM consumes 21% less energy than a traditional ECC with Single Error Correction and Double Error Detection (SEC/DED) capabilities.
2016. Design, Automation and Test in Europe Conference (DATE ’16), Dresden, Germany, March 14-18, 2016. p. 838 – 841.A 4.1 pJ/b 25.6 Gb/s 4-PAM Reduced-State Sliding-Block Viterbi Detector in 14 nm CMOS
The implementation of a digital four-level pulse-amplitude-modulation reduced-state sliding-block Viterbi detector (VD) with two substates and two embedded per-survivor decision-feedback taps operating at one-eighth of the modulation rate is described. Implemented in an experimental chip fabricated in 14nm CMOS, the VD is designed to recover data at 25.6 Gb/s over an emulated time-dispersive channel. The power consumption of the VD together with the test circuitry is 105mW at a supply of 0.7V, achieving an overall energy efficiency of 4.1 pJ/b. At a supply of 0.8V, a data rate of 30.4 Gb/s is achieved with an energy efficiency of 5.3 pJ/b. The VD occupies an area of 0.507 +/- 0.717mm(2). Experimental results showing system performance are obtained using a (2(15)-1)-bit pseudo-random binary sequence. The impact on the bit error rate of the synchronization length for block initialization is also measured.
2016. 46th European Solid-State Device Research Conference (ESSDERC) / 42nd European Solid-State Circuits Conference (ESSCIRC), Lausanne, SWITZERLAND, SEP 12-15, 2016. p. 309 – 312. DOI : 10.1109/ESSCIRC.2016.7598304.Partitioned Successive-Cancellation List Decoding Of Polar Codes
Successive-cancellation list (SCL) decoding is an algorithm that provides very good error-correction performance for polar codes. However, its hardware implementation requires a large amount of memory, mainly to store intermediate results. In this paper, a partitioned SCL algorithm is proposed to reduce the large memory requirements of the conventional SCL algorithm. The decoder tree is broken into partitions that are decoded separately. We show that with careful selection of list sizes and number of partitions, the proposed algorithm can outperform conventional SCL while requiring less memory.
2016. IEEE International Conference on Acoustics, Speech, and Signal Processing, Shanghai, PEOPLES R CHINA, MAR 20-25, 2016. p. 957 – 960. DOI : 10.1109/ICASSP.2016.7471817.Approximate Computing for Unreliable Silicon
2016. 11th IEEE International Conference on Design & Technology of Integrated Systems in Nanoscale Era (DTIS). DOI : 10.1109/DTIS.2016.7483878.Statistical Fault Injection for Impact-Evaluation of Timing Errors on Application Performance
This paper proposes a novel approach to modeling of gate level timing errors during high-level instruction set simulation. In contrast to conventional, purely random fault injection, our physically motivated approach directly relates to the underlying circuit structure, hence allowing for a significantly more detailed characterization of application performance under scaled frequency / voltage (including supply noise). The model uses gate level timing statistics extracted by dynamic timing analysis from the post place & route netlist of a general-purpose processor to perform instruction-aware fault injections. We employ a 28 nm OpenRISC core as a case study, to demonstrate how statistical fault injection provides a more accurate and realistic analysis of power vs. error performance.
2016. 53rd ACM/EDAC/IEEE Design Automation Conference (DAC), Austin, Texas, USA, June 5-9, 2016. p. 13:1 – 13:6. DOI : 10.1145/2897937.2898095.Sliding Window Spectrum Sensing for Full-Duplex Cognitive Radios with Low Access-Latency
2016. IEEE 83rd Vehicular Technology Conference, Nanjing, China, 15-18 May. DOI : 10.1109/VTCSpring.2016.7504477.A Multi-Gbps Unrolled Hardware List Decoder Systematic Polar Code
Polar codes are a new class of block codes with an explicit construction that provably achieve the capacity of various communications channels, even with the low-complexity successive-cancellation (SC) decoding algorithm. Yet, the more complex successive-cancellation list (SCL) decoding algorithm is gathering more attention lately as it significantly improves the error-correction performance of short- to moderate-length polar codes, especially when they are concatenated with a cyclic redundancy check code. However, as SCL decoding explores several decoding paths, existing hardware implementations tend to be significantly slower than SC-based decoders. In this paper, we show how the unrolling technique, which has already been used in the context of SC decoding, can be adapted to SCL decoding yielding a multi-Gbps SCL-based polar decoder with an error-correction performance that is competitive when compared to an LDPC code of similar length and rate. Post-place-and-route ASIC results for 28 nm CMOS are provided showing that this decoder can sustain a throughput greater than 10 Gbps at 468 MHz with an energy efficiency of 7.25 pJ/bit.
2016. Asilomar Conference on Signals, Systems, and Computers, Pacific Grove, CA, USA, Nov. 2016. p. 1194 – 1198. DOI : 10.1109/ACSSC.2016.7869561.A Low-Power Correlator for Wakeup Receivers with Algorithm Pruning through Early Termination
A low-complexity, low-power digital correlator for wakeup receivers is presented. With the proposed algorithm, unnecessary computational cycles are dynamically pruned from the correlation using an early threshold check. For the algorithm, we provide a rigorous mathematical analysis for the associated complexity/performance trade-offs. Furthermore, a low overhead hardware architecture with early-termination capability is developed and implemented in a 0.18um CMOS technology. The post layout power analysis shows that the presented architecture can reduce power by up to 32% when compared to the conventional architecture with negligible degradation in detection probability and without degradation in false-alarm probability.
2016. 2016 IEEE International Symposium on Circuits and Systems (ISCAS), Montreal, Canada, May 22-25, 2016. p. 2667 – 2670. DOI : 10.1109/ISCAS.2016.7539142.A Process Compensated Gain Cell Embedded-DRAM for Ultra-Low-Power Variation-Aware Design
Gain cell embedded DRAM (GC-eDRAM) is a high-density alternative to SRAM for ultra-low-power systems. However, due to its dynamic nature, GC-eDRAM requires power-hungry refresh cycles to ensure data retention. Traditional design approaches dictate configuration of the refresh rate according to the worst bitcell, when biased at low-probability, worst-case conditions. However, due to the process variations and local mismatch that can significantly deteriorate the data retention time of a GC-eDRAM bitcell, this design approach often leads to a large power overhead. In this paper, we present a novel GC-eDRAM architecture, incorporating several techniques for variation-aware operation. The primary feature of this architecture is an improved replica scheme for process compensated access tracking that enables calibration for process variations and adaptive refresh according to the array access statistics. The array is shown to ensure data integrity, providing as much as a 7x reduction in retention power over worst-case refresh-rate design for 20% write activity.
2016. IEEE International Symposium on Circuits and Systems (ISCAS), Montreal, CANADA, MAY 22-25, 2016. p. 1006 – 1009. DOI : 10.1109/ISCAS.2016.7527413.193 MOPS/mW @ 162 MOPS, 0.32V to 1.15V voltage range multi-core accelerator for energy efficient parallel and sequential digital processing
Low power (mW) and high performance (GOPS) are strong requirements for compute-intensive signal processing in E-health, Internet-of-Things, and wearable applications. This work presents a building block for programmable Ultra-Low Power accelerators, namely a tightly-coupled computing cluster that supports parallel and sequential execution at high energy efficiency over a wide range of workload requirements. The cluster, implemented in 28nm UTBB FD-SOI technology, achieves peak energy efficiency in the near-threshold (NVT) operating region: 193 MOPS/mW at 162 MOPS for parallel workloads, and 90 MOPS/mW at 68 MOPS for sequential workloads at 0.46V and 0.5V, respectively. The energy efficient operating range is wide (0.32V to 1.15V), also meeting the design goal of 1 GOPS within a 10 mW power envelope (at 0.66V).
2016. 2016 IEEE Symposium in Low-Power and High-Speed Chips (COOL CHIPS XIX), Yokohama, Japan, April 20-22, 2016. DOI : 10.1109/CoolChips.2016.7503670.Hardware Decoders for Polar Codes: An Overview
Polar codes are an exciting new class of error correcting codes that achieve the symmetric capacity of memoryless channels. Many decoding algorithms were developed and implemented, addressing various application requirements: from error-correction performance rivaling that of LDPC codes to very high throughput or low-complexity decoders. In this work, we review the state of the art in polar decoders implementing the successive-cancellation, belief propagation, and list decoding algorithms, illustrating their advantages.
2016. IEEE International Symposium on Circuits and Systems (ISCAS), Montreal, CANADA, MAY 22-25, 2016. p. 149 – 152. DOI : 10.1109/ISCAS.2016.7527192.Statistical Fault Injection for Impact-Evaluation of Timing Errors on Application Performance
This paper proposes a novel approach to modeling of gate level timing errors during high-level instruction set simulation. hi contrast to conventional, purely random fault injection, our physically motivated approach directly relates to the underlying circuit structure, hence allowing for a significantly more detailed characterization of application performance under scaled frequency / voltage (including supply noise). The model uses gate level timing statistics extracted by dynamic timing analysis from the post place & route netlist of a general-purpose processor to perform instruction aware fault injections. We employ a 28 nm OpenRISC core as a case study, to demonstrate how statistical fault injection provides a more accurate and realistic analysis of power vs. error performance.
2016. 53rd ACM/EDAC/IEEE Design Automation Conference (DAC), Austin, TX, JUN 05-09, 2016. DOI : 10.1145/2897937.2696095.Theses
Microarchitectural Low-Power Design Techniques for Embedded Microprocessors
With the omnipresence of embedded processing in all forms of electronics today, there is a strong trend towards wireless, battery-powered, portable embedded systems which have to operate under stringent energy constraints. Consequently, low power consumption and high energy efficiency have emerged as the two key criteria for embedded microprocessor design. In this thesis we present a range of microarchitectural low-power design techniques which enable the increase of performance for embedded microprocessors and/or the reduction of energy consumption, e.g., through voltage scaling. In the context of cryptographic applications, we explore the effectiveness of instruction set extensions (ISEs) for a range of different cryptographic hash functions (SHA-3 candidates) on a 16-bit microcontroller architecture (PIC24). Specifically, we demonstrate the effectiveness of light-weight ISEs based on lookup table integration and microcoded instructions using finite state machines for operand and address generation. On-node processing in autonomous wireless sensor node devices requires deeply embedded cores with extremely low power consumption. To address this need, we present TamaRISC, a custom-designed ISA with a corresponding ultra-low-power microarchitecture implementation. The TamaRISC architecture is employed in conjunction with an ISE and standard cell memories to design a sub-threshold capable processor system targeted at compressed sensing applications. We furthermore employ TamaRISC in a hybrid SIMD/MIMD multi-core architecture targeted at moderate to high processing requirements (> 1 MOPS). A range of different microarchitectural techniques for efficient memory organization are presented. Specifically, we introduce a configurable data memory mapping technique for private and shared access, as well as instruction broadcast together with synchronized code execution based on checkpointing. We then study an inherent suboptimality due to the worst-case design principle in synchronous circuits, and introduce the concept of dynamic timing margins. We show that dynamic timing margins exist in microprocessor circuits, and that these margins are to a large extent state-dependent and that they are correlated to the sequences of instruction types which are executed within the processor pipeline. To perform this analysis we propose a circuit/processor characterization flow and tool called dynamic timing analysis. Moreover, this flow is employed in order to devise a high-level instruction set simulation environment for impact-evaluation of timing errors on application performance. The presented approach improves the state of the art significantly in terms of simulation accuracy through the use of statistical fault injection. The dynamic timing margins in microprocessors are then systematically exploited for throughput improvements or energy reductions via our proposed instruction-based dynamic clock adjustment (DCA) technique. To this end, we introduce a 6-stage 32-bit microprocessor with cycle-by-cycle DCA. Besides a comprehensive design flow and simulation environment for evaluation of the DCA approach, we additionally present a silicon prototype of a DCA-enabled OpenRISC microarchitecture fabricated in 28 nm FD-SOI CMOS. The test chip includes a suitable clock generation unit which allows for cycle-by-cycle DCA over a wide range with fine granularity at frequencies exceeding 1 GHz. Measurement results of speedups and power reductions are provided.
Lausanne, EPFL, 2016.Modulation, Coding, and Receiver Design for Gigabit mmWave Communication
While wireless communication has become an ubiquitous part of our daily life and the world around us, it has not been able yet to deliver the multi-gigabit throughput required for applications like high-definition video transmission or cellular backhaul communication. The throughput limitation of current wireless systems is mainly the result of a shortage of spectrum and the problem of congestion. Recent advancements in circuit design allow the realization of analog frontends for mmWave frequencies between 30GHz and 300GHz, making abundant unused spectrum accessible. However, the transition to mmWave carrier frequencies and GHz bandwidths comes with new challenges for wireless receiver design. Large variations of the channel conditions and high symbol rates require flexible but power-efficient receiver designs. This thesis investigates receiver algorithms and architectures that enable multi-gigabit mmWave communication. Using a system-level approach, the design options between low-power time-domain and power-hungry frequency-domain signal processing are explored. The system discussion is started with an analysis of the problem of parameter synchronization in mmWave systems and its impact on system design. The proposed synchronization architecture extends known synchronization techniques to provide greater flexibility regarding the operating environments and for system efficiency optimization. For frequency-selective environments, versatile single-carrier frequency domain equalization (SC-FDE) offers not only excellent channel equalization, but also the possibility to integrate additional baseband tasks without overhead. Hence, the high initial complexity of SC-FDE needs to be put in perspective to the complexity savings in the other parts of the baseband. Furthermore, an extension to the SC-FDE architecture is proposed that allows an adaptation of the equalization complexity by switching between a cyclic-prefix mode and a reduced block length overlap-save mode based on the delay spread. Approaching the problem of complexity adaptation from time-domain, a high-speed hardware architecture for the delayed decision feedback sequence estimation (DDFSE) algorithm is presented. DDFSE uses decision feedback to reduce the complexity of the sequence estimation and allows to set the system performance between the performance of full maximum-likelihood detection and pure decision feedback equalization. An implementation of the DDFSE architecture is demonstrated as part of an all-digital IEEE802.11ad baseband ASIC manufactured in 40nm CMOS. A flexible architecture for wideband mmWave receivers based on complex sub-sampling is presented. Complex sub-sampling combines the design advantages of sub-sampling receivers with the flexibility of direct-conversion receivers using a single passive component and a digital compensation scheme. Feasibility of the architecture is proven with a 16Gb/s hardware demonstrator. The demonstrator is used to explore the potential gain of non-equidistant constellations for high-throughput mmWave links. Specifically crafted amplitude phase-shift keying (APSK) modulation achieve 1dB average mutual information (AMI) advantage over quadrature amplitude modulation (QAM) in simulation and on the testbed hardware. The AMI advantage of APSK can be leveraged for a practical transmission using Polar codes which are trained specifically for the constellation.
Lausanne, EPFL, 2016.Hardware implementation aspects of polar decoders and ultra high-speed LDPC decoders
The goal of channel coding is to detect and correct errors that appear during the transmission of information. In the past few decades, channel coding has become an integral part of most communications standards as it improves the energy-efficiency of transceivers manyfold while only requiring a modest investment in terms of the required digital signal processing capabilities. The most commonly used channel codes in modern standards are low-density parity-check (LDPC) codes and Turbo codes, which were the first two types of codes to approach the capacity of several channels while still being practically implementable in hardware. The decoding algorithms for LDPC codes, in particular, are highly parallelizable and suitable for high-throughput applications. A new class of channel codes, called polar codes, was introduced recently. Polar codes have an explicit construction and low-complexity encoding and successive cancellation (SC) decoding algorithms. Moreover, polar codes are provably capacity achieving over a wide range of channels, making them very attractive from a theoretical perspective. Unfortunately, polar codes under standard SC decoding cannot compete with the LDPC and Turbo codes that are used in current standards in terms of their error-correcting performance. For this reason, several improved SC-based decoding algorithms have been introduced. The most prominent SC-based decoding algorithm is the successive cancellation list (SCL) decoding algorithm, which is powerful enough to approach the error-correcting performance of LDPC codes. The original SCL decoding algorithm was described in an arithmetic domain that is not well-suited for hardware implementations and is not clear how an efficient SCL decoder architecture can be implemented. To this end, in this thesis, we re-formulate the SCL decoding algorithm in two distinct arithmetic domains, we describe efficient hardware architectures to implement the resulting SCL decoders, and we compare the decoders with existing LDPC and Turbo decoders in terms of their error-correcting performance and their implementation efficiency. Due to the ongoing technology scaling, the feature sizes of integrated circuits keep shrinking at a remarkable pace. As transistors and memory cells keep shrinking, it becomes increasingly difficult and costly (in terms of both area and power) to ensure that the implemented digital circuits always operate correctly. Thus, manufactured digital signal processing circuits, including channel decoder circuits, may not always operate correctly. Instead of discarding these faulty dies or using costly circuit-level fault mitigation mechanisms, an alternative approach is to try to live with certain malfunctions, provided that the algorithm implemented by the circuit is sufficiently fault-tolerant. In this spirit, in this thesis we examine decoding of polar codes and LDPC codes under the assumption that the memories that are used within the decoders are not fully reliable. We show that, in both cases, there is inherent fault-tolerance and we also propose some methods to reduce the effect of memory faults on the error-correcting performance of the considered decoders.
Lausanne, EPFL, 2016.2015
Journal Articles
An Evolved GSM/EDGE Baseband ASIC Supporting Rx Diversity
In this paper, a baseband ASIC which supports receive diversity and soft-output Viterbi equalization for enhanced 2G networks is presented. It includes a transmitter and receiver with a symbol detector and a decoder with a dedicated incremental redundancy implementation, as well as the necessary control capability to autonomously communicate with the RF-IC. The ASIC is connected to an RF-IC to build a complete Evolved EDGE transceiver system. The transceiver system reaches a measured sensitivity close to -112 dBm for single-antenna GSM voice channels and achieves the reference interference performance for adjacent channels 11.4 dB above 3GPP requirements. It is the first reported solution which fulfills the most demanding 3GPP Downlink Advanced Receive Performance Phase 2 testcases specified for Rx-diversity. The ASIC occupies 6 mm(2) in 130 nm CMOS with a power consumption between 3.9 and 14 mW.
Ieee Journal Of Solid-State Circuits. 2015. Vol. 50, num. 7, p. 1690 – 1701. DOI : 10.1109/Jssc.2015.2417802.Performance estimation for indoor wireless systems using FDTD method
A three-dimensional implementation of the finite-difference time-domain method is used to estimate the down-link outage probability of a direct-sequence code division multiple access system operating in a multi-storey office building in the presence of co-channel interference. The numerical analysis is supported by experimental measurements and good agreement is found for the outage probability. Both simulation and measured results indicate that vertically aligned co-channel base stations have lower outage than a vertically staggered configuration, which is explained by examining the correlation between the desired and interfering signals.
Electronics Letters. 2015. Vol. 51, num. 17, p. 1376 – 1378. DOI : 10.1049/el.2015.1093.LLR-Based Successive Cancellation List Decoding of Polar Codes
We show that successive cancellation list decoding can be formulated exclusively using log-likelihood ratios. In addition to numerical stability, the log-likelihood ratio based formulation has useful properties that simplify the sorting step involved in successive cancellation list decoding. We propose a hardware architecture of the successive cancellation list decoder in the log-likelihood ratio domain which, compared with a log-likelihood domain implementation, requires less irregular and smaller memories. This simplification, together with the gains in the metric sorter, lead to to higher throughput per unit area than other recently proposed architectures. We then evaluate the empirical performance of the CRC-aided successive cancellation list decoder at different list sizes using different CRCs and conclude that it is important to adapt the CRC length to the list size in order to achieve the best error-rate performance of concatenated polar codes. Finally, we synthesize conventional successive cancellation decoders at large block-lengths with the same block-error probability as our proposed CRC-aided successive cancellation list decoders to demonstrate that, while our decoders have slightly lower throughput and larger area, they have a significantly smaller decoding latency.
Ieee Transactions On Signal Processing. 2015. Vol. 63, num. 19, p. 5165 – 5179. DOI : 10.1109/Tsp.2015.2439211.Enhancing Design Space Exploration by Extending CPU/GPU Specifications onto FPGAs
The design cycle for complex special-purpose computing systems is extremely costly and time-consuming. It involves a multiparametric design space exploration for optimization, followed by design verification. Designers of special purpose VLSI implementations often need to explore parameters, such as optimal bitwidth and data representation, through time-consuming Monte Carlo simulations. A prominent example of this simulation-based exploration process is the design of decoders for error correcting systems, such as the Low-Density Parity-Check (LDPC) codes adopted by modern communication standards, which involves thousands of Monte Carlo runs for each design point. Currently, high-performance computing offers a wide set of acceleration options that range from multicore CPUs to Graphics Processing Units (GPUs) and Field Programmable Gate Arrays (FPGAs). The exploitation of diverse target architectures is typically associated with developing multiple code versions, often using distinct programming paradigms. In this context, we evaluate the concept of retargeting a single OpenCL program to multiple platforms, thereby significantly reducing design time. A single OpenCL-based parallel kernel is used without modifications or code tuning on multicore CPUs, GPUs, and FPGAs. We use SOpenCL (Silicon to OpenCL), a tool that automatically converts OpenCL kernels to RTL in order to introduce FPGAs as a potential platform to efficiently execute simulations coded in OpenCL. We use LDPC decoding simulations as a case study. Experimental results were obtained by testing a variety of regular and irregular LDPC codes that range from short/medium (e.g., 8,000 bit) to long length (e.g., 64,800 bit) DVB-S2 codes. We observe that, depending on the design parameters to be simulated, on the dimension and phase of the design, the GPU or FPGA may suit different purposes more conveniently, thus providing different acceleration factors over conventional multicore CPUs.
Acm Transactions On Embedded Computing Systems. 2015. Vol. 14, num. 2, p. 33. DOI : 10.1145/2656207.A Fast Modular Method for True Variation-Aware Separatrix Tracing in Nanoscaled SRAMs
As memory density continues to grow in modern systems, accurate analysis of SRAM stability is increasingly important to ensure high yields. Traditional static noise margin metrics fail to capture the dynamic characteristics of SRAM behavior, leading to expensive over-design and disastrous under-design. One of the central components of more accurate dynamic stability analysis is the separatrix; however, its straightforward extraction is extremely time-consuming, and efficient methods are either non-accurate or extremely difficult to implement. In this paper, we propose a novel algorithm for fast separatrix tracing of any given SRAM topology, designed with industry standard transistor models in nano-scaled technologies. The proposed algorithm is applied to both standard 6T SRAM bitcells, as well as previously proposed alternative sub-threshold bitcells, providing up to three orders-of-magnitude speedup, as compared to brute force methods. In addition, for the first time, statistical Monte Carlo separatrix distributions are plotted.
Ieee Transactions On Very Large Scale Integration (Vlsi) Systems. 2015. Vol. 23, num. 10, p. 2034 – 2042. DOI : 10.1109/TVLSI.2014.2358699.Baseband and RF hardware impairments in full-duplex wireless systems: experimental characterisation and suppression
Hardware imperfections can significantly reduce the performance of full-duplex wireless systems by introducing non-idealities and random effects that make it challenging to fully suppress self-interference. Previous research has mostly focused on analysing the impact of hardware imperfections on full-duplex systems, based on simulations and theoretical models. In this paper, we follow a measurement-based approach to experimentally identify and isolate these hardware imperfections leading to residual self-interference in full-duplex nodes. Our measurements show the important role of images arising from in-phase and quadrature (IQ) imbalance in the transmitter and receiver mixers. We also observe baseband non-linearities in the digital-to-analog converters (DAC), which can introduce strong harmonic components in the transmitted signal that have not been considered previously. A corresponding general mathematical model to suppress these components of the self-interference signal arising from the hardware non-idealities is developed from the observations and measurements. Results from a 10 MHz bandwidth full-duplex system, operating at 2.48 GHz, show that up to 13 dB additional suppression, relative to state-of-the-art implementations, can be achieved by jointly compensating for IQ imbalance and DAC non-linearities.
EURASIP Journal on Wireless Communications and Networking. 2015. Vol. 2015, p. 142. DOI : 10.1186/s13638-015-0350-1.Conference Papers
A 3.52 Gb/s mmWave Baseband with Delayed Decision Feedback Sequence Estimation in 40 nm
We present a digital baseband ASIC for 60 GHz single-carrier (SC) transmission that is optimized for communication scenarios in which most of the energy is concentrated in the first few channel taps. Such scenarios occur for example in office environments with strong reflections. Our circuit targets close-to-optimum maximum-likelihood performance under such conditions. To this end, we show for the first time how a reduced-state-sequence-estimation algorithm can be realized for the 1760 MHz bandwidth of the IEEE 802.11ad standard. The equalizer is complemented in the frontend by a synchronization unit for frequency offset compensation as well as a Golay-sequence based channel estimator and in the backend by an low density parity check (LDPC) decoder. In 40nm CMOS we achieve a measured data rate of up to 3.52 Gb/s using QPSK modulation.
2015. 2015 IEEE Asian Solid-State Circuits Conference (A-SSCC), Xiamen, Fujian, China, November 9-11, 2015. p. 193 – 196. DOI : 10.1109/ASSCC.2015.7387455.Energy-Proportional Single-Carrier Frequency Domain Equalization for mmWave Wireless Communication
mmWave wireless communication is proposed for high-throughput and high-density applications. Due to the large channel bandwidth, mmWave systems face a large variation in the observed delay spread. Many proposed single-carrier (SC) mmWave systems rely on cyclic-prefix (CP) frequency domain equalization (FDE) in order to deal with worst case channel conditions. A downside of CP-FDE receivers is their constant energy consumption independent of the actual channel conditions. The alternative overlap-save (OS) FDE receiver can adapt its complexity, but exhibits an inferior equalization performance. By proposing a hybrid FDE approach the receiver can adapt its complexity and therefor its power consumption dynamically to the given channel conditions. Using the structural similarity of overlap-save and cyclic-prefix FDE architectures the proposed hybrid receiver can switch between the two modes of operation with minimum required hardware overhead. It is shown that the proposed hybrid receiver can significantly reduce its complexity in benign channel conditions while still matching the equalization properties of a conventional CP-FDE receivers in very frequency selective environments.
2015. 49th Asilomar Conference on Signals, Systems, and Computers, Pacific Grove, California, USA, November 8-11, 2015. p. 1133 – 1137. DOI : 10.1109/ACSSC.2015.7421317.Approximate Computing With Unreliable Dynamic Memories
Embedded memories account for a large fraction of the overall silicon area and power consumption in modern SoC(s). While embedded memories are typically realized with SRAM, alternative solutions, such as embedded dynamic memories (eDRAM), can provide higher density and/or reduced power consumption. One major challenge that impedes the widespread adoption of eDRAM is that they require frequent refreshes potentially reducing the availability of the memory in periods of high activity and also consuming significant amount of power due to such frequent refreshes. Reducing the refresh rate while on one hand can reduce the power overhead, if not performed in a timely manner, can cause some cells to lose their content potentially resulting in memory errors. In this paper, we consider extending the refresh period of gain-cell based dynamic memories beyond the worst-case point of failure, assuming that the resulting errors can be tolerated when the use-cases are in the domain of inherently error-resilient applications. For example, we observe that for various data mining applications, a large number of memory failures can be accepted with tolerable imprecision in output quality. In particular, our results indicate that by allowing as many as 177 errors in a 16kB memory, the maximum loss in output quality is 11%. We use this failure limit to study the impact of relaxing reliability constraints on memory availability and retention power for different technologies.
2015. International New Circuits And Systems Conference (NEWCAS), Grenoble, France, June 7-10, 2015. DOI : 10.1109/NEWCAS.2015.7182027.Fundamental Power Limits of SAR and ΔΣ Analog-to-Digital Converters
This work aims at estimating and comparing the power limits of ΔΣ and charge-redistribution successive-approximation register (CR-SAR) analog-to-digital converters (ADCs), in order to identify which topology is the most power-efficient for a target resolution. A power consumption model for mismatch-limited SAR ADCs and for discrete-time (DT) ΔΣ modulators is presented and validated against experimental data. SAR ADCs are found to be the best choice for low-to-medium resolutions, up to roughly 80 dB of dynamic range (DR). At high resolutions, on the other hand, ΔΣ modulators become more power-efficient. This is due to the intrinsic robustness of the ΔΣ modulation principle against circuit imperfections and non-idealities. Furthermore, a comparison of the area occupation of such topologies reveals that, at high resolutions and for a given dynamic range, ΔΣ ADCs result more area-efficient as well. © 2015 IEEE.
2015. IEEE Nordic Circuits and Systems Conference (NORCAS), Oslo, October 2015. DOI : 10.1109/NORCHIP.2015.7364387.Exploiting Dynamic Timing Margins in Microprocessors for Frequency-Over-Scaling with Instruction-Based Clock Adjustment
Static timing analysis provides the basis for setting the clock period of a microprocessor core, based on its worst-case critical path. However, depending on the design, this critical path is not always excited and therefore dynamic timing margins exist that can theoretically be exploited for the benefit of better speed or lower power consumption (through voltage scaling). This paper introduces predictive instruction-based dynamic clock adjustment as a technique to trim dynamic timing margins in pipelined microprocessors. To this end, we exploit the different timing requirements for individual instructions during the dynamically varying program execution flow without the need for complex circuit-level measures to detect and correct timing violations. We provide a design flow to extract the dynamic timing information for the design using post-layout dynamic timing analysis and we integrate the results into a custom cycle-accurate simulator. This simulator allows annotation of individual instructions with their impact on timing (in each pipeline stage) and rapidly derives the overall code execution time for complex benchmarks. The design methodology is illustrated at the microarchitecture level, demonstrating the performance and power gains possible on a 6-stage OpenRISC in-order general purpose processor core in a 28nm CMOS technology. We show that employing instruction-dependent dynamic clock adjustment leads on average to an increase in operating speed by 38% or to a reduction in power consumption by 24%, compared to traditional synchronous clocking, which at all times has to respect the worst-case timing identified through static timing analysis.
2015. The Design, Automation and Test in Europe (DATE), Grenoble, France, March 9-13, 2015. p. 381 – 386. DOI : 10.7873/DATE.2015.0303.An Overlap-Contention Free True-Single-Phase Clock Dual-Edge-Triggered Flip-Flop
Dual-edge-triggered (DET) synchronous operation is a very attractive option for low-power, high-performance designs. Compared to conventional single-edge synchronous systems, DET operation is capable of providing the same throughput at half the clock frequency. This can lead to significant power savings on the clock network that is often one of the major contributors to total system power. However, in order to implement DET operation, special registers need to be introduced that sample data on both clock-edges. These registers are more complex than their single-edge counterparts, and often suffer from a certain amount of clock-overlap between the main clock and the internally generated inverted clock. This overlap can cause contention inside the cell and lead to logic failures, especially when operating at scaled power supplies and under process variations that characterize nanometer technologies. This paper presents a novel, static DET flip-flop (DET-FF) with a true-single-phase clock that completely avoids clock overlap hazards by eliminating the need for an inverted clock edge for functionality. The proposed DET FF was implemented in a standard 40nm CMOS technology, showing full functionality at low-voltage operating points, where conventional DET-FFs fail. Under a near-threshold, 500mV supply voltage, the proposed cell also provides a 35% lower CK-to-Q delay and the lowest power-delay-product compared to all considered DET-FF implementations. © 2015 IEEE.
2015. IEEE International Symposium on Circuits and Systems (ISCAS), Lisbon, Portugal, May, 2015. DOI : 10.1109/ISCAS.2015.7169017.Refresh-Free Dynamic Standard-Cell Based Memories: Application to a QC-LDPC Decoder
The area and power consumption of low-density parity check (LDPC) decoders are typically dominated by embedded memories. To alleviate such high memory costs, this paper exploits the fact that all internal memories of a LDPC decoder are frequently updated with new data. These unique memory access statistics are taken advantage of by replacing all static standard-cell based memories (SCMs) of a prior-art LDPC decoder implementation by dynamic SCMs (D-SCMs), which are designed to retain data just long enough to guarantee reliable operation. The use of D-SCMs leads to a 44% reduction in silicon area of the LDPC decoder compared to the use of static SCMs. The low-power LDPC decoder architecture with refresh-free D-SCMs was implemented in a 90nm CMOS process, and silicon measurements show full functionality and an information bit throughput of up to 600 Mbps (as required by the IEEE 802.11n standard). © 2015 IEEE.
2015. IEEE International Symposium on Circuits and Systems (ISCAS), Lisbon, Portugal, May, 2015. DOI : 10.1109/ISCAS.2015.7168911.Mitigating the Impact of Faults in Unreliable Memories For Error-Resilient Applications
Inherently error-resilient applications in areas such as signal processing, machine learning and data analytics provide opportunities for relaxing reliability requirements, and thereby reducing the overheads incurred by conventional error correction schemes. In this paper, we exploit the tolerable imprecision of such applications by designing an energy-efficient fault-mitigation scheme for unreliable memories to meet target yield. The proposed approach uses a bit-shuffling mechanism to isolate faults into bit locations with lower significance. By doing so, the bit-error distribution is skewed towards the low order bits, substantially limiting the output error magnitude. By controlling the granularity of the shuffling, the proposed technique enables trading-off quality for power, area and timing overhead. Compared to error-correction codes, this can reduce the overhead by as much as 83% in power, 89% in area, and 77% in access time when applied to various data mining applications in 28nm process technology.
2015. Design Automation Conference (DAC’15), San Francisco, California, USA, June 7-11, 2015. p. 1 – 6. DOI : 10.1145/2744769.2744871.Concurrent Spectrum Sensing and Transmission for Cognitive Radio using Self-Interference Cancellation
2015. 16th ACM International Symposium on Mobile Ad Hoc Networking and Computing, Hangzhou, China, 22-25 06 2015. p. 407 – 408. DOI : 10.1145/2746285.2764932.Digital Synchronization for Symbol-spaced IEEE802.11ad Gigabit mmWave Systems
A complete digital synchronization architecture for an IEEE 802.11ad compliant 60 GHz receiver is presented. The characteristics of mmWave systems require a holistic view on the problem of parameter estimation, such that not each parameter is dealt with on its own, but in the context of the complete receiver architecture. To this end the proposed synchronization unit covers packet detection, frequency offset compensation, signal-to-interference-plus-noise (SINR) maximization, frame synchronization, and channel estimation. The presented architecture is especially suitable for low-complexity time domain receivers, which are the most power efficient systems for mmWave, but have high demands in terms of synchronization. A novel two step synchronization procedure takes the specific requirements of the employed equalization and detection stages into account, to maximize the overall system performance. Performance is further improved by a heuristic sampling phase alignment mechanism which search the best sampling phase in order to increase the effective SINR in finite length receivers.
2015. 2015 22nd IEEE International Conference on Electronics, Circuits, and Systems (ICECS), Cairo, Egypt, December 06-09, 2015. p. 637 – 640. DOI : 10.1109/ICECS.2015.7440397.Fractionally Spaced Complex Sub-Nyquist Sampling for Multi-Gigabit 60 GHz Wireless Communication
A novel analog front-end architecture based on complex sub-Nyquist sampling for the intermediate frequency (IF) stage of a mmWave receiver is proposed. With this front-end, the use of a wideband hybrid coupler and two half-rate analog-to-digital converters (ADCs) allow for a flexible placement of the IF. It is shown that digital compensation of the impairments introduced by the non-ideal 90 degree hybrid coupler is required to use high modulation orders. Further a digital signal processing (DSP) architecture is presented which performs equalization of a fractionally spaced sub-sampled IF signal in frequency domain (FD) and integrates the compensation of the impairments with low overhead. Based on this DSP architecture a working 60GHz single-carrier link is demonstrated. Measurement results show the feasibility of 256QAM modulated transmission with a bandwidth of up to 1.8 GHz and a resulting raw data rate of 12.8 Gb/s using our frontend architecture with the digital FD compensation.
2015. Midwest Symposium on Circuits and Systems, Fort Collins, Colorado, USA, August 2-5, 2015. DOI : 10.1109/MWSCAS.2015.7282089.An FPGA-based Accelerator for Rapid Simulation of SC Decoding of Polar Codes
2015. 2015 IEEE International Conference on Electronics, Circuits, and Systems, Cairo, Egypt, December 6-9, 2015. p. 633 – 636. DOI : 10.1109/ICECS.2015.7440396.Energy versus Data Integrity Trade-Offs in Embedded High-Density Logic Compatible Dynamic Memories
Current variation aware design methodologies, tuned for worst-case scenarios, are becoming increasingly pessimistic from the perspective of power and performance. A good example of such pessimism is setting the refresh rate of DRAMs according to the worst-case access statistics, thereby resulting in very frequent refresh cycles, which are responsible for the majority of the standby power consumption of these memories. However, such a high refresh rate may not be required, either due to extremely low probability of the actual occurrence of such a worst-case, or due to the inherent error resilient nature of many applications that can tolerate a certain number of potential failures. In this paper, we exploit and quantify the possibilities that exist in dynamic memory design by shifting to the so-called approximate computing paradigm in order to save power and enhance yield at no cost. The statistical characteristics of the retention time in dynamic memories were revealed by studying a fabricated 2kb CMOS compatible embedded DRAM (eDRAM) memory array based on gain-cells. Measurements show that up to 73% of the retention power can be saved by altering the refresh time and setting it such that a small number of failures is allowed. % can save up to 3.8$\times$ of the retention power
2015. DATE 2015, Grenoble, France, March 9-13, 2015. p. 489 – 494. DOI : 10.7873/DATE.2015.0783.Posters
Circuits and Techniques for Dynamic Timing Monitoring in Microprocessors
Nanotera Annual Meeting 2015, Bern, Switzerland, May 5, 2015.Patents
Method and apparatus for low complexity spectral analysis of bio-signals
A method and device for reducing the computational complexity of a processing algorithm, of a discrete signal, in particular of the spectral estimation and analysis of bio-signals, with minimum or no quality loss, which comprises steps of (a) choosing a domain, such that transforming the signal to the chosen domain results to an approximately sparse representation, wherein at least part of the output data vector has zero or low magnitude elements; (b) converting the original signal in the domain chosen in step (a) through a mathematical transform consisting of arithmetic operations resulting in a vector of output data; (c) reformulating the processing algorithm of the original signal in the original domain into a modified algorithm consisting of equivalent arithmetic operations in the domain chosen in step (a) to yield the expected result with the expected quality quantified in terms of a suitable application metric; (d) combining the mathematical transform of step (b) and the equivalent mathematical operations introduced in step (c) for obtaining the expected result within the original domain with the expected quality; (e) selecting a threshold value based on the difference in the mean magnitude value of the elements of the output data vector of the transform said in step (b) and the preferred complexity reduction and degree of output quality loss that can be tolerated in the expected result within the target application; (f) pruning a number of elements the magnitude of which is less than the threshold value selected in step (e); and/or eliminating arithmetic operations associated with the pruned elements of step (f) either in the mathematical transform of step (b) and/or in the equivalent algorithm of step (c).
US9760536; US2015220486; EP2884884; WO2014027329.
2015.Student Projects
Power analysis and optimization of on-board processing for the EFM32 microprocessor
2015.Automated Performance Characterization of Dynamic Clock Adjustment Techniques on an OpenRISC ISS
2015.2014
Journal Articles
Multi-level wordline driver for robust SRAM design in nano-scale CMOS technology
In this paper, a multi-level wordline driver scheme is presented to improve 6T-SRAM read and write stability. The proposed wordline driver generates a shaped pulse during the read mode and a boosted wordline during the write mode. During read, the shaped pulse is tuned at nominal voltage for a short period of time, whereas for the remaining access time, the wordline voltage is reduced to save the power consumption of the cell. This shaped wordline pulse results in improved read noise margin without any degradation in access time for small wordline load. The improvement is explained by examining the dynamic and nonlinear behavior of the SRAM cell. Furthermore, during the hold mode, for a short time (depending on the size of boosting capacitance), wordline voltage becomes negative and charges up to zero after a specific time that results in a lower leakage current compared to conventional SRAM. The proposed technique results in at least 2 x improvement in read noise margin while it improves write margin by 3 x for lower supply voltages than 0.7 V. The leakage power for the proposed SRAM is reduced by 2% while the total power is improved by 3% in the worst case scenario for an SRAM array. The main advantage of the proposed wordline driver is the improvement of dynamic noise margin with less than 2.5% penalty in area. TSMC 65 nm technology models are used for simulations. (C) 2013 Elsevier Ltd. All rights reserved.
Microelectronics Journal. 2014. Vol. 45, num. 1, p. 23 – 34. DOI : 10.1016/j.mejo.2013.09.009.Density Evolution for Min-Sum Decoding of LDPC Codes Under Unreliable Message Storage
We analyze the performance of quantized min-sum decoding of low-density parity-check codes under unreliable message storage. To this end, we introduce a simple bit-level error model and show that decoder symmetry is preserved under this model. Subsequently, we formulate the corresponding density evolution equations to predict the average bit error probability in the limit of infinite blocklength. We present numerical threshold results and we show that using more quantization bits is not always beneficial in the context of faulty decoders.
IEEE Communications Letters. 2014. Vol. 18, num. 5, p. 849 – 852. DOI : 10.1109/Lcomm.2014.030714.132830.A Lattice Reduction-Aided MIMO Channel Equalizer in 90 nm CMOS Achieving 720 Mb/s
In this paper, a VLSI implementation of a complete MIMO channel equalization ASIC based on lattice reduction-aided linear detection is presented. The architecture performs preprocessing steps at channel rate and low-complexity linear data detection at symbol rate. Preprocessing is based on Seysen’s algorithm for lattice reduction. We present algorithmic improvements of the lattice reduction preprocessing in terms of area and throughput of the VLSI implementation with minor impact on the error-rate. Due to the low-complexity implementation of the lattice reduction-aided data detection stage, our architecture is able to achieve very low power in typical packet-based MIMO wireless data transmission scenarios. The final 90 nm CMOS ASIC achieves an energy efficiency for the detection of 24 pJ/bit at a throughput of 720 Mbps with near-optimal error-rate performance.
IEEE Transactions on Circuits and Systems I: Regular Papers. 2014. Vol. 61, num. 6, p. 1860 – 1871. DOI : 10.1109/TCSI.2013.2295027.A Low-Power Low-Cost 24 GHz RFID Tag With a C-Flash Based Embedded Memory
The key factor in widespread adoption of Radio Frequency Identification (RFID) technology is tag cost minimization. This paper presents the first low-cost, ultra-low power, passive RFID tag, fully integrated on a single substrate in a standard CMOS process. The system combines a 24 GHz, dual on-chip antenna, RF front-end, and a C-Flash based, rewritable, non-volatile memory module to achieve full on-chip system integration. The complete system was designed and fabricated in the TowerJazz 0.18 mu m CMOS technology without any additional mask adders. By embedding the RF, memory, and digital components together upon a single substrate in a standard digital process, the low-cost aspirations of the “5-cent RFID tag” become feasible. Design considerations, analysis, circuit implementations, and measurement results are presented. The entire system was fabricated on a 3.6 mm x 1.6 mm (6.9 mm(2)) die with the integrated antennas comprising 82% of the silicon area. The total read power was measured to be 13.2 mu W, which is sufficiently supplied by the on-chip energy harvesting unit.
Ieee Journal Of Solid-State Circuits. 2014. Vol. 49, num. 9, p. 1942 – 1957. DOI : 10.1109/Jssc.2014.2323352.A Fast and Versatile Quantum Key Distribution System with Hardware Key Distillation and Wavelength Multiplexing
We present a compactly integrated, 625 MHz clocked coherent one-way quantum key distribution system which continuously distributes secret keys over an optical fibre link. To support high secret key rates, we implemented a fast hardware key distillation engine which allows for key distillation rates up to 4 Mbps in real time. The system employs wavelength multiplexing in order to run over only a single optical fibre. Using fast gated InGaAs single photon detectors, we reliably distribute secret keys with a rate above 21 kbps over 25 km of optical fibre. We optimized the system considering a security analysis that respects finite-key-size effects, authentication costs and system errors for a security parameter of epsilon_QKD = 4 x 10^−9.
New Journal of Physics. 2014. Vol. 16, p. 013047. DOI : 10.1088/1367-2630/16/1/013047.Energy/Reliability Trade-Offs in Low-Voltage ReRAM-Based Non-Volatile Flip-Flop Design
The total power budget of Ultra-Low Power (ULP) VLSI Systems-on-Chip (SoCs) is often dominated by the leakage power of embedded memories as well as status registers. On the one hand, supply voltage scaling down to the near-threshold (near-VT) or even to the subthreshold (sub-VT) domain is a commonly used, efficient technique to reduce both leakage power and active energy dissipation. On the other hand, emerging CMOS-compatible device technologies such as Resistive Memories (ReRAMs) enable non-volatile, on-chip data storage and zero-leakage sleep periods. For the first time, we present and compare ReRAM-based Non-Volatile Flip-Flop (NVFF) topologies which are optimized for low-voltage operation (including near-VT and sub-VT operation). Three low-voltage NVFF circuit topologies are proposed and evaluated in terms of energy dissipation and reliability. Using topologies with two complementary programmed ReRAM devices, Monte Carlo simulations accounting for parametric variations confirm reliable data restore operation from the ReRAM devices at a sub- voltage as low as 400 mV. A topology using a single ReRAM device exhibits lower write energy, but requires a near- voltage for robust read. Energy characterization is performed at nominal, near-VT , and sub-VT supply voltages. The minimum energy point is reached for near-VT read operation with a total read+write energy of 735 fJ.
IEEE Transactions on Circuits and Systems Part 1 Regular Papers. 2014. Vol. 61, num. 11, p. 3155 – 3164. DOI : 10.1109/TCSI.2014.2334891.Replica Technique for Adaptive Refresh Timing of Gain-Cell-Embedded DRAM
Gain cells have recently been shown to be a viable alternative to static random access memory in low-power applications due to their low leakage currents and high density. The primary component of power consumption in these arrays is the dynamic power consumed during periodic refresh operations. Refresh timing is traditionally set according to a worst-case evaluation of retention time under extreme process variations, and worst-case access statistics, leading to frequent power-hungry refresh cycles. In this brief we present a replica technique for automatically tracking the retention time of a gain-cell-embedded dynamic-random-access-memory macrocell according to process variations and operating statistics, thereby reducing the data retention power of the array. A 2-kb array was designed and fabricated in a mature 0.18-mu m CMOS process, appropriate for integration in ultralow power applications, such as biomedical sensors. Measurements show efficient retention time tracking across a range of supply voltages and access statistics, lowering the refresh frequency by more than 5x, as compared with traditional worst-case design.
IEEE Transactions on Circuits and Systems II: Express Briefs. 2014. Vol. 61, num. 4, p. 259 – 263. DOI : 10.1109/Tcsii.2014.2305016.Hardware Architecture for List Successive Cancellation Decoding of Polar Codes
This brief presents a hardware architecture and algorithmic improvements for list successive cancellation (SC) decoding of polar codes. More specifically, we show how to completely avoid copying of the likelihoods, which is algorithmically the most cumbersome part of list SC decoding. The hardware architecture was synthesized for a blocklength of N = 1024 bits and list sizes L = 2, 4 using a UMC 90 nm VLSI technology. The resulting decoder can achieve a coded throughput of 181 Mb/s at a frequency of 459 MHz.
IEEE Transactions on Circuits and Systems II: Express Briefs. 2014. Vol. 61, num. 8, p. 609 – 613. DOI : 10.1109/Tcsii.2014.2327336.Energy Efficiency through Significance-Based Computing
An extension of approximate computing, significance-based computing exploits applications’ inherent error resiliency and offers a new structural paradigm that strategically relaxes full computational precision to provide significant energy savings with minimal performance degradation.
Computer. 2014. Vol. 47, num. 7, p. 82 – 85. DOI : 10.1109/MC.2014.182.Conference Papers
Cross Layer Energy-Efficiency Optimization For Cognitive Radio Transceivers
Designing energy-efficient cognitive radio transceivers requires joint optimization of medium access control and the physical layer implementation. In this paper we show an energy efficiency optimization strategy for IEEE 802.11n compliant transceivers in terms of energy consumed by the receiver per successfully received bit. To this end, we propose and explore several modifications of a conventional physical layer implementation, all of which target energy proportional behavior. The proposed modifications intentionally include operation modes and algorithm choices that are suboptimal with respect to throughput and error-rate performance. Yet, we show how (under ideal conditions) the rate adaptation at the medium access control layer can exploit these modifications to achieve superior energy efficiency that is 44% below that of a rate adaptation targeting only maximum goodput.
2014. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). p. 3928 – 3932. DOI : 10.1109/ICASSP.2014.6854338.Correlation Based Phase Noise Compensation in 60 GHz Wireless Systems
2014. 2014 IEEE 28-th Convention of Electrical and Electronics Engineers in Israel, Eilat, Israel, December, 3-5, 2014. DOI : 10.1109/EEEI.2014.7005755.True MIMO transmission using a single RF-chain and antenna: recent developments
2014. Eucap 2014, 8th European Conf. on Antennas and Prop., The Hague, Netherlands, 2014.Robust Asynchronous Indoor Localization Using Led Lighting
We propose a low-cost system for indoor self-localization of mobile devices using modulated LED ceiling lamps that are fully autonomous and broadcast their identifiers without any synchronization. The proposed self-localization method is designed to handle this lack of synchronization as well as the possibility of blocked line-of-sight connections or severe attenuation in real-world environments. This robustness is achieved by applying a suitable Bayesian signal model and by taking into account the inherent sparsity in detecting the concurrently visible lamps. The proposed estimator of the location approximates optimal Bayesian estimation while maintaining low complexity. Simulation results confirm a significant gain in performance compared to a classical matched-filter approach.
2014. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Florence, ITALY, MAY 04-09, 2014. p. 1866 – 1870. DOI : 10.1109/ICASSP.2014.6853922.A Quality-Scalable and Energy-Efficient Approach for Spectral Analysis of Heart Rate Variability
Today there is a growing interest in the integration of health monitoring applications in portable devices necessitating the development of methods that improve the energy efficiency of such systems. In this paper, we present a systematic approach that enables energy-quality trade-offs in spectral analysis systems for bio-signals, which are useful in monitoring various health conditions as those associated with the heart-rate. To enable such trade-offs, the processed signals are expressed initially in a basis in which significant components that carry most of the relevant information can be easily distinguished from the parts that influence the output to a lesser extent. Such a classification allows the pruning of operations associated with the less significant signal components leading to power savings with minor quality loss since only less useful parts are pruned under the given requirements. To exploit the attributes of the modified spectral analysis system, thresholding rules are determined and adopted at design- and run-time, allowing the static or dynamic pruning of less-useful operations based on the accuracy and energy requirements. The proposed algorithm is implemented on a typical sensor node simulator and results show up-to 82% energy savings when static pruning is combined with voltage and frequency scaling, compared to the conventional algorithm in which such trade-offs were not available. In addition, experiments with numerous cardiac samples of various patients show that such energy savings come with a 4.9% average accuracy loss, which does not affect the system detection capability of sinus-arrhythmia which was used as a test case.
2014. Design Automation & Test in Europe (DATE), Dresden, Germany, DOI : 10.7873/DATE.2014.184.Enabling Complexity-Performance Trade-Offs for Successive Cancellation Decoding of Polar Codes
Polar codes are one of the most recent advancements in coding theory and they have attracted significant interest. While they are provably capacity achieving over various channels, they have seen limited practical applications. Unfortunately, the successive nature of successive cancellation based decoders hinders fine-grained adaptation of the decoding complexity to design constraints and operating conditions. In this paper, we propose a systematic method for enabling complexity-performance trade-offs by constructing polar codes based on an optimization problem which minimizes the complexity under a suitably defined mutual information based performance constraint. Moreover, a low-complexity greedy algorithm is proposed in order to solve the optimization problem efficiently for very large code lengths.
2014. IEEE International Symposium on Information Theory, Honolulu, Hawaii, USA, June 29 – July 4, 2014. p. 2977 – 2981. DOI : 10.1109/ISIT.2014.6875380.A Wireless Body Sensor Network For Activity Monitoring With Low Transmission Overhead
Activity recognition has been a research field of high interest over the last years, and it finds application in the medical domain, as well as personal healthcare monitoring during daily home- and sports-activities. With the aim of producing minimum discomfort while performing supervision of subjects, miniaturized networks of low-power wireless nodes are typically deployed on the body to gather and transmit physiological data, thus forming a Wireless Body Sensor Network (WBSN). In this work, we propose a WBSN for online activity monitoring, which combines the sensing capabilities of wearable nodes and the high computational resources of modern smartphones. The proposed solution provides different tradeoffs between classification accuracy and energy consumption, thanks to different workloads assigned to the nodes and to the mobile phone in different network configurations. In particular, our WBSN is able to achieve very high activity recognition accuracies (up to 97.2%) on multiple subjects, while significantly reducing the sampling frequency and the volume of transmitted data with respect to other state-of-the art solutions.
2014. The 12th IEEE International Conference on Embedded and Ubiquitous Computing, Milan, 25-29.08.2014. p. 265 – 272. DOI : 10.1109/EUC.2014.46.Variability-Aware Design Space Exploration of Embedded Memories
With scaling of process technologies and worsening of process variations, embedded memories are susceptible to a large number of failure mechanisms making it hard to achieve high yield. In this paper, by bringing together architecture and circuit-level exploration tools, we analyse the impact of process variations on static random access memory (SRAM) cell stability and determine the impact of SRAM failures on memory functional yield. We then detail the importance of repair mechanisms such as error correcting codes (ECC) and redundancy on improving yield subject to constraints set on power and area. Finally, we show that a design paradigm orthogonal to traditional repair mechanisms involving redefinition of the yield criterion by accepting memories with failures is a promising candidate for improving yield without incurring additional overheads.
2014. 28th IEEE Convention of Electrical and Electronics Engineers in Israel, Eilat, Israel, December 3-5, 2014. DOI : 10.1109/EEEI.2014.7005798.Dynamic stability and noise margins of SRAM arrays in nanoscaled technologies
SRAM stability is one of the primary bottlenecks of current VLSI system design, and the unequivocal supply voltage scaling limiter. Static noise margin metrics have long been the de-facto standard for measuring this stability and estimating the yield of SRAM arrays. However, in modern process technologies, under scaled supply voltages and increased process variations, these traditional metrics are no longer sufficient. Recent research has analyzed the dynamic behavior and stability of SRAM circuits, leading to dynamic stability metrics and dynamic noise margin definition. This paper provides a brief overview of the limitations of static noise margin metrics and the resulting dynamic stability and noise margin concepts that have been proposed to overcome them.
2014. 2014 IEEE Faible Tension Faible Consommation (FTFC), Monaco, Monaco, 4-6 May 2014. p. 1 – 5. DOI : 10.1109/FTFC.2014.6828617.Single event upset mitigation in low power SRAM design
Logic compatible gain cell (GC)-embedded DRAM (eDRAM) arrays are considered an alternative to SRAM due to their small size, nonratioed operation, low static leakage, and two-port functionality. However, traditional GC-eDRAM implementations require boosted control signals in order to write full voltage levels to the cell to reduce the refresh rate and shorten access times. These boosted levels require either an extra power supply or on-chip charge pumps, as well as nontrivial level shifting and toleration of high voltage levels. In this brief, we present a novel, logic compatible, 3T GC-eDRAM bitcell that operates with a single-supply voltage and provides superior write capability to the conventional GC structures. The proposed circuit is demonstrated with a 2-kb memory macro that was designed and fabricated in a mature 0.18-μm CMOS process, targeted at low-power, energy-efficient applications. The test array is powered with a single supply of 900 mV, showing a 0.8-ms worst case retention time, a 1.3-ns write-access time, and a 2.4-pW/bit retention power. The proposed topology provides a bitcell area reduction of 43%, as compared with a redrawn 6-transistor SRAM in the same technology, and an overall macro area reduction of 67% including peripherals.
2014. 2014 IEEE 28th Convention of Electrical & Electronics Engineers in Israel (IEEEI), Eilat, Israel, 3-5 December 2014. p. 1 – 5. DOI : 10.1109/EEEI.2014.7005796.4T Gain-Cell with internal-feedback for ultra-low retention power at scaled CMOS nodes
Gain-Cell embedded DRAM (GC-eDRAM) has recently been recognized as a possible alternative to traditional SRAM. While GC-eDRAM inherently provides high-density, low-leakage, low-voltage, and 2-ported operation, its limited retention time requires periodic, power-hungry refresh cycles. This drawback is further enhanced at scaled technologies, where increased subthreshold leakage currents and decreased in-cell storage capacitances result in faster data deterioration. In this paper, we present a novel 4T GC-eDRAM bitcell that utilizes an internal feedback mechanism to significantly increase the data retention time in scaled CMOS technologies. A 2 kb memory macro was implemented in a low-power 65nm CMOS technology, displaying an over 3× improvement in retention time over the best previous publication at this node. The resulting array displays a nearly 5× reduction in retention power (despite the refresh power component) with a 40% reduction in bitcell area, as compared to a standard 6T SRAM.
2014. 2014 IEEE International Symposium on Circuits and Systems (ISCAS), Melbourne VIC, Australia, 1-5 June 2014. p. 2177 – 2180. DOI : 10.1109/ISCAS.2014.6865600.LLR-based Successive Cancellation List Decoding of Polar Codes
We present an LLR-based implementation of the successive cancellation list (SCL) decoder. To this end, we associate each decoding path with a metric which (i) is a monotone function of the path’s likelihood and (ii) can be computed efficiently from the channel LLRs. The LLR-based formulation leads to a more efficient hardware implementation of the decoder compared to the known log-likelihood based implementation. Synthesis results for an SCL decoder with block-length of N = 1024 and list sizes of L = 2 and L = 4 confirm that the LLR-based decoder has considerable area and operating frequency advantages in the orders of 50% and 30%, respectively.
2014. IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2014), Florence, Italy, May 4-9, 2014. p. 3903 – 3907. DOI : 10.1109/ICASSP.2014.6854333.Data Compression via Logic Synthesis
Nowadays, most software and hardware applications are committed to reduce the footprint and resource usage of data. In this general context, lossless data compression is a beneficial technique that encodes information using fewer (or at most equal number of) bits as compared to the original representation. A traditional compression flow consists of two phases: data decorrelation and entropy encoding. Data decorrelation, also called entropy reduction, aims at reducing the autocorrelation of the input data stream to be compressed in order to enhance the efficiency of entropy encoding. Entropy encoding reduces the size of the previously decorrelated data by using techniques such as Huffman coding, arithmetic coding, and others. When the data decorrelation is optimal, entropy encoding produces the strongest lossless compression possible. While efficient solutions for entropy encoding exist, data decorrelation is still a challenging problem limiting ultimate lossless compression opportunities. In this paper, we use logic synthesis to remove redundancy in binary data aiming to unlock the full potential of lossless compression. Embedded in a complete lossless compression flow, our logic synthesis based methodology is capable to identify the underlying function correlating a data set. Experimental results on data sets deriving from different causal processes show that the proposed approach achieves the highest compression ratio compared to state-of-art compression tools such as ZIP, bzip2 and 7zip.
2014. 19th Asia and South Pacific Design Automation Conference (ASP-DAC 2014), Singapore, January 20-23, 2014. p. 628 – 633. DOI : 10.1109/ASPDAC.2014.6742961.Theses
Novel Approaches Toward Area- and Energy-Efficient Embedded Memories
Lausanne, EPFL, 2014.Energy Efficient VLSI Circuits for MIMO-WLAN
Mobile communication – anytime, anywhere access to data and communication services – has been continuously increasing since the operation of the first wireless communication link by Guglielmo Marconi. The demand for higher data rates, despite the limited bandwidth, led to the development of multiple-input multiple-output (MIMO) communication which is often combined with orthogonal frequency division multiplexing (OFDM). Together, these two techniques achieve a high bandwidth efficiency. Unfortunately, techniques such as MIMO-OFDM significantly increase the signal processing complexity of transceivers. While fast improvements in the integrated circuit (IC) technology enabled to implement more signal processing complexity per chip, large efforts had and have to be done for novel algorithms as well as for efficient very large scaled integration (VLSI) architectures in order to meet today’s and tomorrow’s requirements for mobile wireless communication systems. In this thesis, we will present architectures and VLSI implementations of complete physical (PHY) layer application specific integrated circuits (ASICs) under the constraints imposed by an industrial wireless communication standard. Contrary to many other publications, we do not elaborate individual components of a MIMO-OFDM communication system stand-alone, but in the context of the complete PHY layer ASIC. We will investigate the performance of several MIMO detectors and the corresponding preprocessing circuits, being integrated into the entire PHY layer ASIC, in terms of achievable error-rate, power consumption, and area requirement. Finally, we will assemble the results from the proposed PHY layer implementations in order to enhance the energy efficiency of a transceiver. To this end, we propose a cross-layer optimization of PHY layer and medium access control (MAC) layer.
Lausanne, EPFL, 2014.Posters
Cross-Layer Inexact Design for Low-Power Applications
Approximate and error tolerant circuits are a radical new approach to trade calculation accuracy for better speed, power, area and yield. The IcySoC project platform revisits low-power and low-voltage VLSI design through a cross-layer combined inexact design framework.
NanoTera Annual Meeting, Lausanne, Switzerland,Restructuring of Arithmetic Circuits with Biconditional Binary Decision Diagrams
Biconditional Binary Decision Diagrams (BBDDs) are a novel class of canonical binary decision diagrams where the branching condition, and its associated logic expansion is biconditional on two variables. In this demonstration we use an efficient BBDD manipulation package as front-end to a commercial synthesis tool to restructure arithmetic operations in critical components of telecommunication circuits. We show that our approach meets tight timing constraints otherwise beyond the capabilities of traditional synthesis methods.
University Booth at DATE 2014, Dresden, Germany, March 24-28, 2014.Patents
Ultra-Low Power Multicore Architecture For Parallel Biomedical Signal Processing
A multi-core architecture with ultra-low power consumption is needed for a wide variety of applications, especially in the bio-medical domain. In this patent, an ultra-low power multi-core architecture is presented: it is composed of one or more cores, several (one or more] shared multi-banked instruction and data memories, and flexible crossbar interconnects. The interconnect includes a selective broadcasting mechanism, enabling coordinated multiple accesses to the shared memories, thus energy savings in the memory hierarchy and in the interconnects. The memory hierarchy enables power gating of the unused banks to lower leakage power. The core instruction set of the novel architecture has been customized to exploit the specific features of bio-signal events, as well as the highly parallel computation opportunities of bio-signal processing characteristics. In addition to near threshold computing, the proposed architecture also exploits other advanced low-power features, which lead to further energy savings. The architecture has a synchronization method and unit that allows power- efficient handling of data dependencies when code is parallelized across the multiple processing cores.
WO2013136259; WO2013136259.
2014.Student Projects
Low Power Wake-up Receiver
With more devices becoming mobile, power consumption of communication becomes crucial. Wake-up receivers present an energy-ecient way of detecting incoming transmissions while at the same time the main radio can be fully powered down. After detection the rest of the circuit is activated in order to receive the transmission. Since the digital correlator is the main power consumer of a wake-up receiver, this work compares dierent correlator architectures in terms of their power performance. The correlator architectures are implemented in MATLAB for an analysis of their detection performance and later on in VHDL. A full front and back end implementation is made in order to extract the power performance of the architectures. Large parts of the work ow are automated, allowing to see immediately the impact of multiple parameters such as the sample resolution. Using an equation-based approach, the obtained power estimation is scaled into a sub-threshold region allowing to compare the power performance at very low voltage. The result of our architecture analysis shows that the positioning of the ip- ops is crucial to the power performance: combinational paths should be kept short, in contrast adding additional register to increase speed is benecial. Using large and complex logic decreases the power performance. For the analyzed correlator designs, the inversed structure splitting up the adder chain, which sums the correlation results, performed best. If noise is no issue, using a sample resolution of 1 bit leads to further improvement. Comparing the architectures in the sub-threshold region, fast architectures are less power benecial. A pipelined structure seems to oer a good trade-o between combinational circuit and amount of registers used and has therefore the lowest power consumption.
2014.Retention Time Characterization of Commercial DRAM Modules Using an FPGA-based Test Platform
2014.RSS Range Estimation for Indoor Localization Using LED Lighting
2014.2013
Journal Articles
Impact of body biasing on the retention time of gain-cell memories
Gain-cell-based embedded dynamic random-access memory (DRAMs) are a potential high-density alternative to mainstream static random-access memory (SRAM). However, the limited data retention time of these dynamic bitcells results in the need for power-consuming periodic refresh cycles. This Letter measures the impact of body biasing as a control factor to improve the retention time of a 2 kb memory block, and also examines the distribution of the retention time across the entire gain-cell array. The concept is demonstrated through silicon measurements of a test chip manufactured in a logic-compatible 0.18 μm CMOS process. Although there is a large retention time spread across the measured 2 kb gain-cell array, the minimum, average and maximum retention times are all improved by up to two orders of magnitude when sweeping the body voltage over a range of 375 mV.
The Journal of Engineering. 2013. num. 8, p. 19 – 22. DOI : 10.1049/joe.2013.0057.Exploration of Sub-VT and Near-VT 2T Gain-Cell Memories for Ultra-Low Power Applications under Technology Scaling
Ultra-low power applications often require several kb of embedded memory and are typically operated at the lowest possible operating voltage (VDD) to minimize both dynamic and static power consumption. Embedded memories can easily dominate the overall silicon area of these systems, and their leakage currents often dominate the total power consumption. Gain-cell based embedded DRAM arrays provide a high-density, low-leakage alternative to SRAM for such systems; however, they are typically designed for operation at nominal or only slightly scaled supply voltages. This paper presents a gain-cell array which, for the first time, targets aggressively scaled supply voltages, down into the subthreshold (sub-VT) domain. Minimum VDD design of gain-cell arrays is evaluated in light of technology scaling, considering both a mature 0.18 μm CMOS node, as well as a scaled 40 nm node. We first analyze the trade-offs that characterize the bitcell design in both nodes, arriving at a best-practice design methodology for both mature and scaled technologies. Following this analysis, we propose full gain-cell arrays for each of the nodes, operated at a minimum VDD. We find that an 0.18 μm gain-cell array can be robustly operated at a sub-VT supply voltage of 400mV, providing read/write availability over 99% of the time, despite refresh cycles. This is demonstrated on a 2 kb array, operated at 1 MHz, exhibiting full functionality under parametric variations. As opposed to sub-VT operation at the mature node, we find that the scaled 40 nm node requires a near-threshold 600mV supply to achieve at least 97% read/write availability due to higher leakage currents that limit the bitcell’s retention time. Monte Carlo simulations show that a 600mV 2 kb 40 nm gain-cell array is fully functional at frequencies higher than 50 MHz.
Journal of Low Power Electronics and Applications. 2013. Vol. 3, num. 2, p. 54 – 72. DOI : 10.3390/jlpea3020054.Conference Papers
Efficient VLSI Implementation of Reduced-State Sequence Estimation for Wireless Communications
Modern wireless communication systems require efficient channel equalizer implementations. This paper explores the design space of reduced-state sequence estimation (RSSE). We show how the concept of pre-computation can be applied to greatly reduce computational complexity, such that efficient RSSE architectures can be derived. As a proof of concept, an RSSE was implemented in dedicated hardware, that achieves a 1.6 times higher hardware efficiency when compared to prior art.
2013. 38th International Conference on Acoustics, Speech, and Signal Processing (ICASSP), Vancouver, Canada, May 26-31, 2013. p. 2528 – 2532. DOI : 10.1109/ICASSP.2013.6638111.On Self-interference Suppression Methods for Low-complexity Full-duplex MIMO
Full-duplex wireless communication offers improved spectral efficiency, as well as more efficient relaying and medium access, but requires suppression of self-interference. In this paper we analyze the existing methods for active RF suppression and use the ”Rice architecture” for its low complexity and favorable scaling when applied to multi-antenna systems. We analyze the effects of the different sources of self-interference and quantify the potential for further suppression (genie-aided suppression). Our single-chain implementation using a circulator achieves −48 dB of active RF suppression, but only −66 dB of total suppression in the analog domain. On the other hand, our single-chain implementation using separate antennae reaches −85 dB of total analog suppression, thus reducing the self-interference to the noise floor. Extending these setups, we present a low complexity implementation of a 2 × 2 full-duplex MIMO node, which achieves even higher suppression than the single-chain counterparts.
2013. Asilomar Conference on Signals, Systems, and Computers, Pacific Grove, CA, November 3-6, 2013. p. 992 – 997. DOI : 10.1109/ACSSC.2013.6810439.A ReRAM-Based Non-Volatile Flip-Flop with Sub-VT Read and CMOS Voltage-Compatible Write
The total power budget of Ultra-Low Power (ULP) VLSI Systems-on-Chip (SoCs) is often dominated by the leakage power of embedded memories and pipeline registers, which typically cannot be power-gated during sleep periods as they need to retain data and program state, respectively. On the one hand, supply voltage scaling down to the near-threshold (near-VT) or even to the sub-threshold (sub-VT) domain is a commonly used, efficient technique to reduce both leakage power and active en- ergy dissipation. On the other hand, emerging CMOS-compatible device technologies such as Resistive Memories (ReRAMs) enable non-volatile, on-chip data storage and zero-leakage sleep periods. For the first time, we present a ReRAM-based non-volatile flip- flop which is optimized for sub-VT operation. Writing to the ReRAM devices works with a CMOS-compatible supply voltage. Thanks to near-VT and sub-VT operation and as compared to the write energy, which depends on the ReRAM technology, the read consumes only 5.4% of the total read+write energy. Monte Carlo simulations accounting for parametric variations in both the MOS transistors and the ReRAM devices confirm reliable data restore operation from the ReRAM devices at a sub-VT voltage as low as 400mV, and a standard deviation of up to 5% of the nominal value of the ReRAM resistance.
2013. 11th IEEE International NEWCAS Conference, Paris, France, June 16-19, 2013. DOI : 10.1109/NEWCAS.2013.6573586.Fast and Accurate BER Estimation Methodology for I/O Links based on Extreme Value Theory
2013. Design, Automation & Test in Europe Conference (DATE 2013), Grenoble, France, March 18-22, 2013. p. 503 – 508. DOI : 10.7873/DATE.2013.114.Synchronizing Code Execution on Ultra-Low-Power Embedded Multi-Channel Signal Analysis Platforms
Embedded biosignal analysis involves a considerable amount of parallel computations, which can be exploited by employing low-voltage and ultra-low-power (ULP) parallel computing architectures. By allowing data and instruction broadcasting, single instruction multiple data (SIMD) processing paradigm enables considerable power savings and application speedup, in turn allowing for a lower voltage supply for a given workload. The state-of-the-art multi-core architectures for biosignal analysis however lack a bare, yet smart, synchronization technique among the cores, allowing lockstep execution of algorithm parts that can be performed using the SIMD, even in the presence of data-dependent execution flows. In this paper, we propose a lightweight synchronization technique to enhance an ULP multi-core processor, resulting in improved energy efficiency through lockstep SIMD execution. Our results show that the proposed improvements accomplish tangible power savings, up to 64% for an 8-core system operating at a workload of 89 MOps/s while exploiting voltage scaling.
2013. The Design, Automation and Test in Europe (DATE), 2013, Grenoble, France, p. 396 – 399. DOI : 10.7873/DATE.2013.090.Block-Floating-Point Enhanced MMSE Filter Matrix Computation for MIMO-OFDM Communication Systems
n this paper we present an architecture for an MMSE filter matrix computation unit for signal detection in MIMO-OFDM communication systems. We propose to compute the required matrix inverse based on a Cholesky decomposition, followed by a Gauss-Jordan matrix inversion of the resulting triangular matrix. The high dynamic range required by this approach is traditionally conquered with custom floating-point formats or with fixed-point number representations with a large number of bits. We show in this paper that a block-floating- point scheme with only two normalization steps throughout the computation of the MMSE filter matrix is sufficient to achieve a BER performance close to that of a double precision floating- point implementation for a MIMO-OFDM systems with 64-QAM modulation. The corresponding circuit complexity is superior to that of a pure fixed-point implementation.
2013. 2013 IEEE International Conference on Electronics, Circuits, and Systems, Abu Dhabi, UAE, December 8-11, 2013. p. 787 – 790. DOI : 10.1109/ICECS.2013.6815532.FireBird: PowerPC e200 Based SoC for High Temperature Operation
PowerPC Architecture microcontrollers are commonly used in embedded applications. In this work we present FireBird, the first PowerPC based SoC for reliable operation beyond 200C. It is a significant challenge to design SoCs for reliable operation at high temperatures, due to increased static leakage current, reduced carrier mobility, and increased electromigration. To alleviate the consequences of high temperatures, this paper proposes to customize a PowerPC e200 based SoC by using a dynamically reconfigurable clock frequency, exhaustive clock gating, and electromigration-resistant power supply rings. A 20x9mm2 chip implementing this design has been fabricated in 0.35m CMOS technology. The custom testing procedure showed the expected maximum operating frequency reduction from 43MHz at room-temperature to 33MHz at 200C, which illustrates the importance of an adaptable clock frequency under temperature gradients. At 200C, the maximum power dissipation at 3.3V supply voltage was 1.75W and the idle state static leakage current was 0.2A. Silicon measurements proved that this design outclasses PowerPC based SoCs available in the high-temperature microcontrollers market which are not operational at temperatures above 125C.
2013. IEEE Custom Integrated Circuits Conference (CICC), San Jose, California, USA, September 23-25, 2013. DOI : 10.1109/CICC.2013.6658519.A Parallelized Layered QC-LDPC Decoder for IEEE 802.11ad
We present a doubly parallelized layered quasi-cyclic low density parity-check decoder for the emerging IEEE 802.11ad multigigabit wireless standard. The decoding algorithm is equivalent to a nonparallelized layered decoder and, thus, retains its favorable convergence characteristics, which are known to be superior to those of flooding schedule based decoders. The proposed architecture was synthesized using a TSMC 40 nm CMOS technology, resulting in a cell area of 0.18 mm2 and a clock frequency of 850 MHz. At this clock frequency, the decoder achieves a coded throughput of 3.12 Gbps, thus meeting the throughput requirements when using both the mandatory BPSK modulation and the optional QPSK modulation.
2013. 11th IEEE International NEWCAS Conference, Paris, France, June 16-19, 2013. DOI : 10.1109/NEWCAS.2013.6573590.Theses
Energy-Aware Processing Platform Exploration for Embedded Biosignal Analysis
According to the World Health Organization, lifestyle-related diseases, e.g., cardiovascular diseases are the major cause of mortality worldwide. An accurate and continuous medical supervision is highly required for diagnosis and treatment of such diseases. Our traditional healthcare delivery systems, however can’t cope with consequential increasing healthcare costs and medical management needs. Personal health monitoring systems are poised to offer large-scale and cost-effective solutions to this problem. The use of wearable, miniaturized and autonomous wireless sensor nodes, featuring continuous on-node analysis of biosignals, can indeed provide ambulatory long-term and real-time monitoring required by the patients, and enables faster coordination with medical personnel. In such autonomous nodes, due to very limited available energy resources and costly wireless transmission, an ultra-low-power (ULP) on-node processing platform for advanced biosignal analysis is crucial. In this thesis, I explore ULP processing architectures for on-node biosignal analysis applications; where commonly, moderately complex arithmetic manipulations on single- or multiple- input signals are carried out. To achieve energy efficiency while providing sufficient processing capability to apply advanced biosignal analysis, in this thesis near-threshold (near-Vt h ) computing is exploited. Hence, severe performance degradation and reliability issues, occurring at deeply scaled voltages, can be avoided. In Chapter 3, I introduce a near-Vth computing single-core architecture, consisting of a ULP core, an instruction memory (IM) and a data memory (DM). The ULP core features an instruction set architecture (ISA) customized for biosignal applications. I explore that an ISA with minimal instruction set achieves considerable energy savings compared to the state-of-the-art cores, when executing biosignal applications (i.e., up to 54% compared to an established ISA). The proposed single-core architecture accomplishes high energy efficiency for most of single-input biosignal analysis applications, since it fully exploits near-Vth computing. However, the single-core architecture achieves limited voltage scaling, hence reduced energy awareness, for most of multiple-input biosignal analysis applications, where computational workload requirements are such high that the single-core architecture can’t attain these throughputs in near-Vth regime. To alleviate the performance degradation issue that prevents the single-core architecture from exploiting near-Vt h computing typically for multiple-input biosignal analysis, I propose parallel processing of biosignals on multi-core architectures. To this end, In Chapter 4, a multiple instruction, multiple data (MIMD) multi-core architecture is introduced. The MIMD architecture comprises several ULP cores, individual IMs, and a multi-bank DM shared through a lightweight interconnect between the cores and the DM. I prove that parallel processing of multiple-input biosignals leads to better energy efficiency than the sequential processing (i.e., on a single-core) for moderate and high biosignal workloads. In particular, the MIMD architecture achieves up to 62% power savings with respect to the single-core architecture for high biosignal workloads (i.e., 167 MOps/s). On the other hand, parallel processing of multiple-input biosignals can be penalized at low workloads due to high leakage power dissipation in multi-core architectures. In particular, the MIMD architecture fails against the single-core architecture in terms of energy efficiency for workloads lighter than 1.7 MOps/s. One of the major burden of power dissipation in MIMD architectures is costly multiple instruction fetch. To mitigate this issue, I propose data-level parallelism through single instruction, multiple data (SIMD) paradigm. To this end, in Chapter 4 a novel hybrid multi-core architecture, that supports SIMD and MIMD operations, is introduced. The SIMD operations, coupled with data and instruction broadcasting, enable coordinated multiple accesses to memories, hence reduced instruction fetch power. Additionally, the hybrid multi-core architecture features partial power gating of memories to achieve leakage power savings, vital at low workloads (a few 100 kOps/s). I show that SIMD processing of multiple-input biosignals leads to better energy efficiency compared to the MIMD processing. In particular, when SIMD operations are exploited, the hybrid multi-core architecture achieves up to 45.7% power saving compared to the MIMD architecture for moderate biosignal workloads. I also ascertain that partial power gating of memories is an effective technique to alleviate leakage issue in multi-core architectures. More specifically, partial power gating of the IM in the hybrid multi-core architecture leads to 38.8% power saving at low workloads. Finally, to alleviate issues with applications involving such program parts that limit SIMD execution of applications (i.e., conditional program parts), I propose to resynchronize the cores for stable lockstep code execution in case of synchronization loss. Hence, SIMD operations are exploited even for applications with conditional program parts. To this end, in Chapter 4 a lightweight software-directed hardware synchronizer is introduced. I reveal that for applications with conditional program parts, lockstep SIMD execution accomplishes up to 64% power saving with respect to the elementary SIMD execution at moderate workloads (i.e.,89 MOps/s).
Lausanne, EPFL, 2013.Book Chapters
An Ultra-Low-Power Application-Specific Processor with Sub-VT Memories for Compressed Sensing
Compressed sensing (CS) is a universal low-complexity data compression technique for signals that have a sparse representation in some domain. While CS data compression can be done both in the analog- and digital domain, digital implementations are often used on low-power sensor nodes, where an ultra-low-power (ULP) processor carries out the algorithm on Nyquist-rate sampled data. In such systems an energy-efficient implementation of the CS compression kernel is a vital ingredient to maximize battery lifetime. In this paper, we propose an application-specific instruction-set processor (ASIP) processor that has been optimized for CS data compression and for operation in the subthreshold (sub-VT) regime. The design is equipped with specific sub-VT capable standard-cell based memories, to enable low-voltage operation with low leakage. Our results show that the proposed ASIP accomplishes 62× speed-up and 11.6× power savings with respect to a straightforward CS implementation running on the baseline low-power processor without instruction set extensions.
VLSI-SoC: From Algorithms to Circuits and System-on-Chip Design; Springer, 2013. p. 88 – 106.Posters
Application-Specific Processor Design for Low-Complexity & Low-Power Embedded Systems
PhD Forum: Presentation & Poster Session
Winter School on Design Technologies for Heterogeneous Embedded Systems (FETCH), Leysin, Vaud, Switzerland, January 7-9, 2013.A Multipurpose Testbed for Full-Duplex Wireless Communications
Unlike in the traditional half-duplex (HD) mode of transmission over a wireless link, in the full-duplex (FD) mode both terminals receive and transmit on the same frequency band, simultaneously. Hence, FD gives us up to twice the spectral efficiency of HD, has the potential to resolve the hidden terminal problem, and makes relaying much more efficient. As such, FD is one of the most promising new physical layer techniques being developed at the moment. However, a major difficulty in achieving FD operation is that it requires cancelation of self-interference. In this demonstration, we show various methods of operating in FD mode, offering live, interactive, and visual insight into the operation of the link. The parameters shown include the power budget, signal-to-interference-and-noise ratio, frequency spectrum, constellation graphs, and bit error rate. We compare the operation of the same link in HD and FD mode, highlight the effect of each type of cancelation, as well as show what happens when self-interference is not mitigated at all.
International Conference on Electronics, Circuits, and Systems, Abu Dhabi, UAE, December 8-11 2013.Talks
Near- and Sub-Threshold Design for Ultra-Low-Power Embedded Systems
Ultra-low-power (ULP) software-programmable architectures are gradually replacing dedicated VLSI circuits in many applications, including health care and other critical areas. However, the cost for more flexibility is the less frugal use of energy. This cost can be partially recovered by aggressive supply voltage scaling, often deep into the sub-threshold regime, which, however, raises concerns on performance, standby leakage, and reliability. In this talk, we will discuss some of the issues and possible solutions to ULP computing and embedded systems desigm at scaled voltages. We will discuss architectural choices and circuit level aspects and illustrate them with examples including robust Sub-VT memories, ULP multi-core systems, and Sub-VT application specific processors.
Winter School on Design Technologies for Heterogeneous Embedded Systems (FETCH), Leysin, Vaud, Switzerland, January 7-9, 2013.2012
Journal Articles
VLSI Design of Approximate Message Passing for Signal Restoration and Compressive Sensing
IEEE Journal of Emerging and Selected Topics in Circuits and Systems. 2012. Vol. 2, num. 3, p. 579 – 590. DOI : 10.1109/JETCAS.2012.2214636.Analysis and VLSI Implementation of EWA Rendering for Real-Time HD Video Applications
Non-linear image warping or image resampling is a necessary step in many current and upcoming video applications such as video retargeting, stereoscopic 3D mapping, and multiview synthesis. The challenges for real-time resampling include resampling image quality but also available energy and computational power of the employed device. In this work, we employ an elliptical-weighted average (EWA) rendering approach to 2D image resampling. We extend the classical EWA framework for increased visual quality and provide a VLSI architecture for efficient view rendering. The resulting architecture is able to render high-quality video sequences in real-time targeted for low-power applications in end-user display devices.
IEEE Transactions on Circuits and Systems for Video Technology. 2012. Vol. 22, num. 11, p. 1577 – 1589. DOI : 10.1109/Tcsvt.2012.2201671.Low-power processor architecture exploration for online biomedical signal analysis
In this study, the authors explore sequential and parallel processing architectures, utilising a custom ultra-low-power (ULP) processing core, to extend the lifetime of health monitoring systems, where slow biosignal events and highly parallel computations exist. To this end, a single- and a multi-core architecture are proposed and compared. The single-core architecture is composed of one ULP processing core, an instruction memory (IM) and a data memory (DM), while the multi-core architecture consists of several ULP processing cores, individual IMs for each core, a shared DM and an interconnection crossbar between the cores and the DM. These architectures are compared with respect to power/ performance trade-offs for different target workloads of online biomedical signal analysis, while exploiting near threshold computing. The results show that with respect to the single-core architecture, the multi-core solution consumes 62% less power for high computation requirements (167 MOps/ s), while consuming 46% more power for extremely low computation needs when the power consumption is dominated by leakage. Additionally, the authors show that the proposed ULP processing core, using a simplified instruction set architecture (ISA), achieves energy savings of 54% compared to a reference microcontroller ISA (PIC24).
Circuits, Devices & Systems, IET. 2012. Vol. 6, num. 5, p. 279 – 286. DOI : 10.1049/iet-cds.2012.0011.Logic and Memory Design Based on Unequal Error Protection for Voltage-scalable, Robust and Adaptive DSP Systems
In this paper, we propose a system level design approach considering voltage over-scaling (VOS) that achieves error resiliency using unequal error protection of different computation elements, while incurring minor quality degradation. Depending on user specifications and severity of process variations/channel noise, the degree of VOS in each block of the system is adaptively tuned to ensure minimum system power while providing “just-the-right” amount of quality and robustness. This is achieved, by taking into consideration block level interactions and ensuring that under any change of operating conditions, only the “less-crucial” computations, that contribute less to block/system output quality, are affected. The proposed approach applies unequal error protection to various blocks of a system-logic and memory-and spans multiple layers of design hierarchy-algorithm, architecture and circuit. The design methodology when applied to a multimedia subsystem shows large power benefits ( up to 69% improvement in power consumption) at reasonable image quality while tolerating errors introduced due to VOS, process variations, and channel noise.
Journal Of Signal Processing Systems For Signal Image And Video Technology. 2012. Vol. 68, p. 415 – 431. DOI : 10.1007/s11265-011-0631-9.Distributed static linear Gaussian models using consensus
Algorithms for distributed agreement are a powerful means for formulating distributed versions of existing centralized algorithms. We present a toolkit for this task and show how it can be used systematically to design fully distributed algorithms for static linear Gaussian models, including principal component analysis, factor analysis, and probabilistic principal component analysis. These algorithms do not rely on a fusion center, require only low-volume local (1-hop neighborhood) communications, and are thus efficient, scalable, and robust. We show how they are also guaranteed to asymptotically converge to the same solution as the corresponding existing centralized algorithms. Finally, we illustrate the functioning of our algorithms on two examples, and examine the inherent cost-performance trade-off. (C) 2012 Elsevier Ltd. All rights reserved.
Neural Networks. 2012. Vol. 34, p. 96 – 105. DOI : 10.1016/j.neunet.2012.07.004.Conference Papers
Review and Classification of Gain Cell eDRAM Implementations
With the increasing requirement of a high-density, high-performance, low power alternative to traditional SRAM, Gain Cell (GC) embedded DRAMs have gained a renewed interest in recent years. Several industrial and academic publications have presented GC memory implementations for various target applications, including high-performance processor caches, wireless communication memories, and biomedical system storage. In this paper, we review and compare the recent publications, examining the design requirements and the implementation techniques that lead to achievement of the required design metrics of these applications.
2012. DOI : 10.1109/EEEI.2012.6377022.Successive Interference Cancellation for 3G Downlink: Algorithm and VLSI Architecture
2012. IEEE/IFIP International Conference on Very Large Scale Integration (VLSI-SoC), Santa Cruz, CA, USA, October 7-10, 2012. p. 279 – 282. DOI : 10.1109/VLSI-SoC.2012.7332117.Replica Bit-Line Technique for Embedded Multilevel Gain-Cell DRAM
Multilevel gain-cell DRAMs are interesting to improve the area-efficiency of modern fault-tolerant systems-on-chip implemented in deep-submicron CMOS technologies. This paper addresses the problem of long access times in such multilevel gain-cell DRAMs, which are further aggravated by process parameter variations. A replica bit-line (BL) technique, previously proposed for SRAM, is adapted to speed up the multilevel read operation at a negligible area-increase. Moreover, the same replica column is used to improve the write access time. An 8-kb DRAM macro implemented in 90-nm CMOS technology shows that the replica column is able to successfully track die-to-die process, voltage, and temperature variations to generate control signals with optimum delay. Finally, Monte-Carlo simulations show that a small timing margin of 100 ps is sufficient to also cope with within-die process variations.
2012. IEEE International NEWCAS Conference, Montréal, Canada, June 17-20, 2012. p. 77 – 80. DOI : 10.1109/NEWCAS.2012.6328960.Layered Detection and Decoding in MIMO Wireless Systems
Iterative detection and decoding (IDD) in multiple-input multiple-output (MIMO) wireless systems is known to achieve near channel capacity. The high computational complexity of IDD, however, poses significant challenges for practical implementations (in terms of circuit area, latency, throughput, and power consumption). While the implementation of the involved detector and decoder circuits have received considerable attention in the literature, only little is known about the efficient combination of both blocks in an IDD architecture. In this paper, we propose a novel iterative receiver schedule, which simultaneously performs detection and decoding on the same code block. This novel IDD approach is referred to as layered detection and decoding (LDD) and achieves lower latency and better performance compared to conventional solutions. Moreover, LDD is able to automatically match the decoding effort to the wide range of different modulation schemes and code rates specified in modern MIMO wireless standards. To demonstrate the advantages of LDD, we present an extensive case study based on the characteristics of existing reference designs of a soft-input soft-output MMSE detector and an LDPC decoder.
2012. Conference on Design and Architectures for Signal and Image Processing (DASIP) 2012, Karlsruhe, Germany, October 23-25, 2012.Multi-Core Architecture Design for Ultra-Low-Power Wearable Health Monitoring Systems
Personal health monitoring systems can offer a cost-effective solution for human healthcare. To extend the lifetime of health monitoring systems, we propose a near-threshold ultra-low- power multi-core architecture featuring low-power cores, yet capable of executing biomedical applications, with multiple instruction and data memories, tightly coupled through flexible crossbar interconnects. This architecture also includes broadcasting mechanisms for the data and instruction memories to optimize system energy consumption by tailoring memory sharing to the target application. Moreover, the architecture enables power gating of the unused memory banks to lower leakage power. Our experimental results show that compared to the state-of-the-art, the proposed architecture achieves 39.5% power savings at high workload requirements (637 MOps/s), and 38.8% savings at low workload requirements (5 kOps/s), whereby leakage power consumption dominates.
2012. IEEE/ACM 2012 Design Automation and Test in Europe conference (DATE), Dresden, Germany, March 12-16, 2012. p. 988 – 994. DOI : 10.1109/DATE.2012.6176640.A 2.78 mm2 65 nm CMOS Gigabit MIMO Iterative Detection and Decoding Receiver
Iterative detection and decoding (IDD), combined with spatial-multiplexing multiple-input multiple-output (MIMO) transmission, is a key technique to improve spectral efficiency in wireless communications. In this paper we present the—to the best of our knowledge—first complete silicon implementation of a MIMO IDD receiver. MIMO detection is performed by a multi-core sphere decoder supporting up to 4×4 as antenna configuration and 64-QAM modulation. A flexible low-density parity check decoder is used for forward error correction. The 65 nm CMOS ASIC has a core area of 2.78 mm2 . Its maximum throughput exceeds 1 Gbit/s, at less than 1 nJ/bit. The MIMO IDD ASIC enables more than 2 dB performance gains with respect to non-iterative receivers.
2012. 38th European Solid-State Circuits Conference, Bordeaux, France, September 17-21, 2012.Low-complexity Frequency Synchronization for GSM Systems: Algorithms and Implementation
2012. International Conference on Ultra Modern Telecommunications (ICUMT), St. Petersburg, Russia, October 3-5, 2012.TamaRISC-CS: An Ultra-Low-Power Application-Specific Processor for Compressed Sensing
Compressed sensing (CS) is a universal technique for the compression of sparse signals. CS has been widely used in sensing platforms where portable, autonomous devices have to operate for long periods of time with limited energy resources. Therefore, an ultra-low-power (ULP) CS implementation is vital for these kind of energy-limited systems. Sub-threshold (sub-VT ) operation is commonly used for ULP computing, and can also be combined with CS. However, most established CS implementations can achieve either no or very limited benefit from sub-VT operation. Therefore, we propose a sub-VT application-specific instruction-set processor (ASIP), exploiting the specific operations of CS. Our results show that the proposed ASIP accomplishes 62x speed-up and 11.6x power savings with respect to an established CS implementation running on the baseline low-power processor.
2012. IFIP/IEEE 20th International Conference on Very Large Scale Integration (VLSI-SoC), Santa Cruz, USA, October 7-10, 2012. p. 159 – 164. DOI : 10.1109/VLSI-SoC.2012.7332094.Data Mapping for Unreliable Memories
Future digital signal processing (DSP) systems must provide robustness on algorithm and application level to the presence of reliability issues that come along with corresponding implementations in modern semiconductor process technologies. In this paper, we address this issue by investigating the impact of unreliable memories on general DSP systems. In particular, we propose a novel framework to characterize the effects of unreliable memories, which enables us to devise novel methods to mitigate the associated performance loss. We propose to deploy specifically designed data representations, which have the capability of substantially improving the system reliability compared to that realized by conventional data representations used in digital integrated circuits, such as 2’scomplement or sign-magnitude number formats. To demonstrate the efficacy of the proposed framework, we analyze the impact of unreliable memories on coded communication systems, and we show that the deployment of optimized data representations substantially improves the error-rate performance of such systems.
2012. 50th Annual Allerton Conference on Communication, Control, and Computing, October, 1-5, 2012. p. 679 – 685. DOI : 10.1109/Allerton.2012.6483283.A 500 fW/bit 14 fJ/bit-access 4kb Standard-Cell Based Sub-VT Memory in 65nm CMOS
Ultra-low power (ULP) biomedical implants and sensor nodes typically require small memories of a few kb, while previous work on reliable subthreshold (sub-Vt) memories targets several hundreds of kb. Standard-cell based memories (SCMs) are a straightforward approach to realize robust sub-Vt storage arrays and fill the gap of missing sub-Vt memory compilers. This paper presents an ultra-low-leakage 4kb SCM manufactured in 65nm CMOS technology. To minimize leakage power during standby, a single custom-designed standard-cell (D-latch with 3-state output buffer) addressing all major leakage contributors of SCMs is seamlessly integrated into the fully automated SCM compilation flow. Silicon measurements of a 4kb SCM indicate a leakage power of 500fW per stored bit (at a data-retention voltage of 220mV) and a total energy of 14fJ per accessed bit (at energy-minimum voltage of 500mV), corresponding to the lowest values in 65nm CMOS reported to date.
2012. IEEE European Solid-State Circuits Conference (ESSCIRC), Bordeaux, September 17-21, 2012. p. 321 – 324. DOI : 10.1109/ESSCIRC.2012.6341319.Shortening design time through multiplatform simulations with a portable OpenCL golden-model: the LDPC decoder case
Hardware designers and engineers typically need to explore a multi-parametric design space in order to find the best configuration for their designs using simulations that can take weeks to months to complete. For example, designers of special purpose chips need to explore parameters such as the optimal bitwidth and data representation. This is the case for the development of complex algorithms such as Low-Density Parity-Check (LDPC) decoders used in modern communication systems. Currently, high-performance computing offers a wide set of acceleration options, that range from multicore CPUs to graphics processing units (GPUs) and FPGAs. Depending on the simulation requirements, the ideal architecture to use can vary. In this paper we propose a new design flow based on OpenCL, a unified multiplatform programming model, which accelerates LDPC decoding simulations, thereby significantly reducing architectural exploration and design time. OpenCL-based parallel kernels are used without modifications or code tuning on multicore CPUs, GPUs and FPGAs. We use SOpenCL (Silicon to OpenCL), a tool that automatically converts OpenCL kernels to RTL for mapping the simulations into FPGAs. To the best of our knowledge, this is the first time that a single, unmodified OpenCL code is used to target those three different platforms. We show that, depending on the design parameters to be explored in the simulation, on the dimension and phase of the design, the GPU or the FPGA may suit different purposes more conveniently, providing different acceleration factors. For example, although simulations can typically execute more than 3× faster on FPGAs than on GPUs, the overhead of circuit synthesis often outweighs the benefits of FPGA-accelerated execution.
2012. IEEE 20th International Symposium on Field-Programmable Custom Computing Machines (FCCM), 2012. p. 224 – 231. DOI : 10.1109/FCCM.2012.46.Low Complexity Spectral Analysis of Heart-Rate-Variability through a Wavelet Based FFT
In this paper, a low complexity system for spectral analysis of heart rate variability (HRV) is presented. The main idea of the proposed approach is the implementation of the Fast-Lomb periodogram that is a ubiquitous tool in spectral analysis, using a wavelet based Fast Fourier transform. Interestingly we show that the proposed approach enables the classification of processed data into more and less significant based on their contribution to output quality. Based on such a classification a percentage of less-significant data is being pruned leading to a significant reduction of algorithmic complexity with minimal quality degradation. Indeed, our results indicate that the proposed system can achieve up-to 43% reduction in number of computations with only 4.9% average error in the output quality compared to a conventional FFT based HRV system.
2012. IEEE Computing in Cardiology (CinC), September, 2012. p. 285 – 288.Two-Port Low-Power Gain-Cell Storage Array: Voltage Scaling and Retention Time
The impact of supply voltage scaling on the retention time of a 2-transistor (2T) gain-cell (GC) storage array is investigated, in order to enable low-power/low-voltage data storage. The retention time can be increased when scaling down the supply voltage for a given access statistics and a given write bit-line (WBL) control scheme. Moreover, for a given supply voltage, the retention time can be further increased by controlling the WBL to a voltage level between the supply rails during idle and read states. These two concepts are proved by means of Spectre simulation of a GC-storage array implemented in 180-nm CMOS technology. The proposed 2-kb storage macro is operated at only 40% of the nominal supply voltage and leverages the GCs to enable two-port operation with a negligible area-increase compared to a single-port implementation.
2012. IEEE International Symposium on Circuits and Systems (ISCAS), Seoul, Korea, May 20-23, 2012. p. 2469 – 2472. DOI : 10.1109/ISCAS.2012.6271800.A Sub-VT 2T Gain-Cell Memory for Biomedical Applications
All state-of-the-art subthreshold (sub-VT) memories are based on static bitcells, while the feasibility and limitations of dynamic bitcells operated in the sub-VT regime have not been studied yet. For the first time ever, we examine the sub-VT operation of gain-cells and present a fully functional memory array with data retention times that are 10e4X higher than access times.
2012. IEEE Subthreshold Microelectronics Conference, Boston, Massachusetts, USA, October 9-10, 2012. DOI : 10.1109/SubVT.2012.6404318.Design of Energy Efficient and Dependable Health Monitoring Systems under Unreliable Nanometer Technologies
In this paper we investigate the impact of potential hardware misbehavior induced by reliability issues and scaled voltages in wireless body sensor network (WBSN) nodes. Our study reveals the inherent resilience of popular algorithms in cardiac monitoring applications and argues that by exploiting the unique characteristics of such algorithms the energy efficiency and reliability of such systems can be significantly improved. This is achieved by developing a cross-layer design paradigm that utilizes low cost techniques at the hardware and software layers and by optimizing the synergy between them in order to provide intelligent trade-offs between energy, performance and quality. The main idea of the proposed approach is the selective application of costly robust techniques only to the most critical tasks identified at the application layer that are detrimental for obtaining sufficient output quality. Our results show that by ensuring the correct operation of only 37% of the total computations in an electrocardiogram (ECG) monitoring WBSN node we can achieve up to 70% power savings with only 9% degradation in ECG output quality.
2012. 7th International Conference on Body Area Networks (BodyNets ’12), Oslo, Norway, September 24-26 2012. p. 52 – 58. DOI : 10.4108/icst.bodynets.2012.249935.Instruction Set Extensions for Cryptographic Hash Functions on a Microcontroller Architecture
In this paper, we investigate the benefits of instruction set extensions (ISEs) on a 16-bit microcontroller architecture for software implementations of cryptographic hash functions, using the example of the five SHA-3 final round candidates. We identify the general algorithm bottlenecks, taking into account memory footprints and cycle counts of our optimized reference assembly implementations. We show that our target applications benefit from algorithm-specific ISEs based on finite state machines for address generation, lookup table integration, and extension of computational units through microcoded instructions. The gains in throughput, memory consumption, and the area overhead are assessed, by implementing the modified cores and applications utilizing the developed ISEs. Our results show that with less than 10% additional core area, it is possible to increase the execution speed on average by 172% (ranging from 21% to 703%), while reducing memory requirements on average by more than 40%.
2012. 23rd IEEE International Conference on Application-specific Systems, Architectures and Processors (ASAP), Delft, The Netherlands, July 9-11, 2012. p. 117 – 124. DOI : 10.1109/ASAP.2012.13.On the Exploitation of the Inherent Error Resilience of Wireless Systems under Unreliable Silicon
In this paper, we investigate the impact of circuit misbehavior due to parametric variations and voltage scaling on the performance of wireless communication systems. Our study reveals the inherent error resilience of such systems and argues that sufficiently reliable operation can be maintained even in the presence of unreliable circuits and manufacturing defects. We further show how selective application of more robust circuit design techniques is sufficient to deal with high defect rates at low overhead and improve energy efficiency with negligible system performance degradation.
2012. IEEE Design Automation Conference (DAC), San Francisco, June 3-7, 2012. p. 510 – 515. DOI : 10.1145/2228360.2228451.Book Chapters
VLSI Implementation of Hard- and Soft-Output Sphere Decoding for Wide-Band MIMO Systems
VLSI-SoC: Forward-Looking Trends in IC and Systems Design; Berlin, Heidelberg: Springer Berlin Heidelberg, 2012. p. 128 – 154.Working Papers
Investigating the Potential of Custom Instruction Set Extensions for SHA-3 Candidates on a 16-bit Microcontroller Architecture
In this paper, we investigate the benefit of instruction set extensions for software implementations of all five SHA-3 candidates. To this end, we start from optimized assembly code for a common 16-bit microcontroller instruction set architecture. By themselves, these implementations provide reference for complexity of the algorithms on 16-bit architectures, commonly used in embedded systems. For each algorithm, we then propose suitable instruction set extensions and implement the modified processor core. We assess the gains in throughput, memory consumption, and the area overhead. Our results show that with less than 10% additional area, it is possible to increase the execution speed on average by almost 40%, while reducing memory requirements on average by more than 40%. In particular, the Grostl algorithm, which was one of the slowest algorithms in previous reference implementations, ends up being the fastest implementation by some margin, once minor (but dedicated) instruction set extensions are taken into account.
2012
Posters
Standard-Cell Based Memories (SCMs): from Sub-VT to Error-Resilient Systems
Embedded memories consume an increasingly dominant part of the overall area and power of a large variety of systems-on-chip [ITRS’09]: 1) biomedical implants and wireless sensor networks require robust memories operating in the sub-VT domain; 2) many handheld devices and microprocessors are operated near to threshold-voltage; and 3) fault-tolerant systems/error-resilient computing has attracted interest due to increaing process variations. Standard-cell based memories (SCMs) entail minimum design effort and are immediately functional in any system from reliable sub-VT to error-resilient high-performance. In particular, sub-VT SCMs ensure robustness and improve access bandwidth and energy-efficiency compared to sub-VT SRAM macros. Adding only one custom cell (low-leakage latch) to a commercial standard-cell library further improves energy-efficiency of sub-VT SCMs. In fault-tolerant systems requiring small data retention times, a small amount of errors in the memory content does not severely impede system functionality, and dynamic latches yield SCMs smaller than commercial 6T SRAM macros for storage capacities up to at least 2kb. Various silicon-prooven SCM architectures are presented, and the best-practice SCM implementations for both sub-VT and above-VT applications are derived. To reduce leakage power in sub-VT SCMs, a latch with few highly resistive VDD-ground path is designed using transistor stacking and stretching. For the benefit of smaller silicon area, but at the cost of reduced robustness, various dynamic latches are integrated in the SCM compilation flow.
IEEE International Solid-State Circuits Conference (ISSCC), Student Research Preview (SRP) session, San Francisco, California, USA, February 17-21, 2012.Talks
1 Mbps coherent one-way QKD with dense wavelength division multiplexing and hardware key distillation
We present the latest results obtained with a quantum cryptography prototype based on a coherent-one way quantum key distribution (QKD) scheme. To support its continuous high rate secret key generation we developed different low-noise single photon detectors for telecom wavelength based on a sine gating and low-pass-filtering technique, as well as a negative feedback APD in an active hold-off circuit. A newly developed hardware distillation engine allows for continuous operation of secret key distribution up to 1 Mbps. We also present results of our system in a DWDM (dense wavelength-division multiplexing) configuration where only one single fiber is needed to interconnect Alice’ and Bob’s systems. The final prototype is fully compatible to serve a high-speed encryption device developed in parallel which provides encrypted communication of up to 100 Gbps.
2nd Annual Conference on Quantum Cryptography (QCRYPT 2012), Singapore, September 10th-14th, 2012.Student Projects
Building a visible light communication system
2012.2011
Journal Articles
An FPGA-based processing pipeline for high definition stereo video
This paper presents a real-time processing platform for high definition stereo video. The system is capable to process stereo-video streams at resolutions up to 1920×1080 at 30 frames per second (1080p30). In the hybrid FPGA-GPU-CPU system, a high-density FPGA is used to perform not only the low-level image processing tasks such as color interpolation and cross-image color correction, but also to carry out radial undistortion, image rectification, and disparity estimation. We show how the corresponding algorithms can be implemented very efficiently in programmable hardware, relieving the GPU from the burden of these tasks. Our FPGA implementation results are compared to corresponding GPU implementations and to other implementations reported in the literature.
EURASIP Journal on Image and Video Processing. 2011. Vol. 2011, p. 18. DOI : 10.1186/1687-5281-2011-18.Benchmarking of standard-cell based memories in the sub-VT domain in 65-nm CMOS technology
In this paper, standard-cell based memories (SCMs) are proposed as an alternative to full-custom sub-VT SRAM macros for ultra-low-power systems requiring small memory blocks. The energy per memory access as well as the maximum achievable throughput in the sub-VT domain of various SCM architectures are evaluated by means of a gate-level sub-VT characterization model, building on data extracted from fully placed, routed, and back-annotated netlists. The reliable operation at the energy-minimum voltage of the various SCM architectures in a 65-nm CMOS technology considering within-die process parameter variations is demonstrated by means of Monte Carlo circuit simulation. Finally, the energy per memory access, the achievable throughput, and the area of the best SCM architecture are compared to recent sub-VT SRAM designs.
IEEE Journal of Emerging and Selected Topics in Circuits and Systems. 2011. Vol. 1, num. 2, p. 173 – 182. DOI : 10.1109/JETCAS.2011.2162159.Conference Papers
Significance Driven Computation on Next-Generation Unreliable Platforms
In this paper, we propose a design paradigm for energy efficient and variation-aware operation of next-generation multicore heterogeneous platforms. The main idea behind the proposed approach lies on the observation that not all operations are equally important in shaping the output quality of various applications and of the overall system. Based on such an observation, we suggest that all levels of the software design stack, including the programming model, compiler, operating system (OS) and run-time system should identify the critical tasks and ensure correct operation of such tasks by assigning them to dynamically adjusted reliable cores/units. Specifically, based on error rates and operating conditions identified by a sense-and-adapt (SeA) unit, the OS selects and sets the right mode of operation of the overall system. The run-time system identifies the critical/less-critical tasks based on special directives and schedules them to the appropriate units that are dynamically adjusted for highly-accurate/approximate operation by tuning their voltage/frequency. Units that execute less significant operations can operate at voltages less than what is required for correct operation and consume less power, if required, since such tasks do not need to be always exact as opposed to the critical ones. Such scheme can lead to energy efficient and reliable operation, while reducing the design cost and overheads of conventional circuit/micro-architecture level techniques.
2011. 48th ACM/IEEE/EDAC Design Automation Conference (DAC), San Diego, CA, Jun 05-09, 2011. p. 290 – 291. DOI : 10.1145/2024724.2024794.Synthesis strategies for sub-VT systems
Various synthesis strategies relying on conventional standard-cell libraries (SCLs) are evaluated in order to minimize the energy dissipation per operation in sub-threshold (sub-VT) systems. First, two sub-V T analysis methods are reviewed, both of which allow to evaluate the energy dissipation and performance in the sub-VT regime for designs which have been synthesized using a 65-nm CMOS SCL, characterized at nominal supply voltage. Both analysis methods are able to predict the energy minimum supply voltage (EMV) of any given design. Next, the results of a sub-V T synthesis at EMV using re-characterized SCLs are compared to the initial synthesis results. Finally, the results of timing-driven synthesis in both the above-VT and the sub-VT domain are compared to the results of power-driven synthesis.
2011. IEEE European Conference on Circuit Theory and Design (ECCTD), Linköping, Sweden, August 29-31, 2011. p. 552 – 555. DOI : 10.1109/ECCTD.2011.6043593.Power/Performance Exploration of Single-core and Multi-core Processor Approaches for Biomedical Signal Processing
This study presents a single-core and a multi-core processor architecture for health monitoring systems where slow biosignal events and highly parallel computations exist. The single-core architecture is composed of a processing core (PC), an instruction memory (IM) and a data memory (DM), while the multi-core architecture consists of PCs, individual IMs for each core, a shared DM and an interconnection crossbar between the cores and the DM. These architectures are compared with respect to power vs performance trade-offs for a multi-lead electrocardiogram signal conditioning application exploiting near threshold computing. The results show that the multi-core solution consumes 66% less power for high computation requirements (50.1 MOps/s), whereas 10.4% more power for low computation needs (681 kOps/s).
2011. Workshop on Power and Timing Modeling, Optimization and Simulation (PATMOS ‘11), Madrid, Spain, September 26-29, 2011. p. 102 – 111. DOI : 10.1007/978-3-642-24154-3_11.Area, throughput, and energy-efficiency trade-offs in the VLSI implementation of LDPC decoders
2011. IEEE International Symposium on Circuits and Systems (ISCAS), Rio de Janeiro, Brazil, May 15-18, 2011. p. 1772 – 1775. DOI : 10.1109/ISCAS.2011.5937927.A 772 Mbit/s 8.81 bit/nJ 90 nm CMOS Soft-Input Soft-Output Sphere Decoder
2011. IEEE Asian Solid-State Circuits Conference, Jeju, Korea, November 14-16, 2011. p. 297 – 300. DOI : 10.1109/ASSCC.2011.6123571.System-level implications of residual transmit-RF impairments in MIMO systems
2011. European Conference on Antennas and Propagation, Rome, Italy, April 11-15, 2011. p. 2686 – 2689.Computational stereo camera system with programmable control loop
2011. ACM SIGGRAPH, Vancouver, Canada, August 7-11, 2011. DOI : 10.1145/2010324.1964989.Random Sampling ADC for Sparse Spectrum Sensing
Scanning large bandwidths (spectrum sensing) pushes today’s analog hardware to its limits since periodic sampling at Nyquist rate with sufficient resolution is often prohibitively complex. In this paper, we consider a scenario where the signal to be acquired is sparse in the frequency domain (e.g., spectrum sensing in cognitive radio applications) and we are interested in identifying the sparse support of the signal. For this type of applications, we describe a new analog-to-digital converter (ADC) architecture that acquires unequally spaced samples based on a slope ADC, which is one of the least complex ADC architectures available. For the signal reconstruction, we employ algorithms from compressed sensing for the recovery of the dominant spectral components. The performance of the proposed design is compared to more traditional designs with comparable or higher hardware complexity.
2011. European Signal Processing Conference, Barcelona, Spain, August 29 – September 2, 2011.Design and failure analysis of logic-compatible multilevel gain-cell-based DRAM for fault-tolerant VLSI systems
This paper considers the problem of increasing the storage density in fault-tolerant VLSI systems which require only limited data retention times. To this end, the concept of storing many bits per memory cell is applied to area-efficient and fully logic-compatible gain-cell-based dynamic memories. A memory macro in 90-nm CMOS technology including multilevel write and read circuits is proposed and analyzed with respect to its read failure probability due to within-die process variations by means of Monte Carlo simulations.
2011. IEEE 21st Edition of the Great Lakes Symposium on VLSI (GLSVLSI), Lausanne, Switzerland, May 2-4, 2011. p. 343 – 346. DOI : 10.1145/1973009.1973078.2010
Journal Articles
Simulation and Emulation of MIMO Wireless Baseband Transceivers
The development of state-of-the-art wireless communication transceivers in semiconductor technology is a challenging process due to complexity and stringent requirements of modern communication standards such as IEEE 802.11n. This tutorial paper describes a complete design, verification, and performance characterization methodology that is tailored to the needs of the development of state-of-the-art wireless baseband transceivers for both research and industrial products. Compared to the methods widely used for the development of communication research testbeds, the described design flow focuses on the evolution of a given system specification to a final ASIC implementation through multiple design representations. The corresponding verification and characterization environment supports rapid floating-point and fixed-point performance characterization and ensures consistency across the entire design process and across all design representations. This framework has been successfully employed for the development and verification of an industrial-grade, fully standard compliant, 4-stream IEEE 802.11n MIMOOFDM baseband transceiver. Copyright (C) 2010 Pierre Greisen et al.
Eurasip Journal On Wireless Communications And Networking. 2010. p. 196796. DOI : 10.1155/2010/196796.Conference Papers
MIMO transmission with residual transmit-RF impairments
2010. 2010 International ITG Workshop on Smart Antennas (WSA), Bremen, Germany, 23-24 02 2010. p. 189 – 196. DOI : 10.1109/WSA.2010.5456453.Area- and throughput-optimized VLSI architecture of sphere decoding
2010. 2010 18th IEEE/IFIP International Conference on VLSI and System-on-Chip (VLSI-SoC), Madrid, Spain, 27-29 09 2010. p. 189 – 194. DOI : 10.1109/VLSISOC.2010.5642593.The effect of unreliable LLR storage on the performance of MIMO-BICM
2010. 2010 44th Asilomar Conference on Signals, Systems and Computers, Pacific Grove, CA, USA, 7-10 November 2010. p. 736 – 740. DOI : 10.1109/ACSSC.2010.5757661.Low-complexity Seysen’s algorithm based lattice reduction-aided MIMO detection for hardware implementations
2010. 2010 44th Asilomar Conference on Signals, Systems and Computers, Pacific Grove, CA, USA, 7-10 November 2010. p. 1468 – 1472. DOI : 10.1109/ACSSC.2010.5757780.Matching pursuit: Evaluation and implementatio for LTE channel estimation
2010. 2010 IEEE International Symposium on Circuits and Systems – ISCAS 2010, Paris, France, 30 05 – 2 06 2010. p. 589 – 592. DOI : 10.1109/ISCAS.2010.5537528.A 15.8 pJ/bit/iter quasi-cyclic LDPC decoder for IEEE 802.11n in 90 nm CMOS
We present a low-power quasi-cyclic (QC) low density parity check (LDPC) decoder that meets the throughput requirements of the highest-rate (600 Mbps) modes of the IEEE 802.11n WLAN standard. The design is based on the layered offset-min-sum algorithm and is runtime-programmable to process different code matrices (including all rates and block lengths specified by IEEE 802.11n). The register-transfer-level implementation has been optimized for best energy efficiency. The corresponding 90nm CMOS ASIC has a core area of 1.77mm2 and achieves a maximum throughput of 680 Mbps at 346MHz clock frequency and 10 decoding iterations. The measured energy efficiency is 15.8 pJ/bit/iteration at a nominal operating voltage of 1.0V.
2010. 2010 IEEE Asian Solid-State Circuits Conference (A-SSCC), Beijing, China, November 8-10, 2010. p. 1 – 4. DOI : 10.1109/ASSCC.2010.5716618.VLSI Implementation of a Low-Complexity LLL Lattice Reduction Algorithm for MIMO Detection
Lattice-reduction (LR)-aided successive interference cancellation (SIC) is able to achieve close-to optimum error-rate performance for data detection in multiple-input multiple-output (MIMO) wireless communication systems. In this work, we propose a hardware-efficient VLSI architecture of the Lenstra-Lenstra-Lovasz (LLL) LR algorithm for SIC-based data detection. For this purpose, we introduce various algorithmic modifications that enable an efficient hardware implementation. Comparisons with existing FPGA implementations show that our design outperforms state-of-the-art LR implementations in terms of hardware-efficiency and throughput. We finally provide reference ASIC implementation results for 130 nm CMOS technology.
2010. International Symposium on Circuits and Systems Nano-Bio Circuit Fabrics and Systems (ISCAS 2010), Paris, FRANCE, May 30-Jun 02, 2010. p. 3745 – 3748. DOI : 10.1109/ISCAS.2010.5537742.Towards generic low-power area-efficient standard cell based memory architectures
Digital IC designers often use SRAM macrocells to implement on-chip memory functionality. In this paper we argue that in several situations, standard cell based memories (SCMs) can have advantages over SRAM macrocells. Various ways to implement SCMs are presented and compared to each other for different CMOS technologies and standard cell libraries and to corresponding macrocells, aiming for finding the most adequate memory option for each application. The benefits and drawbacks of SCMs compared to macrocells are illustrated with the example of a low-power low-density parity check (LDPC) decoder.
2010. 2010 53rd IEEE International Midwest Symposium on Circuits and Systems (MWSCAS), Seattle, WA, USA, August 1-4, 2010. p. 129 – 132. DOI : 10.1109/MWSCAS.2010.5548579.Systolic-array based regularized QR-decomposition for IEEE 802.11n compliant soft-MMSE detection
2010. 2010 International Conference on Microelectronics (ICM), Cairo, Egypt, 19-22 December 2010. p. 391 – 394. DOI : 10.1109/ICM.2010.5696169.2009
Journal Articles
Design and Optimization of an HSDPA Turbo Decoder ASIC
IEEE Journal of Solid-State Circuits. 2009. Vol. 44, num. 1, p. 98 – 106. DOI : 10.1109/JSSC.2008.2007166.Conference Papers
Implementation of a 2×2 MIMO-OFDM receiver on an application specific processor
This paper describes the implementation of the hard computational kernels required for the baseband (BB) processing of a 2×2 multiple-input multiple-output (MIMO)-OFDM receiver on a design-framework for application specific processors. The employed low-complexity BB algorithms are described and their computational complexity is derived. The receiver is split into two parts which are mapped onto two application specific processors, each tailored to the computational needs of the associated digital signal processing kernels. The first processor performs the per stream MIMO-OFDM processing. The second processor handles the MIMO detection. Finally, the 0.18um 1P/6M CMOS technology layout of both fabricated application specific processors is presented. Real-time BB processing is possible on these engines running at a clock frequency of 250 MHz.
2009. International Conference on Microelectronics, Cairo, EGYPT, Dec 29-31, 2007. p. 1642 – 1649. DOI : 10.1016/j.mejo.2009.02.005.A 4-Stream 802.11n Baseband Transceiver in 0.13 mu m CMOS
An IEEE 802.11n baseband transceiver ASIC is implemented in 0.13 mu m CMOS technology. The implementation has a core area of 14.4 mm(2) and is the first to support the optional 3- and 4-stream MIMO transmission modes of the standard for data rates up to 600 Mbps.
2009. Symposium on VLSI Circuits, Kyoto, JAPAN, Jun 16-18, 2009. p. 282 – 283.2008
Journal Articles
A real-time 4-stream MIMO-OFDM transceiver: System design, FPGA implementation, and characterization
When designing complex communication systems, such as MIMO-OFDM transceivers, prototypes have become an important tool for understanding the implementation trade-offs and the system behavior. This paper presents a real-time FPGA prototype for a 4-stream MIMO-OFDM transceiver capable of transmitting 216 Mbit/s in 20MHz bandwidth. The paper covers all parts of the system from RF to channel decoding and considers both algorithm and implementation aspects. In particular, we discuss the initial parameter estimation, channel estimation, MIMO detection, parameter tracking, and channel, decoding. FPGA implementation results are reported along with measurements that demonstrate the throughput of spatial multiplexing with four spatial streams.
IEEE Journal on Selected Areas in Communications. 2008. Vol. 26, p. 877 – 889. DOI : 10.1109/JSAC.2008.080805.Conference Papers
Hardware-efficient steering matrix computation architecture for MIMO communication systems
Beamforming (BF) improves the error rate performance of multiple-input multiple-output (MIMO) wireless communication systems by spatial separation of the transmitted data streams. Spatial separation is achieved by multiplication of the transmit vector by a steering matrix, which is obtained through the singular value decomposition (SVD) of the channel matrix. In this paper, we describe a hardware-efficient VLSI architecture for steering matrix computation using a hardware- optimized SVD algorithm. Our architecture contains a high-speed Givens rotation unit which achieves high processing throughput at low area. The resulting VLSI implementation requires 3.3 mus per steering matrix computation at an expense of 41.3 kGEs and shows a 3.5-fold hardware-efficiency gain compared to a reference SVD implementation.
2008. IEEE International Symposium on Circuits and Systems, 2008. ISCAS 2008., Seattle, WA, USA, May 18-21 2008. p. 304 – 307. DOI : 10.1109/ISCAS.2008.4541415.Configurable High-Throughput Decoder Architecture for Quasi-Cyclic LDPC Codes
We describe a fully reconfigurable low-density parity check (LDPC) decoder for quasi-cyclic (QC) codes. The proposed hardware architecture is able to decode virtually any QC-LDPC code that fits into the allocated memories while achieving high decoding throughput. Our VLSI implementation has been optimized for the IEEE 802.11n standard and achieves a throughput of 780 M bit/s with a core area of 3.39 mm(2) in 0.18 mu m CMOS technology.
2008. 42nd Asilomar Conference on Signals, Systems and Computers, Pacific Grove, CA, Oct 26-29, 2008. p. 1137 – 1142. DOI : 10.1109/ACSSC.2008.5074592.A 58mW 1.2 mm 2 HSDPA Turbo Decoder ASIC in 0.13 μm CMOS
This paper presents the implementation of the 1.2 mm2 HSDPA turbo decoder ASIC in 0.13 mum CMOS achieves a measured maximum frequency of 246 MHz, which translates to a maximum throughput of 20.2 Mb/s at 5.5 iterations. The peak throughput of 10.8 Mb/s required for HSDPA is achieved at 58 mW and an energy efficiency of 0.7 nJ/b/iter. The number of iterations versus input SNR, as determined by the implemented stopping criterion, and corresponding power measurements.
2008. IEEE International Solid-State Circuits Conference (ISSCC), San Francisco, CA, Feb 03-07, 2008. p. 264 – 265+612. DOI : 10.1109/ISSCC.2008.4523158.Soft-output sphere decoding: Algorithms and VLSI implementation
Multiple-input multiple-output (MIMO) detection algorithms providing soft information for a subsequent channel decoder pose significant implementation challenges due to their high computational complexity. In this paper, we show how sphere decoding can be used as an efficient tool to implement soft-output MIMO detection with flexible trade-offs between computational complexity and (error rate) performance. In particular, we provide VLSI implementation results which demonstrate that single tree-search, sorted QR-decomposition, channel matrix regularization, log-likelihood ratio clipping, and imposing runtime constraints are the key ingredients for realizing soft-output MIMO detectors with near max-log performance at a chip area that is only 58% higher than that of the best-known hard-output sphere decoder VLSI implementation.
2008. 40th Asilomar Conference on Signals, Systems and Computers, Pacific Grove, CA, Oct 29-Nov 02, 2006. p. 290 – 300. DOI : 10.1109/JSAC.2008.080206.2007
Conference Papers
FFT processor for OFDM channel estimation
Pilot-assisted channel estimation for communication systems employing orthogonal frequency division multiplexing modulation requires significant signal processing at the receiver if the correlation among the frequency-domain channel coefficients is to be exploited in order to improve accuracy. In this work, a conventional FFT processor is extended to support all operations required by a selected channel estimation algorithm, so that both OFDM de/modulation and channel estimation can be efficiently performed on the same hardware unit. The silicon complexity of the extended processor, which was prototyped in a real-time testbed using FPGAs, is compared to a conventional FFT processor.
2007. IEEE International Symposium on Circuits and Systems, New Orleans, LA, May 27-30, 2007. p. 1417 – 1420. DOI : 10.1109/ISCAS.2007.378494.VLSI implementation of a high-speed iterative sorted MMSE QR decomposition
The QR decomposition is an important, but often underestimated prerequisite for pseudo- or non-linear detection methods such as successive interference cancellation or sphere decoding for multiple-input multiple-output (MIMO) systems. The ability of concurrent iterative sorting during the QR decomposition introduces a moderate overall latency, but provides the base for an improved layered stream decoding. This paper describes the architecture and results of the first VLSI implementation of an iterative sorted QR decomposition preprocessor for MIMO receivers. The presented architecture performs MIMO channel preprocessing using Givens rotations in order to compute the minimum mean squared error QR decomposition.
2007. IEEE International Symposium on Circuits and Systems, New Orleans, LA, May 27-30, 2007. p. 1421 – 1424. DOI : 10.1109/ISCAS.2007.378495.An IEEE 802.11a baseband receiver implementation on an application specific processor
This paper describes the implementation of the complete baseband processing of an IEEE 802.11a receiver on a design-framework for application specific processors. The underlying generic architecture is described, the computational kernels required for an IEEE 802.11a receiver are analyzed, and suitable processing units and architecture-configurations, to be defined at design-time, are identified. The discussion of the receiver implementation shows that the proposed architecture can meet real-time requirements on a 0.13 mu m CMOS process using a dock frequency of 160 MHz. The design demonstrates how the proposed standard-specific reconfigurable architecture is a valid alternative to ASIC and DSP implementations when looking for a balance between performance and flexibility.
2007. 50th Midwest Symposium on Circuits and Systems, Montreal, CANADA, Sep 05, 2007-Aug 08, 2008. p. 1066 – 1069.Regularized frequency domain equalization algorithm and its VLSI implementation
Approximation of Toeplitz matrices with circulant matrices is a well-known approach to reduce the computational complexity of linear equalizers. This paper presents a novel technique to compute linear equalizer coefficients in the frequency domain. It is shown how a regularization term can help to reduce the error caused by the frequency domain approximation. A corresponding VLSI implementation provides reference for the true silicon complexity and for the complexity increase associated with the proposed algorithm.
2007. IEEE International Symposium on Circuits and Systems, New Orleans, LA, May 27-30, 2007. p. 3530 – 3533. DOI : 10.1109/ISCAS.2007.378444.VLSI implementation of a lattice-reduction algorithm for multi-antenna broadcast precoding
This paper describes the first VLSI implementation of lattice reduction (LR) aided multi-antenna broadcast precoding with vector perturbation. The considered LR scheme is based on Brun’s algorithm for finding integer relations. We analyze its high-level architectural issues, we devise a corresponding low-complexity implementation, and, finally, we develop a suitable VLSI architecture. The resulting circuit provides reference for the true silicon complexity of LR for broadcast precoding with vector perturbation.
2007. IEEE International Symposium on Circuits and Systems, New Orleans, LA, May 27-30, 2007. p. 673 – 676. DOI : 10.1109/ISCAS.2007.377898.2006
Conference Papers
VLSI implementation of the list sphere algorithm
Sphere decoding (SD) is widely considered as one of the most promising detection schemes for multiple-input multiple-output (MIMO) communication systems. The recently proposed list sphere-decoding (LSD) algorithm is an extension of the original SD algorithm that improves the error rate performance of wireless communication systems considerably by providing soft-outputs instead of binary decisions. This paper addresses the VLSI implementation of the LSD algorithm. To this end, algorithm optimizations suitable for efficient hardware implementations are developed. The implemented circuits achieve a gain of up to 3 dB in SNR compared to hard output SDs and a throughput of up to 272 Mbps at 20 dB SNR in a 0.25 mu m technology for 4×4 MIMO systems with 16-QAM modulation.
2006. 24th Norchip Conference, Linkoping, SWEDEN, Nov 20-21, 2006. p. 107 – 110. DOI : 10.1109/NORCHP.2006.329255.A frame-start detector for a 4×4 MIMO-OFDM system
Future wireless LANs will increase the peak data rate by employing multiple antennas at both transmitter and receiver.
2006. 31st IEEE International Conference on Acoustics, Speech and Signal Processing, Toulouse, FRANCE, May 14-19, 2006. p. 4095 – 4098. DOI : 10.1109/ICASSP.2006.1660996.Advanced receiver algorithms for MIMO wireless communications
We describe the VLSI implementation of MIMO detectors that exhibit close-to optimum error-rate performance, but still achieve high throughput at low silicon area. In particular algorithms and VLSI architectures for sphere decoding (SD) and K-best detection are considered, and the corresponding trade-offs between uncoded error-rate performance, silicon area, and throughput are explored. We show that SD with a per-block run-time constraint is best suited for practical implementations.
2006. Design, Automation and Test in Europe Conference and Exhibition (DATE 06), Munich, GERMANY, Mar 06-10, 2006. p. 591 – 596. DOI : 10.1109/DATE.2006.243974.K-Best MIMO detection VLSI architectures achieving up to 424 mbps
From an error rate performance perspective, maximum likelihood (ML) detection is the preferred detection method for multiple-input multiple-output (MIMO) communication systems. However, for high transmission rates a straight forward exhaustive search implementation suffers from prohibitive complexity. The K-best algorithm provides close-to-ML bit error rate (BIER) performance, while its circuit complexity is reduced compared to an exhaustive search. In this paper, a new VLSI architecture for the implementation of the K-best algorithm is presented. Instead of the mostly sequential processing that has been applied in previous VLSI implementations of the algorithm, the presented solution takes a more parallel approach. Further-more, the application of a simplified norm is discussed. The implementation in an ASIC achieves up to 424 Mbps throughput with an area that is almost on par with current state-of-the-art implementations.
2006. IEEE International Symposium on Circuits and Systems, Kos Isl, GREECE, May 21-24, 2006. p. 1151 – 1154. DOI : 10.1109/ISCAS.2006.1692794.Silicon implementation of an MMSE-based soft demapper for MIMO-BICM
The performance of systems employing bit-interleaved coded modulation (BICM) critically depends on the availability of soft information. In the multi-antenna case, the extraction of optimum bit-metrics becomes prohibitively complex, so that suboptimal solutions need to be adopted for practical implementation. Instead of considering all the spatially multiplexed streams jointly, the implementation presented in this paper computes the soft information on each data stream separately, based on the output of an MMSE equalizer.
2006. IEEE International Symposium on Circuits and Systems, Kos Isl, GREECE, May 21-24, 2006. p. 2597 – 2600. DOI : 10.1109/ISCAS.2006.1693155.A unification of ML-optimal tree-search decoders
2006. 40th Asilomar Conference on Signals, Systems and Computers, Pacific Grove, CA, Oct 29-Nov 01, 2006. p. 2185 – 2189. DOI : 10.1109/ACSSC.2006.355156.Algorithm and VLSI architecture for linear MMSE detection in MIMO-OFDM systems
The paper describes an algorithm and a corresponding VLSI architecture for the implementation of linear MMSE detection in packet-based MIMO-OFDM communication systems. The advantages of the presented receiver architecture are low latency, high-throughput, and efficient resource utilization, since the hardware required for the computation of the MMSE estimators is reused for the detection. The algorithm also supports the extraction of soft information for channel decoding.
2006. IEEE International Symposium on Circuits and Systems, Kos Isl, GREECE, May 21-24, 2006. p. 4102 – 4105. DOI : 10.1109/ISCAS.2006.1693531.Soft-output sphere decoding: Performanceand implementation aspects
Multiple-input multiple-output (MIMO) detection algorithms providing soft information for a subsequent channel decoder pose significant implementation challenges due to their high computational complexity. In this paper, we show how sphere decoding can be used as an efficient tool to implement soft-output MIMO detection with flexible trade-offs between computational complexity and (error rate) performance. In particular, we demonstrate that single tree search, ordered QR decomposition, channel matrix regularization, and log-likelihood ratio clipping are the key ingredients for realizing soft-output MIMO detectors with near max-log performance at a computational complexity that is reasonably close to that of hard-output sphere decoding.
2006. 40th Asilomar Conference on Signals, Systems and Computers, Pacific Grove, CA, Oct 29-Nov 01, 2006. p. 2071 – 2076. DOI : 10.1109/ACSSC.2006.355132.2005
Conference Papers
VLSI implementation of MIMO detection using the sphere decoding algorithm
Multiple-input multiple-output (MIMO) techniques are a key enabling technology for high-rate wireless communications. This paper discusses two ASIC implementations of MIMO sphere decoders. The first ASIC attains maximum-likelihood performance with an average throughput of 73 Mb/s at a signal-to-noise ratio (SNR) of 20 dB; the second ASIC shows only a negligible bit-error-rate degradation and achieves a throughput of 170 Mb/s at the same SNR. The three key contributing factors to high throughput and low complexity are: depth-first tree traversal with radius reduction, implemented in a one-node-per-cycle architecture, the use of the l(infinity)–instead of l(2)-norm, and, finally, the efficient implementation of the enumeration approach recently proposed in [1]. The resulting ASICs currently rank among the fastest reported MIMO detector implementations.
2005. 30th European Solid-State Circuits Conference (ESSCIRC 2004), Leuven, BELGIUM, Sep 21-23, 2004. p. 1566 – 1577. DOI : 10.1109/JSSC.2005.847505.ASIC implementation of a MIMO-OFDM transceiver for 192 Mbps WLANs
Next generation wireless local area networks (WLANs) such as the IEEE 802.11n standard are expected to rely on multiple antennas at both transmitter and receiver to increase throughput and link reliability. However, these improvements come at a significant increase in signal processing and hence hardware complexity compared to existing single-antenna systems. This paper presents, to the best of the authors’ knowledge, the first 4 x 4 MIMO-OFDM WLAN physical layer ASIC based on the OFDM specifications of the IEEE 802.11a standard. The ASIC achieves an uncoded throughput of 192 Mbps in a 20 MHz channel resulting in a spectral efficiency of 9.6 bits/s/Hz. We describe the hardware architectural differences to single-antenna OFDM systems as well as the extensions made necessary by the use of multiple antennas. Our implementation provides reference for the silicon complexity of MIMO-OFDM systems.
2005. 31st European Solid-State Circuits Conference, Grenoble, FRANCE, Sep 12-16, 2005. p. 215 – 218. DOI : 10.1109/ESSCIR.2005.1541598.Receiver design for multi-antenna wireless communications
Algorithm choices and corresponding VLSI architectures for spatial multiplexing receivers in multiple-input multiple-output (MIMO) communication systems are described in this paper. Implementations of linear and successive interference cancellation receivers are compared to implementations of maximum likelihood (ML) decoders that attain optimum bit error rate performance. The presented designs provide reference for the true silicon complexity of the algorithms under consideration and help to identify the limits with respect to practical implementations.
2005. International Conference on PhD Research in Microelectronics and Electronics (PRIME 2005), Lausanne, SWITZERLAND, 2005. p. 231 – 234. DOI : 10.1109/RME.2005.1542930.FPGA implementation of Viterbi decoders for MIMO-BICM
The FPGA implementation of Viterbi decoders for multiple-input multiple-output (MIMO) wireless communication systems with bit-interleaved coded modulation (BICM) and perantenna coding is considered. The paper describes how the recursive add-compare-select (ACS) unit, which constitutes the performance bottleneck of the circuit, can be pipelined to increase the throughput. As opposed to employing multiple parallel decoders, silicon area (resource utilization on the FPGA) is significantly reduced. The proposed optimizations lead to an implementation that achieves a throughput of 216 Mbps in a 4 x 4 MIMO-WLAN system prototype based on IEEE 802.11a.
2005. 39th Asilomar Conference on Signals, Systems and Computers, Pacific Grove, CA, Oct 30-Nov 02, 2005. p. 734 – 738. DOI : 10.1109/ACSSC.2005.1599849.Interpolation-based QR decomposition in MIMO-OFDM systems
The extension of multiple-input multiple-output (MIMO) sphere decoding from the narrowband case to wideband systems based on orthogonal frequency division multiplexing (OFDM) requires the computation of a QR decomposition for each of the data-carrying OFDM tones. Since the number of data-carrying tones ranges from 48 (as in the IEEE 802.11 a/g, standards) to 6817 (as in the DVBT standard), the corresponding computational complexity will in general be significant. This paper presents two algorithms for interpolation-based QR decomposition in MIMO-OFDM systems. An in-depth computational complexity analysis shows that the proposed algorithms, for a sufficiently high number of data-carrying tones and small channel order, exhibit significantly smaller complexity than brute-force per-tone QR decomposition.
2005. 6th IEEE Workshop on Signal Processing Advances in Wireless Communications, New York, NY, Jun 05-08, 2005. p. 945 – 949. DOI : 10.1109/SPAWC.2005.1506279.2004
Conference Papers
Low complexity frequency domain equalization of MIMO channels with applications to MIMO-CDMA systems
CDMA and MIMO-CDMA systems using RAKE receivers are heavily limited by self- and multiple-access-interference. Linear equalization is a means to remove this interference, however it is often not practical due to the enormous complexity, especially in the MIMO case. This paper presents an approach to greatly reduce the complexity of linear MIMO equalizers. It discusses the complexity reduction of the equalizer itself and describes a suboptimal low-complexity method to compute its coefficients. The application of frequency domain equalization using the overlap/add FFT method to MIMO systems is suggested. The coefficients of the joint equalizer/MIMO detector are also derived in the frequency domain, based on an approximation of an MMSE criterion. Performance results in terms of BER are quantified through simulations of a MIMO-extended UMTS-FDD downlink.
2004. 58th IEEE Vehicular Technology Conference (VTC 2003), Orlando, FL, Oct 06-09, 2003. p. 468 – 472. DOI : 10.1109/VETECF.2003.1285061.VLSI implementation of the sphere decoding algorithm
Maximum likelihood detection is an essential part of high-performance multiple-input-multiple-output (MIMO) communication systems. While it is attractive due to its superior performance (in terms of BER) its complexity using a straightforward exhaustive search grows exponentially with the number of antennas and the order of the modulation scheme. Sphere decoding is a promising method to reduce the average decoding complexity significantly without compromising performance. This paper discusses the VLSI implementation of the sphere decoder and presents the first implementation of the algorithm that does not compromise BER performance.
2004. 30th European Solid-State Circuits Conference (ESSCIRC 2004), Leuven, BELGIUM, Sep 21-23, 2004. p. 303 – 306. DOI : 10.1109/ESSCIR.2004.1356678.Performance of MIMO-extended UMTS-FDD downlink comparing space-time RAKE and linear equalizer
The paper discusses the performance of a MIMO-extended UMTS-FDD downlink with a 2×2 and a 4×4 antenna configuration. An analytical expression for the average signal to interference ratio (SIR) of a space-time (ST) RAKE receiver in an interference dominated environment is derived. Simulations are used to assess the BER performance with the number of concurrent users as a parameter. The results are compared to a single antenna link with equivalent data rate. In addition to that, a linear zero forcing (ZF) and an MMSE MIMO equalizer receiver is evaluated and compared to the ST-RAKE. It is shown that both methods can successfully mitigate the significant performance Imitation due to self interference (SI) and multiple access interference (MAI).
2004. 58th IEEE Vehicular Technology Conference (VTC 2003), Orlando, FL, Oct 06-09, 2003. p. 473 – 477. DOI : 10.1109/VETECF.2003.1285062.2003
Journal Articles
Rapid prototyping for wireless designs: the five-ones approach
In a highly innovative market, wireless systems nowadays undergo very short production cycles. Due to these tough timing constraints, the time-consuming process of prototyping is often neglected, jeopardizing the entire product becoming successful. Heavy application of automatic tools can allow for rapid prototyping overcoming this unfortunate situation and de-risking the product challenge. However, the application of automatic tools alone does not speed up the prototyping process sufficiently. By reflecting on current design processes, several paradigms for faster prototyping are concluded, named the Five-Ones Approach: One team, One environment, One code, One documentation and One code revision tool. Based on such a Five-Ones Approach, a consistent prototyping environment to implement a prototyping design from first idea to final implementation is presented in this paper. In particular, the design of a prototype for a MIMO system with four transmit and four receive antennas, based on the current UMTS FDD downlink standard is reported. (C) 2003 Elsevier Science B.V. All rights reserved.
Signal Processing. 2003. Vol. 83, p. 1427 – 1444. DOI : 10.1016/S0165-1684(03)00090-2.Conference Papers
Efficient ASIC implementation of a real-time depth mapping stereo vision system
This paper presents a fast and area-efficient implementation of a real-time stereo vision algorithm for spatial depth mapping. The design combines two well-known area-based approaches to stereo thatching and includes an occlusion detection method. Hardware efficiency is achieved by storing only partial images on-chip, avoiding full-sized frame buffers. A low-latency dataflow-oriented structure makes it possible to process 256 x 192 pixel Input streams with a rate In excess of 50 frames per second, amounting to more than 54 million pixel x disparity measurements per second (PDS) (for a 25-pixel disparity range), or roughly 18 GOPS. The design has been Integrated In a 0.25 mu m standard CMOS technology and occupies an area of less than 3 mm(2).
2003. 46th IEEE International Midwest Symposium on Circuits and Systems, Cairo, EGYPT, Dec 27-30, 2003. p. 1478 – 1481. DOI : 10.1109/MWSCAS.2003.1562575.Variable delay ripple carry adder with carry chain interrupt detection
A statistical approach for the area efficient implementation of fast wide operand adders using early termination detection is described and analyzed. It is shown that high throughput can be achieved based on area- and routing-efficient ripple-carry adders with only marginal overhead. They share a low AT-product with Brent-Kung adders but provide designers with totally different area/delay tradeoffs. The circuit does not require full-custom design and fits well into both self-timed and synchronous designs.
2003. IEEE International Symposium on Circuits and Systems, BANGKOK, THAILAND, May 25-28, 2003. p. 113 – 116. DOI : 10.1109/ISCAS.2003.1206202.A 50 MBPS 4×4 maximum likelihood decoder for multiple-input multiple-output systems with QPSK modulation
In this article we present an efficient approach for the implementation of optimum maximum likelihood decoding of QPSK modulated multiple-input-multiple-output data streams. The proposed method does not compromise optimality of the detection algorithm. Instead it uses the special properties of QPSK modulation together with algebraic transformations and architectural optimizations to achieve very low hardware complexity and high speed. To our knowledge it is the fastest and most area efficient reported VLSI implementation of a hard decission Maximum Likelihood decoder for QPSK based MIMO systems. It can be applied to a variety of narrow band and wideband systems in many different configurations, including different degrees of spatial multiplexing and receive diversity.
2003. 10th IEEE International Conference on Electronics, Circuits and Systems, Sharjah, U ARAB EMIRATES, Dec 14-17, 2003. p. 332 – 335. DOI : 10.1109/ICECS.2003.1302044.Programmable code processor for software defined radio
This paper describes a flexible processor capable of producing binary codes for various standards such as UNITS and 802.11b. Its field of application lies in base-stations and in future software defined radio terminals. Because of its flexibility just one or two instances may be integrated into a SoC (System on Chip) where multiple codes are needed. This approach adds flexibility compared to a dedicated code generating structure where the class of codes is fixed. Specialized Bit-ALUs allow to address multiple bits of a register and to operate simultaneously on them. Instructions tailored for code generation enhance its efficiency considerably. The processor was successfully integrated and tested in a 0.25mum 5ML CMOS process.
2003. 37th Asilomar Conference on Signals, Systems and Computers, Pacific Grove, CA, Nov 09-12, 2003. p. 2156 – 2160. DOI : 10.1109/ACSSC.2003.1292362.Practical low complexity linear equalization for MIMO-CDMA systems
This article first reviews recently proposed techniques for adaptive and direct linear MIMO equalization in the context of MIMO-CDMA systems and in particular with application to a MIMO-extended UMTS-FDD downlink. The focus is thereby mainly on the complexity of the algorithms. The second part of the paper proposes frequency domain (FD) MIMO equalization using the overlap/add FFT method in conjunction with two different low-complexity FD-deconvolution techniques to obtain the equalizer coefficients based on explicit channel impulse response estimates. The effects of imperfect channel estimation are discussed. An architecture for the VLSI implementation of the proposed method is suggested and an estimate of the complexity of the proposed circuit is given in the conclusions.
2003. 37th Asilomar Conference on Signals, Systems and Computers, Pacific Grove, CA, Nov 09-12, 2003. p. 1266 – 1272. DOI : 10.1109/ACSSC.2003.1292192.2002
Conference Papers
An ASIC implementation of adaptive arithmetic coding
In this work, we present an improved version of an ASIC implementation of the adaptive arithmetic coding algorithm which uses a two-level memory hierarchy. We propose algorithmic modifications and a special hardware structure to speed-up the design without degrading the compression ratio obtained using this memory hierarchy. Moreover, several new features which increase the compression efficiency are introduced. Finally, a VLSI implementation based on the results of our work is presented.
2002. 36th Asilomar Conference on Signals, Systems and Computers, PACIFIC GROVE, CA, Nov 03-06, 2002. p. 1078 – 1083. DOI : 10.1109/ACSSC.2002.1196950.2001
Conference Papers
Rapid prototyping design of a 4×4 BLAST-over-UMTS system
BLAST techniques to increase the utilized bandwidth in commercial systems are currently feasible. This paper describes the design of a UMTS prototype, supporting four transmit and four receive antennas, achieving almost four times the capacity of a conventional system. Various parts of the transmitter and receiver are mapped on FPGAs and fast DSPs communicating via a specially designed communication link. The system was entirely designed using C code, embedded in SIMULINK S-functions for simulation, and, after validation, automatically mapped onto the hardware platform.
2001. 35th Asilomar Conference on Signals, Systems and Computers, PACIFIC GROVE, CA, Nov 04-07, 2001. p. 1256 – 1260. DOI : 10.1109/ACSSC.2001.987692.2000
Conference Papers
A 3D-DCT real-time video compression system for low complexity single-chip VLSI implementation
2000. Mobile Multimedia Conference, MoMuC 2000, Tokyo, Japan, November, 2000.From basic concept to real-time implementation: Prototyping WCDMA downlink receiver algorithms – A case study
In this paper we present an approach to rapid prototyping of advanced signal processing techniques for future wireless applications currently being adopted within Bell Labs Research. The aim of the “Bell Labs Algorithm Development and Evaluation ” (BLADE) initiative is to devise a design framework specifically targeting the needs land capabilities) of high-level algorithm designers (viz. communication engineers), which enables them to quickly, “translate” a research idea into a working real-time system for practical experimentation purposes. The mixed DSP/FPGA implementation of a WCDMA testbed is used as an example to describe our initial experience with the proposed design methodology, which is based on a set of commercially available software tools and a platform consisting of off the shelf hardware modules.
2000. 34th Asilomar Conference on Signals, Systems, and Computers, PACIFIC GROVE, CA, Oct 29-Nov 01, 2000. p. 84 – 88. DOI : 10.1109/ACSSC.2000.910922.