1 of 15

94

Optimized Multi-Processor System-on-Chip (MPSoC) Design for Low-Resource JPEG Encoding

Kanishka Gunawardana*, Chandula Adhikari*, Isuru Nawinne*� *Department of Computer Engineering, Faculty of Engineering, University of Peradeniya

2 of 15

JPEG Encoding

  • Method of lossy compression for digital images.
  • JPEG compression reduces file size by discarding some image data, which can result in a slight loss of quality.
  • Supports 8-bit and 12-bit per color channel, but 8-bit is more common.

3 of 15

Multi-Processor Systems

  • Two or more processors that work together to perform tasks.
  • By dividing the workload, multiprocessor systems can perform encoding tasks more efficiently, saving energy and resources.

4 of 15

Research Problem

Computational Intensity: multiple complex steps which require significant computational power

Hardware Limitations: Embedded systems and low-power devices often have limited hardware capabilities

Memory Requirements: The process needs substantial memory to store intermediate data and perform operations

5 of 15

Research Objective

Primary Goal: Optimize MPSoC Architecture for JPEG Encoding

Specific Objectives:

  • Reduce resource usage
  • Address performance bottlenecks
  • Enhance memory management

6 of 15

Literature Review

Reviewed multiple research studies focusing on heterogeneous

multiprocessor configurations

eg:-

  • Synthesis of heterogeneous pipelined multiprocessor systems using ILP: JPEG case study - H. Javaid and S. Parameswaran
  • Design Exploration for FPGA-Based Multiprocessor Architecture: JPEG Encoding Case Study - J. Wu, J. Williams, N. Bergmann and P. Sutton

7 of 15

MPSoC Architecture for JPEG Encoding

8 of 15

Methodology - Custom Hardware Components

Three Key Optimizations:

  1. Custom Instruction for DCT
  2. Custom Multiplication Instruction for Quantization
  3. Custom FIFOs for Level Shifting

9 of 15

Methodology - Memory Management

Two-Step Memory Optimization Strategy

  1. Reducing Code Memory Footprint
  2. On-Chip Memory Utilization
  3. Relocate memory-intensive stages to on-chip memory
  4. Replace SDRAM with on-chip memory for Color Space Conversion

10 of 15

Methodology - Pipeline Enhancements

Superscalar Pipelines

  • Introduced for DCT and Quantization stages

FIFO Depth Adjustment

  • Set to 128 to accommodate 8x8 block processing
  • Minimize idle times and Improve data flow

11 of 15

12 of 15

13 of 15

Results and Performance Improvement

Initial System Performance

  • Throughput: 11,100 bytes/second
  • DCT stage identified as major bottleneck

Optimized System Performance

  • Throughput increased to 31,000 bytes/second
  • 2.78x performance improvement

Resource Utilization

  • Logic Elements: 22% of total
  • Memory Bits: 79% of total
  • Embedded Multiplier: 48% utilized

14 of 15

Conclusion

Key Achievements

  • Significant performance improvements
  • Resource-efficient design

Future Work

  • Explore further optimizations (Huffman Encoding)

15 of 15

THANK YOU