imageresizer_web_logo

ImageResizer C++ library

v2.1.0

Table of contents

Overview

ImageResizer C++ library provides image resizing. The library contains two mutually exclusive implementations. The software implementation uses SIMD-optimized bicubic interpolation (Keys cubic, a = -0.75). The library implements a two-pass separable filter with platform-specific vectorization: SSSE3 + AVX2 on x86/x64, NEON on ARM64, and a scalar fallback for other architectures. Tile-based processing ensures L2 cache locality for upscale paths, while a row-cache strategy is used for downscale. The G2D implementation is a wrapper around the G2D API, which is used to work with the i.MX GPU. Resizing is single threaded by default and can optionally be spread over image rows with OpenMP, see Multithreading. Apart from OpenMP (supplied with the compiler) the library doesn’t have third-party dependencies. Input data must be 3-byte-per-pixel interleaved format (BGR24, RGB24 and YUV24). The library uses C++17 standard and is compatible with Linux and Windows.

Versions

Table 1 - Library versions.

Version Release date What’s new
1.0.0 03.04.2026 First version of the library.
1.0.1 05.04.2026 Fixed alignments in data structures.
2.0.0 23.08.2026 - setMaxThreads(…) and getMaxThreads() methods added to control how many CPU cores the resizer may use.
- Breaking change: the numThreads parameter of resize(…) is removed, the thread count is set with setMaxThreads(…) only.
- OpenMP is optional now.
- Resizing to the same size is a plain copy now instead of a full filter pass (~8 times faster).
- Fixed out-of-bounds read in the ARM64 NEON horizontal pass for source images narrower than 6 pixels.
- Fixed undefined behaviour (misaligned 32-bit store) in the x86 horizontal pass.
- Fixed the build in the Debug, RelWithDebInfo and MinSizeRel configurations on GCC/Clang: the SSSE3/AVX2 flags were applied to Release only while the kernels use the intrinsics unconditionally.
2.1.0 04.09.2026 - Added a G2D pipeline for i.MX.

Library files

The library supplied by source code only. The user would be given a set of files in the form of a CMake project (repository). The repository structure is shown below:

CMakeLists.txt ------------------- Main CMake file of the library.
src ------------------------------ Folder with source code of the library.
    CMakeLists.txt --------------- CMake file of the library.
    ImageResizer.cpp ------------- C++ implementation file.
    ImageResizer.h --------------- Header file which includes ImageResizer class declaration.
    ImageResizerVersion.h -------- Header file which includes version of the library.
    ImageResizerVersion.h.in ----- CMake service file to generate version file.
example -------------------------- Folder for example application.
    CMakeLists.txt --------------- CMake file for example application.
    main.cpp --------------------- Source code file of example application.
static --------------------------- Folder with static resources.
    imageresizer_web_logo.png ---- Web logo image.

ImageResizer class description

ImageResizer class declaration

The ImageResizer class declared in ImageResizer.h file. Class declaration:

namespace cr
{
namespace video
{
/// Image resizer class.
class ImageResizer
{
public:

    /// Get the version of the ImageResizer class.
    static std::string getVersion();

    /// Resize image.
    bool resize(uint8_t* src, int srcWidth, int srcHeight,
                uint8_t* dst, int dstWidth, int dstHeight);

    /// Set upper limit for the number of threads used by resize(...).
    void setMaxThreads(int maxThreads);

    /// Get the number of threads resize(...) may use at most.
    int getMaxThreads() const;
};
}
}

getVersion method

The getVersion() static method returns string of current class version. Method declaration:

static std::string getVersion();

Method can be used without ImageResizer class instance:

std::cout << "ImageResizer version: " << cr::video::ImageResizer::getVersion();

Console output:

ImageResizer version: 2.1.0

resize method

The resize(…) method performs image resizing with bicubic interpolation. Method declaration:

bool resize(uint8_t* src, int srcWidth, int srcHeight,
            uint8_t* dst, int dstWidth, int dstHeight);
Parameter Description
src Pointer to source image data. Only 3-byte-per-pixel interleaved data (BGR24, RGB24, YUV24).
srcWidth Width of source image in pixels. Must be > 0.
srcHeight Height of source image in pixels. Must be > 0.
dst Pointer to destination image data buffer. Must be pre-allocated with size >= dstWidth * dstHeight * 3 bytes.
dstWidth Width of destination image in pixels. Must be > 0.
dstHeight Height of destination image in pixels. Must be > 0.

Returns: TRUE if resizing was successful, FALSE otherwise (invalid parameters).

By default the method resizes the whole image on the calling thread. Row-level parallelism is enabled with setMaxThreads, which is the only way to set the thread count, see Multithreading.

If the source and the destination sizes are equal the method just copies the image data: at scale 1.0 the cubic kernel degenerates to the weights {0, 1, 0, 0}, so the filtered result is byte identical to the source and the copy is the fastest way to produce it.

Example:

// Init image resizer.
ImageResizer resizer;

// Init source and destination image buffers.
std::vector<uint8_t> src(640 * 480 * 3);
std::vector<uint8_t> dst(1920 * 1080 * 3);

// Resize on the calling thread.
resizer.resize(src.data(), 640, 480, dst.data(), 1920, 1080);

// Allow 4 threads and resize again.
resizer.setMaxThreads(4);
resizer.resize(src.data(), 640, 480, dst.data(), 1920, 1080);

setMaxThreads method

The setMaxThreads(…) method enables parallel resizing and sets an upper limit for the number of threads that resize(…) may use. Parallel resizing is disabled by default, so this method has to be called to make the resizer use more than one thread. Method declaration:

void setMaxThreads(int maxThreads);
Parameter Description
maxThreads Maximum number of threads. Value 1 (the default state of the object) keeps resizing single threaded. Values <= 0 mean “as many threads as OpenMP makes available”. The limit is never larger than omp_get_max_threads().

Returns: none. The setting belongs to the ImageResizer object and must be applied before resizing starts (it is not synchronised with resize calls that are already running). The method also starts the OpenMP worker threads, so the first resize does not pay for the team start-up.

Example:

ImageResizer resizer;        // single threaded, nothing to configure

resizer.setMaxThreads(0);    // use every thread OpenMP offers
resizer.setMaxThreads(4);    // never use more than 4 threads
resizer.setMaxThreads(1);    // back to single threaded resizing

The numThreads parameter of resize(…) was removed in v2.0.0: setMaxThreads(…) is the only way to set the thread count, exactly as in the FormatConverter library.

To keep some CPU cores free for other software, pass the number of cores the resizer is allowed to occupy:

#include <thread>

// Leave 2 cores of the machine free.
const int cores = (int)std::thread::hardware_concurrency();
resizer.setMaxThreads(cores > 2 ? cores - 2 : 1);

getMaxThreads method

The getMaxThreads() method returns the thread limit that is in effect for the resizer. Method declaration:

int getMaxThreads() const;

Returns: the number of threads resize(…) may use at most, after setMaxThreads(…) is applied to the value provided by OpenMP. The value is 1 until setMaxThreads(…) is called and is always >= 1. The team size of a particular resize can be smaller, because it also depends on the destination image height, see Multithreading.

Example:

ImageResizer resizer;
std::cout << resizer.getMaxThreads() << std::endl;   // 1

resizer.setMaxThreads(18);
std::cout << resizer.getMaxThreads() << std::endl;   // 18 on a 20 core machine

Multithreading

Parallel resizing is disabled by default. A freshly constructed ImageResizer resizes the whole image on the calling thread and does not create any OpenMP team. Call setMaxThreads to enable row-level parallelism. The thread control is identical to the one of the FormatConverter library, so both libraries are configured the same way in the same application.

resize(…) distributes the destination rows over the team with schedule(static), so each thread gets one contiguous band of rows and threads never share output bytes. The result does not depend on the number of threads: the same input always produces byte-identical output whether the resize runs on 1 thread or on 64.

The library never calls omp_set_num_threads(…) or omp_set_dynamic(…). The team size is passed to every parallel region with the OpenMP num_threads clause, which affects that region only. As a consequence:

  • the OpenMP configuration of your application is never modified by the library;
  • resize(…) may be called from several threads at the same time, but not through one ImageResizer object shared between threads, because the object caches the interpolation tables of the last used geometry. Give every thread its own ImageResizer instance;
  • calling resize(…) from inside your own #pragma omp parallel region is safe.

Number of threads

For every resize the team size is calculated as:

threads = min( getMaxThreads(),   // limit of the resizer
               dstHeight )        // cannot exceed the number of destination rows

Since getMaxThreads() returns 1 until setMaxThreads(…) is called, the default resizer always ends up with a single thread.

Table 2 - Team size on a 20 core machine after setMaxThreads(0).

Destination size Destination rows Threads used Threads used by default
64 x 8 8 8 1
320 x 240 240 20 1
1920 x 1080 1080 20 1
3840 x 2160 2160 20 1

Enabling and limiting the number of threads

Table 3 - Ways to control the thread count, from the widest scope to the narrowest.

Way Scope Description
Build without OpenMP Whole library If the compiler is invoked without OpenMP support the library builds and runs single threaded; setMaxThreads(…) is accepted but getMaxThreads() always returns 1.
OMP_NUM_THREADS environment variable Whole process Standard OpenMP setting. It defines the value returned by omp_get_max_threads(), which is the upper bound of what the resizer can request.
setMaxThreads(n) One ImageResizer Enables parallel resizing and sets a hard upper limit for that resizer object. n = 0 means all available threads, n = 1 means single threaded.

The limit is always at least 1 and never larger than omp_get_max_threads().

Example - a 20 core machine, 2 cores must stay free for other software:

#include <thread>
#include "ImageResizer.h"

cr::video::ImageResizer resizer;

const int cores = (int)std::thread::hardware_concurrency();
resizer.setMaxThreads(cores > 2 ? cores - 2 : 1);

// 18
std::cout << resizer.getMaxThreads() << std::endl;

Example - ImageResizer and FormatConverter sharing one thread budget:

cr::video::ImageResizer    resizer;
cr::video::FormatConverter converter;

const int cores = (int)std::thread::hardware_concurrency();
const int budget = cores > 2 ? cores - 2 : 1;

resizer.setMaxThreads(budget);
converter.setMaxThreads(budget);

Example - limiting the whole process from the command line:

OMP_NUM_THREADS=6 ./YourApplication

Performance

The library’s software implementation was compared against cv::resize(..., cv::INTER_CUBIC), which implements the same Keys cubic kernel (a = -0.75) with the same (dst + 0.5) * scale - 0.5 sample mapping.

The comparison is made against two OpenCV builds, because the difference between them is larger than the difference between OpenCV and this library:

  • OpenCV 5.0.0 official wheel, built with -O3 and Intel IPP 2026.0.0. cv::resize is dispatched to ippiResize here, which is the fastest resize implementation OpenCV can offer. This is the honest opponent and the numbers below should be read from this column.
  • OpenCV 4.6.0 distribution package (Ubuntu libopencv-dev). It is built without IPP, with -O2 and with hardening flags (-D_FORTIFY_SOURCE=3, -fstack-protector-strong, -fcf-protection, -fno-omit-frame-pointer), and falls back to OpenCV’s own SIMD kernels. Enabling IPP alone accounts for a 1.8x - 2.8x difference, so measuring against a distribution package overstates any advantage.

Both libraries were given the same number of threads (cv::setNumThreads(n) and setMaxThreads(n)) and were measured in separate processes, because letting the OpenMP team and the OpenCV thread pool run interleaved makes them fight for the same cores and distorts both results.

Test conditions: Intel Core i7-13700H (10 cores / 20 threads), Ubuntu 24.04, GCC 13.3, -O3 -march=native -funroll-loops -mssse3 -mavx2, BGR24 input, median of 51 runs.

Table 4 - Single thread, time per frame in microseconds.

Conversion ImageResizer OpenCV 5 + IPP Speedup OpenCV, no IPP Speedup
640x512 → 1920x1080 1418 2397 1.7x 5369 3.8x
1280x720 → 3840x2160 5197 7142 1.4x 17234 3.3x
320x240 → 1920x1080 1122 1790 1.6x 3306 2.9x
1920x1080 → 640x512 490 916 1.9x 2546 5.2x
3840x2160 → 1920x1080 3408 6068 1.8x 16506 4.8x
3840x2160 → 640x360 797 1255 1.6x 3182 4.0x
1920x1080 → 1920x1080 247 239 1.0x 217 0.9x

Table 5 - 4 threads, time per frame in microseconds.

Conversion ImageResizer OpenCV 5 + IPP Speedup OpenCV, no IPP Speedup
640x512 → 1920x1080 396 740 1.9x 1804 4.6x
1280x720 → 3840x2160 1478 2772 1.9x 5563 3.8x
320x240 → 1920x1080 341 786 2.3x 1182 3.5x
1920x1080 → 640x512 134 288 2.1x 1147 8.5x
3840x2160 → 1920x1080 1538 1946 1.3x 4755 3.1x
3840x2160 → 640x360 186 360 1.9x 953 5.1x
1920x1080 → 1920x1080 66 253 3.8x 267 4.0x

Against the strongest OpenCV build the library is roughly 1.4x - 2.3x faster. The margin comes from specialisation rather than from better vectorisation: this library only handles 8-bit 3-channel interleaved data in contiguous buffers with one interpolation mode, so it can keep the whole horizontal pass in 8-bit PMADDUBSW form and the vertical pass in int16, while cv::resize supports every depth, channel count, ROI stride and interpolation mode and pays for that generality. The same-size case is a plain copy in both libraries, so at one thread it is a tie and the difference at 4 threads only reflects how the copy is split.

Table 6 - Thread scaling of the library itself, time per frame in microseconds.

Conversion 1 thr 2 thr 4 thr 8 thr 20 thr Best speedup
640x512 → 1920x1080 1378 745 376 393 242 5.7x
1280x720 → 3840x2160 5082 2826 1418 1447 888 5.7x
1920x1080 → 640x512 533 285 125 145 86 6.2x
3840x2160 → 1920x1080 3428 2041 1185 853 539 6.4x

Scaling saturates around 6x on a 10 core machine because the filter is memory bandwidth bound rather than compute bound: at 4K the two passes move well over 100 MB per frame.

Accuracy versus OpenCV

The implementations use different fixed-point precision for the filter coefficients (6 fractional bits here, 11 in OpenCV), so the results differ by a fraction of an intensity level. The lower precision is part of the reason this library is faster: 6-bit weights fit in a byte and allow the 8-bit multiply-accumulate path. OpenCV’s own IPP and generic kernels differ from each other by up to 1 level, so the numbers below are the same against either.

Table 7 - Difference against cv::resize(..., cv::INTER_CUBIC), BGR24.

Conversion Max difference Mean difference PSNR
640x512 → 1920x1080 5 0.41 51.1 dB
1280x720 → 3840x2160 3 0.34 52.2 dB
320x240 → 1920x1080 5 0.54 49.1 dB
1920x1080 → 640x512 4 0.20 53.9 dB
3840x2160 → 1920x1080 1 0.65 50.0 dB
3840x2160 → 640x360 1 0.65 50.0 dB
1920x1080 → 1920x1080 0 0.00 exact

The output is also bit identical regardless of the thread count and of the SIMD backend: the x86 SSSE3/AVX2, the ARM64 NEON and the scalar kernels produce the same bytes for the same input.

Several libraries in one application

Different thread limits in different libraries do not conflict. The limit set with setMaxThreads(…) is a field of the library object, not a global setting, and it reaches OpenMP only through the num_threads clause of a particular parallel region. A num_threads clause applies to the region it is attached to and to nothing else, so an ImageResizer configured for 4 threads and a FormatConverter configured for 16 threads simply request different team sizes from the same shared thread pool:

cr::video::ImageResizer    resizer;
cr::video::FormatConverter converter;

resizer.setMaxThreads(4);     // resize(...) regions run on 4 threads
converter.setMaxThreads(16);  // convert(...) regions run on 16 threads

// 4 and 16, neither call changes the setting of the other library.
std::cout << resizer.getMaxThreads() << " " << converter.getMaxThreads();

This holds because neither library calls omp_set_num_threads(…) or omp_set_dynamic(…). Those functions write the process wide nthreads-var / dyn-var internal control variables, which is exactly what would make two libraries overwrite each other’s setting and would also silently change the OpenMP configuration of the host application.

Table 4 - What is shared between the libraries and what is not.

Item Shared Comment
m_maxThreads of the object No Private field, one per ImageResizer / FormatConverter object.
num_threads clause No Applies to a single parallel region only.
OpenMP worker thread pool Yes The runtime keeps one pool per thread. Teams are formed from it on demand, so different team sizes are fine.
omp_get_max_threads() Yes Upper bound for both libraries, controlled by OMP_NUM_THREADS or by the application. Neither library modifies it.

Two points still deserve attention:

  1. Oversubscription. The limits are independent, so they add up when the libraries work at the same time from different application threads. On a 20 core machine resizer.setMaxThreads(16) plus converter.setMaxThreads(16) may put 32 threads on 20 cores. This is a scheduling inefficiency, not a correctness problem, but for a pipeline that runs both libraries concurrently give them a shared budget:

    const int budget = (int)std::thread::hardware_concurrency() / 2;
    resizer.setMaxThreads(budget);
    converter.setMaxThreads(budget);
    

    When the libraries are used one after another in the same pipeline stage they never run at the same time and both can be given the full budget.

  2. One OpenMP runtime per process. Mixing several OpenMP runtimes in one binary (for example GCC libgomp together with LLVM libomp, or MSVC vcomp together with Intel libiomp5) gives each runtime its own thread pool and leads to oversubscription and, in some combinations, to crashes. This is a linking issue and not related to setMaxThreads(…): build all the libraries of the application with the same compiler and OpenMP implementation, which is what the CMake target OpenMP::OpenMP_CXX used by both libraries does.

Nested parallelism follows the usual OpenMP rules: if resize(…) is called from inside a #pragma omp parallel region of the application, the inner region is executed by a single thread unless the application enables nesting.

Build and connect to your project

Typical commands to build ImageResizer:

cd ImageResizer
mkdir build
cd build
cmake ..
make

By default the library is built in Release mode with aggressive compiler optimizations:

  • MSVC: /O2 /Ob3 /GL /Gy /LTCG
  • GCC/Clang: -O3 -funroll-loops -march=native -mtune=native

On x86/x64 the SSSE3 and AVX2 kernels are compiled in every configuration (/arch:AVX2 for MSVC, -mssse3 -mavx2 for GCC/Clang), because the intrinsics are used unconditionally. AVX2 is therefore a hard requirement of the x86 build; on ARM64 the NEON kernels are used instead and no extra flag is needed. -march=native tunes the Release build for the machine it is compiled on, so the resulting binary is not meant to be copied to a machine with an older CPU.

If you want to cross-build the library for i.MX, you need to execute the following commands:

source /opt/tdx-xwayland/7.3.0/environment-setup-cortexa55-tdx-linux
cd ImageResizer
mkdir build
cd build
cmake ..
make

To build and install the BSP, follow the instructions.

If you want to add ImageResizer to your CMake project as source code, you can do the following. For example, if your repository has the following structure:

CMakeLists.txt
src
    CMakeLists.txt
    yourLib.h
    yourLib.cpp

Create folder 3rdparty in your repository and copy ImageResizer repository folder there. New structure of your repository:

CMakeLists.txt
src
    CMakeLists.txt
    yourLib.h
    yourLib.cpp
3rdparty
    ImageResizer

Create CMakeLists.txt file in 3rdparty folder. CMakeLists.txt should contain:

cmake_minimum_required(VERSION 3.13)

################################################################################
## 3RD-PARTY
## dependencies for the project
################################################################################
project(3rdparty LANGUAGES CXX)

################################################################################
## SETTINGS
## basic 3rd-party settings before use
################################################################################
# To inherit the top-level architecture when the project is used as a submodule.
SET(PARENT ${PARENT}_YOUR_PROJECT_3RDPARTY)
# Disable self-overwriting of parameters inside included subdirectories.
SET(${PARENT}_SUBMODULE_CACHE_OVERWRITE OFF CACHE BOOL "" FORCE)

################################################################################
## INCLUDING SUBDIRECTORIES
## Adding subdirectories according to the 3rd-party configuration
################################################################################
if (${PARENT}_SUBMODULE_IMAGE_RESIZER)
    add_subdirectory(ImageResizer)
endif()

File 3rdparty/CMakeLists.txt adds folder ImageResizer to your project and excludes the example application from compiling (by default it is excluded when ImageResizer included as sub-repository). Your repository new structure will be:

CMakeLists.txt
src
    CMakeLists.txt
    yourLib.h
    yourLib.cpp
3rdparty
    CMakeLists.txt
    ImageResizer

Next you need include folder 3rdparty in main CMakeLists.txt file of your repository. Add string at the end of your main CMakeLists.txt:

add_subdirectory(3rdparty)

Next you have to include ImageResizer library in your src/CMakeLists.txt file:

target_link_libraries(${PROJECT_NAME} ImageResizer)

Done!

Example

The example resizes a synthetic image from 640x512 to 1920x1080 and measures execution time:

#include <iostream>
#include <cstdint>
#include <vector>
#include <chrono>
#include <algorithm>
#include "ImageResizer.h"

using namespace std;
using namespace cr::video;
using namespace std::chrono;

int main(void)
{
    cout << "ImageResizer v" << ImageResizer::getVersion() << endl;

    // Source image size.
    const int srcWidth = 640;
    const int srcHeight = 512;

    // Destination image size.
    const int dstWidth = 1920;
    const int dstHeight = 1080;

    // Create source image with deterministic pattern.
    vector<uint8_t> srcImg(srcWidth * srcHeight * 3);
    for (int y = 0; y < srcHeight; ++y)
        for (int x = 0; x < srcWidth; ++x) {
            int idx = (y * srcWidth + x) * 3;
            srcImg[idx + 0] = (uint8_t)((x * 37 + y * 13) & 0xFF);
            srcImg[idx + 1] = (uint8_t)((x * 59 + y * 7)  & 0xFF);
            srcImg[idx + 2] = (uint8_t)((x * 23 + y * 41) & 0xFF);
        }

    // Create destination image buffer.
    vector<uint8_t> dstImg(dstWidth * dstHeight * 3);

    // Image resizer.
    ImageResizer imageResizer;

    // Allow the resizer to occupy up to 4 CPU cores. Without this call the
    // resize(...) method runs entirely on the calling thread.
    imageResizer.setMaxThreads(4);
    cout << "Max threads: " << imageResizer.getMaxThreads() << endl;

    // Warmup run (includes one-time table build).
    bool result = imageResizer.resize(srcImg.data(), srcWidth, srcHeight,
                                      dstImg.data(), dstWidth, dstHeight);

    // Benchmark: run multiple times and report median.
    const int N = 21;
    vector<long long> times(N);
    for (int i = 0; i < N; i++) {
        auto begin = steady_clock::now();
        imageResizer.resize(srcImg.data(), srcWidth, srcHeight,
                            dstImg.data(), dstWidth, dstHeight);
        auto end = steady_clock::now();
        times[i] = duration_cast<microseconds>(end - begin).count();
    }
    sort(times.begin(), times.end());

    cout << "Result: " << (result ? "OK" : "FAIL") << endl;
    cout << "Resize time (median of " << N << " runs): " <<
    times[N / 2] << " us" << endl;

    return 0;
}

Table of contents