A modern, high-performance C++20 library designed for game hacking in user space. Features an incredibly fast SIMD-accelerated signature scanner, and a set of platform-agnostic APIs for interfacing with loaded modules, modifying virtual memory protections, and more.
- Vectorized scanning for byte patterns
- SSE 4.1 and AVX2 on x86/x64
- AVX-512 on x64
- Neon on ARM/ARM64
- RAII memory protector
- Convenience wrappers over OS APIs
- Language bindings (C, Java, Python, C#)
- Supports Windows, Linux, macOS, and Android
- General-purpose library features
hat::cowhat::cstring_viewhat::fixed_stringhat::string_literal
This project adheres to semantic versioning. Any declaration that
is within a detail or experimental namespace is not considered part of the public API, and usage
may break at any time without the MAJOR version number being incremented.
The best supported way to use libhat is via FetchContent or CPM in a CMake project:
FetchContent_Declare(
libhat
GIT_REPOSITORY https://github.com/BasedInc/libhat.git
GIT_TAG v0.11.1
)
FetchContent_MakeAvailable(libhat)
target_link_libraries(my_target libhat::libhat)CPMAddPackage("gh:BasedInc/libhat#v0.11.1")
target_link_libraries(my_target libhat::libhat)If you are using xmake, an official package is available:
add_requires("libhat v0.10.0")
target("my_target")
add_packages("libhat")If you are using MSBuild or another build system, a vcpkg port is available:
{
"dependencies": [
"libhat"
]
}Pre-built binaries are also available with releases, which include
archives with the full set of install files for libhat and libhat_c (shared).
If you are exclusively building libhat from source, and do not need tests or examples, set the following options when generating the buildsystem:
-DLIBHAT_TESTING=OFF -DLIBHAT_EXAMPLES=OFF
If you want to include the C bindings in your build, this must be done via the root CMake project. Specify
either -DLIBHAT_STATIC_C_LIB=ON or -DLIBHAT_SHARED_C_LIB=ON to build libhat_c as a static or shared library,
respectively. (Note that only one version of the library can be built per build configuration).
If you are using a multi-config generator, such as Visual Studio:
cmake -B build
cmake --build build --config Release
cmake --install build --prefix installIf you are using a single-config generator, such as Ninja:
cmake -B build -DCMAKE_BUILD_TYPE=Release -G Ninja
cmake --build build
cmake --install build --prefix installBoth options will output the headers, libraries, and binaries into a local install directory.
Currently, libhat only supports static linking (dynamic linking can be achieved using the C bindings). As a result, features are configured through CMake options. By default, all features are enabled:
Enables support for the SSE vectorized find_pattern implementation, allowing CPUs that don't support newer instruction
sets (AVX families) to get higher throughput than the base standard library implementation. If compile-time support is
not present (-march or /ARCH), runtime support is validated through cpuid. If the target architecture is not x86
or x86_64, enabling this option has no effect.
CPU features: sse4.1
Enables support for the AVX2 vectorized find_pattern implementation. If compile-time support is not present
(-march or /ARCH), runtime support is validated through cpuid. If the target architecture is not x86
or x86_64, enabling this option has no effect.
CPU features: avx avx2 bmi
Enables support for the AVX512 (BW+F) vectorized find_pattern implementation. If compile-time support is not present
(-march or /ARCH), runtime support is validated through cpuid. If the target architecture is not x86_64,
enabling this option has no effect.
CPU features: avx512bw avx512f bmi
Enables support for hat::scan_hint::x86_64, which allows informed anchor selection when searching for patterns in
x86_64 machine code, greatly reducing search time at the cost of a small data table (see Benchmarks).
This option is always supported regardless of the target architecture.
Enables support for hat::scan_hint::aarch64, which allows informed anchor selection when searching for patterns in
aarch64 machine code, greatly reducing search time at the cost of a small data table. This option is always supported
regardless of the target architecture.
The table below compares the single threaded throughput in bytes/s (real time) between
libhat and two other commonly used implementations for pattern
scanning. The input buffers were randomly generated using a fixed seed, and the pattern
scanned does not contain any match in the buffer. The benchmark was compiled on Windows
with clang-cl 22.1.1, using the MSVC 14.51.36231 toolchain and the default release mode
flags (/GR /EHsc /MD /O2 /Ob2). The benchmark was run on a system with an i7-14700K
(supporting AVX2) and 64GB (4x16GB) DDR5 6000 MT/s (30-38-38-96).
The full source code is available here.
---------------------------------------------------------------------------------------------------
Benchmark Time CPU Iterations bytes_per_second
---------------------------------------------------------------------------------------------------
libhat+AVX2/4MiB 70877 ns 70998 ns 79007 55.1129Gi/s
libhat+AVX2/8MiB 142373 ns 141935 ns 38750 54.8734Gi/s
libhat+AVX2/16MiB 310238 ns 308405 ns 17783 50.3645Gi/s
libhat+AVX2/32MiB 1002988 ns 999485 ns 5581 31.1569Gi/s
libhat+AVX2/64MiB 2452875 ns 2449913 ns 2296 25.4803Gi/s
libhat+AVX2/128MiB 5223758 ns 5183858 ns 1064 23.9291Gi/s
libhat+AVX2/256MiB 10713488 ns 10637019 ns 520 23.3351Gi/s
std::search/4MiB 1348946 ns 1342598 ns 4178 2.89578Gi/s
std::search/8MiB 2690336 ns 2688011 ns 2081 2.90391Gi/s
std::search/16MiB 5383973 ns 5353287 ns 1042 2.90213Gi/s
std::search/32MiB 10807674 ns 10715996 ns 522 2.89146Gi/s
std::search/64MiB 21520222 ns 21454327 ns 260 2.90425Gi/s
std::search/128MiB 43072097 ns 42999031 ns 129 2.90211Gi/s
std::search/256MiB 86326666 ns 85817308 ns 65 2.89598Gi/s
csgosimple/4MiB 1704462 ns 1695994 ns 3289 2.29178Gi/s
csgosimple/8MiB 3437443 ns 3420064 ns 1631 2.27277Gi/s
csgosimple/16MiB 6867328 ns 6835938 ns 816 2.27527Gi/s
csgosimple/32MiB 13715049 ns 13686131 ns 411 2.27852Gi/s
csgosimple/64MiB 27225719 ns 27210366 ns 205 2.29562Gi/s
csgosimple/128MiB 54692310 ns 54227941 ns 102 2.28551Gi/s
csgosimple/256MiB 108922873 ns 108762255 ns 51 2.29520Gi/s
learn_more/4MiB 4091135 ns 4090909 ns 1375 977.724Mi/s
learn_more/8MiB 8160467 ns 8143248 ns 685 980.336Mi/s
learn_more/16MiB 16449846 ns 16403959 ns 341 972.653Mi/s
learn_more/32MiB 33114440 ns 33099112 ns 169 966.346Mi/s
learn_more/64MiB 66329900 ns 65848214 ns 84 964.874Mi/s
learn_more/128MiB 133070590 ns 132812500 ns 42 961.895Mi/s
learn_more/256MiB 265652538 ns 264880952 ns 21 963.665Mi/s
Using the appropriate configuration, libhat is able to maintain its high throughput when searching
machine code at a speed comparable to searching uniform buffers. The table below once again compares
the single threaded throughput in bytes/s (real time) against the same two alternative pattern scanners.
The buffer being scanned is chrome.dll from a Chromium
snapshot
(~227MiB of code), and the pattern matches at the first instruction of DllMain.
The full source code is available here.
---------------------------------------------------------------------------------------------------
Benchmark Time CPU Iterations bytes_per_second
---------------------------------------------------------------------------------------------------
BM_find 36383054 ns 36254085 ns 153 6.09333Gi/s
BM_find_align 11898416 ns 11876327 ns 471 18.6322Gi/s
BM_find_hint 9687339 ns 9644397 ns 580 22.8849Gi/s
BM_find_align_hint 9689014 ns 9636324 ns 574 22.8810Gi/s
BM_UC1 236431167 ns 236979167 ns 24 960.172Mi/s
BM_UC2 427872923 ns 427884615 ns 13 530.565Mi/s
Below is a summary of the current support for libhat's platform-dependent APIs:
| Windows | Linux | macOS | Android | |
|---|---|---|---|---|
hat::get_system |
✅ | ✅ | ✅ | ✅ |
hat::memory_protector |
✅ | ✅ | ✅ | ✅ |
hp::get_process_module |
✅ | ✅ | ✅ | ✅ |
hp::get_module |
✅ | ✅ | ✅ | ✅ |
hp::module_at |
✅ | ✅ | ✅ | ✅ |
hp::is_readable |
✅ | ✅ | ✅ | ✅ |
hp::is_writable |
✅ | ✅ | ✅ | ✅ |
hp::is_executable |
✅ | ✅ | ✅ | ✅ |
hp::module::get_symbol |
✅ | ✅ | ✅ | ✅ |
hp::module::get_module_data |
✅ | ✅ | ✅ | ✅ |
hp::module::get_executable_data |
✅ | ✅ | ✅ | ✅ |
hp::module::get_section_data |
✅ | ✅ | ✅ | ✅ |
hp::module::for_each_section |
✅ | ✅ | ✅ | ✅ |
hp::module::for_each_segment |
✅ | ✅ | ✅ | ✅ |
libhat's signature syntax consists of space-delimited tokens and is backwards compatible with IDA syntax:
- 8 character sequences are interpreted as binary
- 2 character sequences are interpreted as hex
- 1 character must be a wildcard (
?)
Any digit can be substituted for a wildcard, for example:
????1111is a binary sequence, and matches any byte with all ones in the lower nibbleA?is a hex sequence, and matches any byte of the form1010????- Both
????????and??are equivalent to?, and will match any byte
A complete pattern might look like AB ? 12 ?3. This matches any 4-byte
subrange s for which all the following conditions are met:
s[0] == 0xABs[2] == 0x12s[3] & 0x0F == 0x03
As a scanning optimization, all patterns are recommended to have at least one fully masked byte. Attempting to find a pattern that does not meet this requirement results in significant performance degredation. Additionally, it is recommended (but not required) that patterns contain at least 2 consecutive fully masked bytes, as this will greatly speed up the vectorized scanning algorithms.
?1 02is allowed?? 02is allowed01 02is allowed (and recommended)
In code, there are a few ways to initialize a signature from its string representation:
#include <libhat/scanner.hpp>
// Parse a pattern's string representation to an array of bytes at compile time
constexpr hat::fixed_signature pattern = hat::compile_signature<"48 8D 05 ? ? ? ? E8">();
// Parse using the UDLs at compile time
using namespace hat::literals;
constexpr hat::fixed_signature pattern = "48 8D 05 ? ? ? ? E8"_sig; // stack owned
constexpr hat::signature_view pattern = "48 8D 05 ? ? ? ? E8"_sigv; // static lifetime (requires C++23)
// Parse it at runtime
using parsed_t = hat::result<hat::signature, hat::signature_parse_error>;
parsed_t runtime_pattern = hat::parse_signature("48 8D 05 ? ? ? ? E8");#include <libhat/scanner.hpp>
// Scan for this pattern using your CPU's vectorization features
auto begin = /* a contiguous iterator over std::byte */;
auto end = /* ... */;
hat::scan_result result = hat::find_pattern(begin, end, pattern);
// Scan a section in the process's base module
hat::scan_result result = hat::find_pattern(pattern, ".text");
// Or another module loaded into the process
std::optional<hat::process::module> ntdll = hat::process::get_module("ntdll.dll");
assert(ntdll.has_value());
hat::scan_result result = hat::find_pattern(pattern, ".text", *ntdll);
// Get the address pointed at by the pattern
const std::byte* address = result.get();
// Resolve an RIP relative address at a given offset
//
// | signature matches here
// | | relative address located at +3
// v v
// 48 8D 05 BE 53 23 01 lea rax, [rip+0x12353be]
//
const std::byte* relative_address = result.rel(3);libhat has a few optimizations for searching for patterns in x86_64 and AArch64 machine code:
#include <libhat/scanner.hpp>
// Compilers will often align the start address of a function on 16-bytes. Scanning for patterns that
// match the start of a function can take advantage of this by specifying the defaulted `alignment`
// parameter (all overloads have this parameter):
std::span<std::byte> range = /* ... */;
hat::signature_view pattern = /* ... */;
hat::scan_result result = hat::find_pattern(range, pattern, hat::scan_alignment::X16);
// Or, if the architecture has byte-aligned instructions (such as ARM and AArch64):
hat::scan_result result = hat::find_pattern(range, pattern, hat::scan_alignment::X4);
// Additionally, machine code contains a non-uniform distribution of bytes. By passing the respective
// scan hint (either `x86_64` or `aarch64`), the search anchor can be tuned to the least frequent
// bytes that are present in the pattern.
hat::scan_result result = hat::find_pattern(range, pattern, hat::scan_alignment::X1, hat::scan_hint::x86_64);#include <libhat/access.hpp>
// An example struct and it's member offsets
struct S {
uint32_t a{}; // 0x0
uint32_t b{}; // 0x4
uint32_t c{}; // 0x8
uint32_t d{}; // 0xC
};
S s;
// Obtain a mutable reference to 's.b' via it's offset
uint32_t& b = hat::member_at<uint32_t>(&s, 0x4);
// If the provided pointer is const, the returned reference is const
const uint32_t& b = hat::member_at<uint32_t>(&std::as_const(s), 0x4);#include <libhat/memory_protector.hpp>
uintptr_t* vftable = ...; // Pointer to a virtual function table in read-only data
size_t target_func_index = ...; // Index to an interesting function
// Use memory_protector to enable write protections
hat::memory_protector prot{
(uintptr_t) &vftable[target_func_index], // a pointer to the target memory
sizeof(uintptr_t), // the size of the memory block
hat::protection::Read | hat::protection::Write // the new protection flags
};
// Overwrite function table entry to redirect to a custom callback
vftable[target_func_index] = (uintptr_t) my_callback;
// On scope exit, original protections will be restored
prot.~memory_protector(); // compiler generated