Know what the words mean.
Short definitions, with a path to the full mechanism. Search in English or Arabic.
ABI
Rules for calling functions, preserving registers, and connecting compiled code.
ALU
An execution unit for integer arithmetic and bit operations.
Arithmetic intensity
Useful operations per byte moved at a chosen memory boundary.
Atomic
An access with defined indivisibility and concurrency semantics.
Bandwidth
The amount of data transferred per unit time.
BAR
A PCI device configuration resource describing a memory or I/O window.
Binning
Grouping manufactured parts by tested operating characteristics.
Branch predictor
Hardware that guesses future control flow before it is resolved.
BVH
Nested bounds that narrow a ray’s search for intersections.
Cache line
A block transferred and tracked as a unit by a cache.
Chiplet
A die designed to participate in a multi-die system.
CMOS
Logic built from complementary n-channel and p-channel devices.
Coalescing
Combining nearby GPU lane accesses into fewer memory transactions.
Coherence
Rules that coordinate cached copies of a memory location.
CPI / IPC
Average cycles per instruction and instructions per cycle for a defined measurement.
Die
One manufactured piece of processed silicon.
DMA
A device transfers data through a mapped memory path without CPU byte-by-byte copying.
DRAM
Charge-based memory that needs refresh to retain data.
DVFS
Changing operating voltage and frequency under platform policy.
ECC
Redundant information that detects or corrects specified error patterns.
False sharing
Different writable data shares a coherence line and causes ownership contention.
FLOPS
Floating-point operations per second, for a specified format and counting convention.
FPGA
Programmable logic and routing configured to implement a circuit.
FTL
An SSD mapping from logical storage addresses to physical flash pages.
IOMMU
Translates and constrains device-visible memory addresses.
ISA
The software-visible instruction and architectural state contract.
JIT
Generating machine code while a program is running.
Latency
Delay between a request and its usable result.
LUT
A configurable truth-table resource in programmable logic.
MAC
Multiply operands and add the product into an accumulator.
MMIO
Device registers exposed through mapped CPU address ranges.
MT/s
Transfer rate, which is distinct from the memory clock frequency.
NUMA
Memory cost varies with processor and physical placement.
NVMe
A queue-based command interface for nonvolatile storage.
Occupancy
Resident GPU work relative to a supported residency limit.
Page fault
An address translation or access condition that requires OS handling.
PCIe
A packet-based serial interconnect with negotiated speed and lane width.
Register renaming
Mapping architectural register names onto physical value storage.
Retirement
Making completed instruction effects part of architectural state.
Roofline
An upper bound connecting compute rate, bandwidth, and arithmetic intensity.
SIMD / SIMT
Vector operations and a thread abstraction using grouped lane issue, respectively.
SM / CU
Vendor-specific structures combining scheduling, lanes, and local resources.
SMT
Multiple hardware contexts share resources in one physical core.
SoC
Integration of processors, controllers, and other system blocks.
SRAM
Feedback-based state storage that retains data while powered without DRAM-style refresh.
TLB
Cached address translations and associated permissions.
TOPS
An operation rate requiring stated precision, counting, and sparsity conditions.
Warp / wavefront
A group of GPU threads participating in grouped instruction issue.
Write amplification
Physical flash writes exceed host writes because of internal management.