Machine model for scheduling, bundling, and heuristics.
Declared in <llvm/MC/MCSchedule.h>
struct MCSchedModel;
The machine model directly provides basic information about the microarchitecture to the scheduler in the form of properties. It also optionally refers to scheduler resource tables and itinerary tables. Scheduler resource tables model the latency and cost for each instruction type. Itinerary tables are an independent mechanism that provides a detailed reservation table describing each cycle of instruction execution. Subtargets may define any or all of the above categories of data depending on the type of CPU and selected scheduler.
The machine independent properties defined here are used by the scheduler as an abstract machine model. A real micro-architecture has a number of buffers, queues, and stages. Declaring that a given machine-independent abstract property corresponds to a specific physical property across all subtargets can't be done. Nonetheless, the abstract model is useful. Futhermore, subtargets typically extend this model with processor specific resources to model any hardware features that can be exploited by scheduling heuristics and aren't sufficiently represented in the abstract.
The abstract pipeline is built around the notion of an "issue point". This is merely a reference point for counting machine cycles. The physical machine will have pipeline stages that delay execution. The scheduler does not model those delays because they are irrelevant as long as they are consistent. Inaccuracies arise when instructions have different execution delays relative to each other, in addition to their intrinsic latency. Those special cases can be handled by TableGen constructs such as, ReadAdvance, which reduces latency when reading data, and ReleaseAtCycles, which consumes a processor resource when writing data for a number of abstract cycles.
TODO: One tool currently missing is the ability to add a delay to ReleaseAtCycles. That would be easy to add and would likely cover all cases currently handled by the legacy itinerary tables.
A note on out-of-order execution and, more generally, instruction buffers. Part of the CPU pipeline is always in-order. The issue point, which is the point of reference for counting cycles, only makes sense as an in-order part of the pipeline. Other parts of the pipeline are sometimes falling behind and sometimes catching up. It's only interesting to model those other, decoupled parts of the pipeline if they may be predictably resource constrained in a way that the scheduler can exploit.
The LLVM machine model distinguishes between in-order constraints and out-of-order constraints so that the target's scheduling strategy can apply appropriate heuristics. For a well-balanced CPU pipeline, out-of-order resources would not typically be treated as a hard scheduling constraint. For example, in the GenericScheduler, a delay caused by limited out-of-order resources is not directly reflected in the number of cycles that the scheduler sees between issuing an instruction and its dependent instructions. In other words, out-of-order resources don't directly increase the latency between pairs of instructions. However, they can still be used to detect potential bottlenecks across a sequence of instructions and bias the scheduling heuristics appropriately.
| Name | Description |
|---|---|
computeInstrLatency | computeInstrLatency overloads |
getExtraProcessorInfo | Return the extra processor info for this model. |
getNumProcResourceKinds | Return the number of processor resource kinds in this model. |
getProcResource | Return the processor resource descriptor at ProcResourceIdx. |
getProcessorID | Return the TableGen processor ID for this model. |
getReciprocalThroughput | Return the reciprocal throughput for instruction Inst on STI. |
getResourceBufferSize | Return the buffer size of the resource. If a positive scale factor is provided and the original buffer size is > 1, the size is scaled accordingly. |
getSchedClassDesc | Return the scheduling class descriptor at SchedClassIdx. |
getSchedClassName | Return the name of the scheduling class at SchedClassIdx. |
hasExtraProcessorInfo | Return true if extra processor info is available for this model. |
hasInstrSchedModel | Does this machine model include instruction-level scheduling. |
isComplete | Return true if this machine model data for all instructions with a scheduling class (itinerary class or SchedRW list). |
isOutOfOrder | Return true if machine supports out of order execution. |
| Name | Description |
|---|---|
computeInstrLatency | Returns the latency value for the scheduling class. |
getBypassDelayCycles | Returns the bypass delay cycle for the maximum latency write cycle. |
getForwardingDelayCycles | Returns the maximum forwarding delay for register reads dependent on writes of scheduling class WriteResourceIdx. |
getReciprocalThroughput | getReciprocalThroughput overloads |
| Name | Description |
|---|---|
CompleteModel | True if this model has scheduling data for all instructions with a class. |
EnableIntervals | Whether MachineScheduler should track resource usage with intervals. |
ExtraProcessorInfo | Optional extra processor details for tools such as llvm-mca. |
HighLatency | Expected latency of "very high latency" operations. |
InstrItineraries | Instruction itinerary tables used by InstrItineraryData. |
IssueWidth | Maximum number of instructions that may be scheduled in one cycle group. |
LoadLatency | Expected latency of load instructions. |
LoopMicroOpBufferSize | Number of micro-ops the processor may buffer for optimized loop execution. |
MicroOpBufferSize | Number of micro-ops the processor may buffer for out-of-order execution. |
MispredictPenalty | Typical extra cycles to recover from a branch misprediction. |
NumProcResourceKinds | Number of processor resource kinds in ProcResourceTable. |
NumSchedClasses | Number of scheduling classes in SchedClassTable. |
PostRAScheduler | True if post-RA scheduling should be enabled (default false). |
ProcID | TableGen processor ID for this scheduling model. |
ProcResourceTable | Table of processor resource descriptors. |
SchedClassNames | String table of scheduling class names (debug/dump builds). |
SchedClassTable | Table of scheduling class descriptors. |
| Name | Description |
|---|---|
Default | Returns the default initialized model. |
DefaultHighLatency | Default value for HighLatency. |
DefaultIssueWidth | Default value for IssueWidth. |
DefaultLoadLatency | Default value for LoadLatency. |
DefaultLoopMicroOpBufferSize | Default value for LoopMicroOpBufferSize. |
DefaultMicroOpBufferSize | Default value for MicroOpBufferSize. |
DefaultMispredictPenalty | Default value for MispredictPenalty. |
| Name | Description |
|---|---|
llvm::InstrItineraryData | Itinerary data supplied by a subtarget to be used by a target. |
| Name | Description |
|---|---|
mca::computeBlockRThroughput | Compute the reciprocal block throughput for a code block. |
mca::computeProcResourceMasks | Populates vector Masks with processor resource masks. |
mca::dumpProcResourceMasks | Dump processor resource masks from SM for debugging. |