Abstract: An apparatus includes a memory and one or more processors. The memory is configured to store tensors for Machine Learning (ML) processing. The one or more processors are configured to receive a work plan associated with a subgraph of a ML graph of a ML model, the work plan supports processing of tensors having respective shapes in a selected range of shapes. A shape of a tensor specifies respective sizes of dimensions of that tensor. The one or more processors are further configured to receive from the memory an input tensor having an actual shape, to modify the work plan based on the actual shape to produce a modified work plan for processing the input tensor in accordance with the subgraph, and to process the input tensor in accordance with the subgraph by submitting the modified work plan for execution by one or more of the processors.