Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
36 commits
Select commit Hold shift + click to select a range
906f354
feat(insert-sync): add three absolute synchronization gates
Adamkh329 Aug 3, 2026
633d97c
fix(insert-sync): decode every legal buffer-token pipe spelling in G1
Adamkh329 Aug 5, 2026
aabe641
test(insert-sync): cover both buffer-token pipe spellings in G1
Adamkh329 Aug 5, 2026
c566b1d
fix(insert-sync): require an event pair to be equally guarded in G2
Adamkh329 Aug 5, 2026
5593193
test(insert-sync): cover event-pair guard placement in G2
Adamkh329 Aug 5, 2026
1ffec3f
fix(insert-sync): excuse a same-pipe PIPE_S dependency in G1
Adamkh329 Aug 5, 2026
8e5462e
fix(insert-sync): credit a coverage edge only where the sync is alway…
Adamkh329 Aug 5, 2026
97dc595
test(insert-sync): cover sync reachability in G1
Adamkh329 Aug 5, 2026
77fc28f
test(insert-sync): mark the section of the gate fixtures' sync ops
Adamkh329 Aug 5, 2026
57e3c7b
feat(insert-sync): add two differential synchronization gates
Adamkh329 Aug 3, 2026
3983d6e
feat(insert-sync): allocate event ids, buffer ids and barriers from o…
Adamkh329 Aug 4, 2026
93dc003
test(insert-sync): cover the unified allocator's paths and their nega…
Adamkh329 Aug 4, 2026
e80750a
docs(insert-sync): describe the unified sync allocator
Adamkh329 Aug 4, 2026
8405b4a
fix(insert-sync): close five write-only or unused members in the allo…
Adamkh329 Aug 4, 2026
16c0410
docs(insert-sync): separate what is proven from what is measured
Adamkh329 Aug 4, 2026
8c3011d
fix(insert-sync): assert the barrier-unbounded invariant where the sp…
Adamkh329 Aug 6, 2026
5ff666b
docs(insert-sync): correct what SyncOpRecord::mechanism is
Adamkh329 Aug 6, 2026
24fd079
fix(insert-sync): read the leak scan's open interval by field, not pa…
Adamkh329 Aug 6, 2026
139530f
test(insert-sync): record the new self-ordered field in the allocator…
Adamkh329 Aug 6, 2026
b121265
docs(insert-sync): correct false claims in the unified allocator comm…
Adamkh329 Aug 6, 2026
9582e55
docs(insert-sync): compress the unified allocator comment blocks
Adamkh329 Aug 6, 2026
1ba8a51
docs(insert-sync): correct seven claims contradicted by the code
Adamkh329 Aug 6, 2026
f8515c3
docs(insert-sync): put the G1-self section banner on the G1-self section
Adamkh329 Aug 6, 2026
6307327
fix(insert-sync): bracket every routed hazard, and refuse the ones a …
Adamkh329 Aug 7, 2026
d33e584
feat(insert-sync): record on each hazard whether its endpoints share …
Adamkh329 Aug 10, 2026
5e31dbe
feat(ptoas): add --check-addr-reuse-war to list the orderings an addr…
Adamkh329 Aug 10, 2026
2c12592
fix(ptoas): reject pto.tassign under --enable-unified-sync
Adamkh329 Aug 11, 2026
a41229e
fix(insert-sync): emit the unrealised-hazard diagnostic instead of as…
Adamkh329 Aug 11, 2026
f4ce8ff
fix(insert-sync): give persisted coverage profiles a stable digest an…
Adamkh329 Aug 11, 2026
9e06a9d
fix(insert-sync): fail the pass on the buffer-id conditions it only r…
Adamkh329 Aug 11, 2026
150ac57
fix(insert-sync): refuse a routing plan containing a split hazard
Adamkh329 Aug 11, 2026
becb096
docs(insert-sync): correct the record on BaseMemInfo::hasKnownPhysica…
Adamkh329 Aug 11, 2026
de48e6d
test(insert-sync): pin that both ends of a routed hazard hold the sam…
Adamkh329 Aug 12, 2026
3610b6f
chore: ignore every out-of-tree build directory, not just build/
Adamkh329 Aug 12, 2026
4a6a295
fix(bufid_sync): make optimizeSamePipeMerge deterministic
Adamkh329 Aug 16, 2026
f155840
feat(ptoas): route single-core-dispatch kernels to whole-kernel buffe…
Adamkh329 Aug 19, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 4 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,10 @@
build/
build_plain/
build_plan/
# Any out-of-tree build directory, however named. `build/` alone let a second
# configuration (build-assert/) be added by `git add -A`.
build*/
build-assert/
install/

# TileLang ST standalone build outputs (see temp_docs/standalone_st.md)
Expand Down
525 changes: 525 additions & 0 deletions docs/designs/ptoas-unified-sync-design.md

Large diffs are not rendered by default.

7 changes: 7 additions & 0 deletions include/PTO/Transforms/InsertSync/SyncCodegen.h
Original file line number Diff line number Diff line change
Expand Up @@ -38,8 +38,15 @@ class SyncCodegen {

/// 入口函数:执行代码生成
void Run();

/// True if a set/wait hazard reached emission with NO event id, i.e. with no
/// mechanism realised. `Run()` emits a diagnostic for each such hazard; the caller
/// decides whether that fails the pass. See the guard in `SyncInsert`.
bool sawUnrealisedHazard() const { return sawUnrealisedHazard_; }

private:
bool sawUnrealisedHazard_ = false;

// --- 核心插入逻辑 ---
void SyncInsert(IRRewriter &rewriter, Operation *op, SyncOperation *sync,
bool beforeInsert);
Expand Down
77 changes: 72 additions & 5 deletions include/PTO/Transforms/InsertSync/SyncCommon.h
Original file line number Diff line number Diff line change
Expand Up @@ -163,8 +163,48 @@ class SyncOperation {
SYNC_BLOCK_WAIT,
SYNC_BLOCK_ALL,
};


/// Which hardware resource backs this sync -- and therefore, per the buf-id
/// design revision, WHERE it may be placed.
///
/// This is deliberately a DIFFERENT axis from TYPE. `TYPE` is the op *shape*
/// (a set, a wait, a barrier) and is decided by the analysis. `MECHANISM` is
/// the resource *class* that backs it, and is decided by the allocator:
///
/// EVENT Per-(src,dst) private pool of 8 flags. Carried at a SyncIR
/// position, so `MoveSyncState` may hoist it out of a branch or
/// loop. This is what every in-tree sync is today.
/// BUFID Shared rotating pool of K buffer-ID tokens (K=0 on A3, 32 on A5).
/// A token's get/rel counters are BIDIRECTIONAL: one token discharges
/// the forward RAW *and* the reverse WAR. Two consequences, and they
/// are the reason this enum exists rather than a bool:
/// - no backward-matched pair is needed (`UpdateBackwardMatchSync`
/// lives in SyncEventIdAllocation, which the unified allocator
/// never runs -- so there is nothing to delete);
/// - it MUST be emitted OP-ANCHORED, bracketing the access, not at
/// a hoisted SyncIR position.
/// Set by the allocator when it routes a hazard off events.
/// BARRIER Spill. Carries no id.
/// BLOCK Cross-core block sync (reserved ids 14/15). Not routed by this
/// allocator; present so a block sync can never be mistaken for an
/// event.
///
/// Every value EXCEPT BUFID is derivable from TYPE, and the constructor does
/// derive it (see `DefaultMechanismFor`), so no construction site has to
/// remember to set it and a barrier can never silently claim to be an event.
/// BUFID is the one value that is *not* derivable -- buffer-ID sync does not
/// go through `SyncOperation` at all today -- and it is precisely the routing
/// decision the unified allocator exists to make.
enum class MECHANISM { EVENT, BUFID, BARRIER, BLOCK };

bool isCompensation = false;
/// For a compensation op (head-set / tail-wait), the `kSyncIndex_` of the in-body
/// hazard it primes and drains; -1 otherwise. Needed because a compensation pair is
/// created as its OWN sync group, so `GetSyncIndex()` does not identify its owner --
/// and G2 must be able to tell "these two records are halves of one hazard" from
/// "two different hazards collide on an id". Without it the exemption would have to
/// excuse every compensation record, which would drop real conflicts.
int compensationOf = -1;

static const int kNullEventId{-1};

Expand All @@ -173,7 +213,20 @@ class SyncOperation {
// Root buffers participating in the dependency that created this sync pair.
// These are kept for allocation/widening heuristics and debug printing; set/
// wait redundancy pruning is based on the pipe pair semantics.
//
// `rootBuffer` is the ROOT ALLOCATION, which is much COARSER than a buffer: many
// distinct tiles share one root. It is therefore NOT a buffer identity, and must
// not be used as one -- see `depMemInfos` below.
SmallVector<Value> depRootBuffers;
// The BaseMemInfo objects behind this dependency, kept so a consumer can
// recover the TILE identity `(scope, baseAddr, size)` rather than the coarse
// root above. Owned by the translator's `Buffer2MemInfoMap`; these are raw
// pointers into that arena, which every analysis in the pipeline shares.
//
// This exists because `InsertSyncAnalysis` HAS these pointers when it builds a
// sync pair and used to discard them, keeping only the roots -- which left the
// unified allocator unable to tell which buffer a hazard was actually on.
SmallVector<const BaseMemInfo *> depMemInfos;
bool uselessSync{false};
int eventIdNum{1};
// For multi-buffer dyn-event sync: the slot SSA expression at this access
Expand All @@ -192,15 +245,25 @@ class SyncOperation {
SyncOperation(TYPE type, pto::PipelineType srcPipe, pto::PipelineType dstPipe,
unsigned kSyncIndex, unsigned syncIRIndex,
std::optional<int> forEndIndex, bool isComp = false)
: isCompensation(isComp), eventIds({}), type_(type), srcPipe_(srcPipe),
: isCompensation(isComp), eventIds({}), type_(type),
mechanism_(DefaultMechanismFor(type)), srcPipe_(srcPipe),
dstPipe_(dstPipe), kSyncIndex_(kSyncIndex), syncIRIndex_(syncIRIndex),
forEndIndex_(forEndIndex) {};

~SyncOperation() = default;

std::unique_ptr<SyncOperation> GetMatchSync(unsigned index) const;

TYPE GetType() const { return type_; }

/// The resource class backing this sync. Defaults from TYPE; only the
/// allocator moves a sync off its default (today: onto BUFID).
MECHANISM GetMechanism() const { return mechanism_; }
void SetMechanism(MECHANISM m) { mechanism_ = m; }

/// The mechanism a given TYPE implies. Kept next to the enum so the two can
/// never drift; `SetPipeAll` and the constructor are its only callers.
static MECHANISM DefaultMechanismFor(TYPE t);
pto::PipelineType GetSrcPipe() const { return srcPipe_; }
pto::PipelineType GetDstPipe() const { return dstPipe_; }

Expand All @@ -223,6 +286,7 @@ class SyncOperation {
bool isBarrierType() const;

static std::string TypeName(TYPE t);
static std::string MechanismName(MECHANISM m);
std::string GetCoreTypeName(TCoreType t) const;

using SyncOperations =
Expand All @@ -235,6 +299,9 @@ class SyncOperation {

private:
TYPE type_;
// Private on purpose: it must stay consistent with type_ (see SetPipeAll,
// which downgrades both together). Write it through SetMechanism only.
MECHANISM mechanism_;
pto::PipelineType srcPipe_;
pto::PipelineType dstPipe_;
const unsigned kSyncIndex_;
Expand Down
7 changes: 7 additions & 0 deletions include/PTO/Transforms/InsertSync/SyncEventIdAllocation.h
Original file line number Diff line number Diff line change
Expand Up @@ -40,6 +40,12 @@ using SyncCycle = DenseMap<int, EventCyclePool>;

class SyncEventIdAllocation {
public:
/// Number of event ids reserved at the tail of the intra-core pool
/// [0, kTotalEventIdNum) for the direction `src -> dst`; 0 when the direction
/// reserves nothing. Exposed so the device-id-legality gate reads the same
/// reservation map this allocator obeys, rather than duplicating it.
static uint64_t GetReservedEventIdNum(PipelineType src, PipelineType dst);

SyncEventIdAllocation(SyncIRs &syncIR, SyncOperations &syncOperations)
: syncIR_(syncIR), syncOperations_(syncOperations) {
reserveBlockAllEventIds();
Expand Down Expand Up @@ -124,6 +130,7 @@ class SyncEventIdAllocation {
static const llvm::DenseMap<std::pair<PipelineType, PipelineType>, uint64_t>
reservedEventIdNum;
uint64_t reservedBlockSyncEventIdNum{0};

};

} // namespace pto
Expand Down
136 changes: 136 additions & 0 deletions include/PTO/Transforms/InsertSync/SyncOracleExtract.h
Original file line number Diff line number Diff line change
@@ -0,0 +1,136 @@
// Copyright (c) 2026 Huawei Technologies Co., Ltd.
// This program is free software, you can redistribute it and/or modify it under the terms and conditions of
// CANN Open Software License Agreement Version 2.0 (the "License").
// Please refer to the License for details. You may not use this file except in compliance with the License.
// THIS SOFTWARE IS PROVIDED ON AN "AS IS" BASIS, WITHOUT WARRANTIES OF ANY KIND, EITHER EXPRESS OR IMPLIED,
// INCLUDING BUT NOT LIMITED TO NON-INFRINGEMENT, MERCHANTABILITY, OR FITNESS FOR A PARTICULAR PURPOSE.
// See LICENSE in the root of the software repository for the full text of the License.

//===- SyncOracleExtract.h - Shared sync-op extractor (oracle) ---*- C++ -*-===//
//
// The shared walk-surface the oracle gates ride on: the emitted `pto.set_flag` /
// `pto.wait_flag` / `pto.get_buf` / `pto.rls_buf` / `pto.barrier` ops of one
// function. G3 reads the ids, G2 reads their live ranges.
//
// `SyncOpRecord` describes the same synchronization before codegen, as an allocator
// holds it. Only G3 has an overload taking it, and only the gate self-test feeds it,
// because nothing else in this build produces pre-codegen records.
//
// Extraction is deterministic (program order), so a rendering is stable across runs.
//===----------------------------------------------------------------------===//

#ifndef MLIR_DIALECT_PTO_TRANSFORMS_INSERTSYNC_SYNCORACLEEXTRACT_H
#define MLIR_DIALECT_PTO_TRANSFORMS_INSERTSYNC_SYNCORACLEEXTRACT_H

#include "PTO/Transforms/InsertSync/SyncCommon.h"
#include "mlir/Dialect/Func/IR/FuncOps.h"
#include "llvm/ADT/ArrayRef.h"
#include "llvm/ADT/STLFunctionalExtras.h"
#include "llvm/ADT/SmallVector.h"
#include "llvm/ADT/StringRef.h"
#include "llvm/Support/raw_ostream.h"

#include <string>

namespace mlir {
namespace pto {
namespace oracle {

/// One emitted synchronization op, read out of the IR after codegen.
/// `eventId` is set only for set_flag/wait_flag; `bufId` only for
/// get_buf/rls_buf; both are -1 when not applicable.
struct IRSyncRecord {
std::string opName;
std::string srcPipe;
std::string dstPipe;
/// Typed pipes for the gates. The strings above are for rendering only;
/// gates must never string-match a pipe name.
PipelineType srcPipeType = PipelineType::PIPE_UNASSIGNED;
PipelineType dstPipeType = PipelineType::PIPE_UNASSIGNED;
int64_t eventId = -1;
int64_t bufId = -1;
/// get_buf/rls_buf `mode`. 0 = the bidirectional token (the only mode PTOAS
/// emits today). Non-zero = the directional mode (set_flagV2/wait_flagV2),
/// deferred -- see the buffer-ID legality notes in SyncOracleGates.h, which
/// applies the strict pairing rules only to mode 0.
int64_t mode = 0;
/// Enclosing block. A get_buf and its rls_buf must sit in the SAME block, or
/// the pair may not both execute (counter desync). Null for non-buf ops.
mlir::Block *block = nullptr;
/// The op itself, for gates that must ask where it sits in the control flow. An
/// event pair is checked on its enclosing conditionals, which a block pointer
/// alone cannot answer: two blocks differ for loop nesting too, and that case is
/// legitimate.
mlir::Operation *op = nullptr;
unsigned order = 0; // program order within the function
};

/// One entry of the in-memory SyncOperation set, before codegen. `eventIds` is
/// empty until an allocator assigns them.
struct SyncOpRecord {
std::string type;
/// `static_cast<int>(SyncOperation::TYPE)`; -1 when unset. Gates switch on
/// this, never on the `type` string.
int typeCode = -1;
/// `SyncOperation::MechanismName` of the resource class the allocator routed this
/// sync to -- and so whether it is SyncIR-positioned (event/barrier) or op-anchored
/// (bufid). Empty until the pre-codegen extractor fills it.
///
/// RENDERING ONLY. Unlike `type`, this has no numeric companion, so nothing can
/// switch on it without string-matching a name. A gate that needs the mechanism as
/// a decision should get a `mechanismCode` beside this, the way `typeCode` sits
/// beside `type`, rather than compare these strings.
std::string mechanism;
int srcPipe = -1;
int dstPipe = -1;
llvm::SmallVector<int, 2> eventIds;
unsigned syncIndex = 0;
unsigned irIndex = 0;
/// Mirrors `SyncOperation::isCompensation` / `compensationOf`. G2's syncops scan
/// needs both: a compensation record shares its ids with the in-body pair it primes
/// BY DESIGN, so that pairing must not read as two hazards colliding -- while a
/// collision with any OTHER hazard still must.
bool isCompensation = false;
int compensationOf = -1;
};

/// Human-readable name of an InsertSync PipelineType ("MTE2", "V", "ALL", ...).
llvm::StringRef pipelineTypeName(PipelineType pipe);

/// The pipe a `get_buf`/`rls_buf` `op_type` names.
///
/// `PTO_PipeLikeAttr` admits three spellings -- pipe_event_type, sync_op_type, and a
/// plain PipeAttr. A decoder handling only the first two returns PIPE_UNASSIGNED for
/// the third rather than failing, so its caller silently loses the op.
pto::PIPE bufSyncPipe(mlir::Attribute opTypeAttr);

/// Render a report and write it to stderr atomically.
///
/// STDERR, NEVER STDOUT: stdout carries the emitted kernel source, so a report
/// written there is spliced into the C++ and it stops compiling.
///
/// MLIR runs nested func::FuncOp passes in PARALLEL and ptoas does not disable
/// threading, so unsynchronized writes from several functions interleave mid-line.
/// Every oracle pass must print through here. Ordering *between* functions stays
/// nondeterministic, so no consumer may depend on it.
void emitReport(llvm::function_ref<void(llvm::raw_ostream &)> body);

/// Extract the emitted sync ops from `func`, in program order.
llvm::SmallVector<IRSyncRecord> extractFromIR(func::FuncOp func);

/// The same synchronization before codegen, as an allocator holds it: the ids it has
/// just written, before anything is emitted.
llvm::SmallVector<SyncOpRecord> extractFromSyncOps(const SyncOperations &ops);


/// Deterministic, diff-friendly rendering (one record per line).
void printIRRecords(llvm::raw_ostream &os, llvm::StringRef funcName,
llvm::ArrayRef<IRSyncRecord> records);
void printSyncOpRecords(llvm::raw_ostream &os, llvm::StringRef funcName,
llvm::ArrayRef<SyncOpRecord> records);

} // namespace oracle
} // namespace pto
} // namespace mlir

#endif // MLIR_DIALECT_PTO_TRANSFORMS_INSERTSYNC_SYNCORACLEEXTRACT_H
Loading