Executing Linear Temporal Logic (LTL) instructions with learned robot policies faces two practical challenges: demonstrations may cover
individual behaviors without containing the temporal compositions requested at deployment, and semantic success predicates often provide
little useful motor guidance through quantitative robustness. We address these challenges by separating behavior learning from temporal
composition. From the same offline demonstrations, we train an unconditioned multimodal diffusion policy and a semantic predictor that
estimates which semantic outcomes are likely to follow a proposed action sequence. These predictions provide a learned guidance signal
without requiring hand-designed robustness measures for semantic task objectives. At deployment, an LTL automaton tracks instruction
progress and scores the policy's proposed actions according to their predicted semantic outcomes. Neither learned component receives
the specification during training, allowing demonstrated behaviors to be reused in longer, previously unseen temporal compositions without
retraining. For additional runtime safety constraints with informative continuous margins, an optional robustness-based controller locally
refines the selected actions. Experiments in a navigation environment, CALVIN manipulation, and on a real robot demonstrate reliable
execution of complex temporal instructions and additional safety constraints, with substantial improvements over temporal-logic baselines.