Hi, thanks for great work!
DSpark uses the same target feature layers as the comparison drafters, but does not report a placement ablation. For Qwen3.5-style GDN/full-attention backbones, should features be selected uniformly by depth, by layer type, or at group boundaries? Have you tested whether Markov/confidence training changes the optimal selection?
Thanks again!
Hi, thanks for great work!
DSpark uses the same target feature layers as the comparison drafters, but does not report a placement ablation. For Qwen3.5-style GDN/full-attention backbones, should features be selected uniformly by depth, by layer type, or at group boundaries? Have you tested whether Markov/confidence training changes the optimal selection?
Thanks again!