Domain-Specificity Inducing Transformers for Source-Free Domain Adaptation
Conventional Domain Adaptation (DA) methods aim to learn ___domain-invariant feature representations to improve the target adaptation performance. However, we motivate that ___domain-specificity is equally important since in-___domain trained models hold crucial ___domain-specific properties that are beneficial for adaptation. Hence, we propose to build a framework that supports disentanglement and learning of ___domain-specific factors and task-specific factors in a unified model. Motivated by the success of vision transformers in several multi-modal vision problems, we find that queries could be leveraged to extract the ___domain-specific factors. Hence, we propose a novel Domain-specificity-inducing Transformer (DSiT) framework for disentangling and learning both ___domain-specific and task-specific factors. To achieve disentanglement, we propose to construct novel Domain-Representative Inputs (DRI) with ___domain-specific information to train a ___domain classifier with a novel ___domain token. We are the first to utilize vision transformers for ___domain adaptation in a privacy-oriented source-free setting, and our approach achieves state-of-the-art performance on single-source, multi-source, and multi-target benchmarks
