Associate Editor Comments to the Author (required): Thank you for the revised version of the paper. You have addressed all the remaining requests of the reviewers. I believe that your paper can now be accepted. Reviewer(s)' Comments to Author: Reviewer: 1 Recommendation: Accept Comments: The revised version addressed all the concerns I raised in the previous version. Additional Questions: What is the key contribution of this paper?: This paper trains a recommender system to jointly suggest items and explain its reasoning via subjective rationales. It fine-tunes this model to incorporate iterative user feedback via self-supervised bot-play using product reviews for training. The major benefit is training the conversational recommendation model doesn’t require dialog transcripts. The work is evaluated based on simulated users and hired user subjects (Turks) and the results are very promising. Relevance to the journal (TORS Scope Statement): Excellent Comparison with previous and related works (Consider, for example: Are previous and related works adequately discussed? Are the relations of previous work to the current work clear?): Good a) Novelty (Comment on the novelty of the idea, and the methods and techniques proposed): Good Novelty and potential for innovation details: b) Soundness (Comment on the soundness of the idea, and the methods and techniques proposed.): Good soundness details: c) Experimental evaluation [Important guidelines]: (Do the authors use an appropriate methodology and suitable statistical methods? Is the choice of datasets, baseline methods and evaluation metrics in empirical comparisons well justified? Have all compared models in offline experiments been systematically tuned?): Good Experimental evaluation details: d) Theoretical background and evaluation: (Have the research questions and hypotheses been made explicit? Have all claims been substantiated through empirical evidence, theoretical considerations, or proofs?): Good Theoretical background and evaluation details: e) Reproducibility [Important guidelines] (In case of offline experiments, do the authors share the code and data to reproduce the findings? Does the shared material include scripts for preprocessing and tuning, are hyper-parameters documented? Does the shared code include the code for the baselines? In case of user studies, is the collected data shared?): Excellent Reproducibility details: Presentation: (Comment about the quality of English and the structure of the paper: Are there parts that should be rewritten, expanded upon, or dropped? Is the paper concise?): Good Reasons to accept (Provide 1-3 concise bulleted-list items.): Reasons to revise/reject (Provide 1-3 concise bulleted-list items.): Does this paper have potential real world significance?: Yes ==================== Associate Editor Comments to the Author (required): The reviewers have recognised the improvements of the newer version with respect to the original submission. There are just a couple of final observations that can be easily fixed in the final version. Please find below the comments from the reviewers of your paper: Reviewer: 1 Recommendation: Accept Comments: This version is much better and the authors have addressed all my major concerns. Minor suggestions: 1) Ɣ_u and Ɣ_u^{MF} are so different and important, thus better to have them separately in Figure 2(a) 2) missing \ in table 2 Additional Questions: What is the key contribution of this paper?: This paper trains a recommender system to jointly suggest items and explain its reasoning via subjective rationales. It fine-tunes this model to incorporate iterative user feedback via self-supervised bot-play using product reviews for training. The major benefit is training the conversational recommendation model doesn’t require dialog transcripts. The work is evaluated based on simulated users and hired user subjects (Turks) and the results are very promising. Relevance to the journal (TORS Scope Statement): Excellent Comparison with previous and related works (Consider, for example: Are previous and related works adequately discussed? Are the relations of previous work to the current work clear?): Good a) Novelty (Comment on the novelty of the idea, and the methods and techniques proposed): Good Novelty and potential for innovation details: This is an extended and improved version of a prior Recsys paper, thus not very novel. However I assume that's fine for the journal. b) Soundness (Comment on the soundness of the idea, and the methods and techniques proposed.): Good soundness details: c) Experimental evaluation [Important guidelines]: (Do the authors use an appropriate methodology and suitable statistical methods? Is the choice of datasets, baseline methods and evaluation metrics in empirical comparisons well justified? Have all compared models in offline experiments been systematically tuned?): Good Experimental evaluation details: d) Theoretical background and evaluation: (Have the research questions and hypotheses been made explicit? Have all claims been substantiated through empirical evidence, theoretical considerations, or proofs?): Good Theoretical background and evaluation details: e) Reproducibility [Important guidelines] (In case of offline experiments, do the authors share the code and data to reproduce the findings? Does the shared material include scripts for preprocessing and tuning, are hyper-parameters documented? Does the shared code include the code for the baselines? In case of user studies, is the collected data shared?): Fair Reproducibility details: Presentation: (Comment about the quality of English and the structure of the paper: Are there parts that should be rewritten, expanded upon, or dropped? Is the paper concise?): Good Reasons to accept (Provide 1-3 concise bulleted-list items.): Reasons to revise/reject (Provide 1-3 concise bulleted-list items.): Does this paper have potential real world significance?: Yes Reviewer: 2 Recommendation: Minor Revision Comments: The revised version of the paper shows significant improvement. I strongly recommend that the authors enhance the transparency and reproducibility of their work by providing access to the source code of their model. This step will contribute to the overall credibility of the research and facilitate the replication of the study for validation purposes. Additional Questions: What is the key contribution of this paper?: The paper presents a framework for training conversational recommender systems using bot-play on historical user reviews without the need for large collections of human dialogues. Relevance to the journal (TORS Scope Statement): Excellent Comparison with previous and related works (Consider, for example: Are previous and related works adequately discussed? Are the relations of previous work to the current work clear?): Good a) Novelty (Comment on the novelty of the idea, and the methods and techniques proposed): Good Novelty and potential for innovation details: b) Soundness (Comment on the soundness of the idea, and the methods and techniques proposed.): Good soundness details: c) Experimental evaluation [Important guidelines]: (Do the authors use an appropriate methodology and suitable statistical methods? Is the choice of datasets, baseline methods and evaluation metrics in empirical comparisons well justified? Have all compared models in offline experiments been systematically tuned?): Good Experimental evaluation details: d) Theoretical background and evaluation: (Have the research questions and hypotheses been made explicit? Have all claims been substantiated through empirical evidence, theoretical considerations, or proofs?): Good Theoretical background and evaluation details: e) Reproducibility [Important guidelines] (In case of offline experiments, do the authors share the code and data to reproduce the findings? Does the shared material include scripts for preprocessing and tuning, are hyper-parameters documented? Does the shared code include the code for the baselines? In case of user studies, is the collected data shared?): Poor Reproducibility details: Presentation: (Comment about the quality of English and the structure of the paper: Are there parts that should be rewritten, expanded upon, or dropped? Is the paper concise?): Good Reasons to accept (Provide 1-3 concise bulleted-list items.): Reasons to revise/reject (Provide 1-3 concise bulleted-list items.): Does this paper have potential real world significance?: Maybe ==================== Associate Editor comments: Associate Editor Comments to the Author (required): The reviewers agree that the paper is sound but requires some significant improvement, especially in better explaining the assumptions and limitations of the proposed approach, and the evaluation methodologies. In particular, it is clear that it’s hard to simulate real users, but the paper should at least make clear the assumptions and limitations of the simulation. Moreover one reviewer is rather surprised to see that SR@1 is improved so much over MM-VAE. Hence, it would be helpful to explain why, maybe with some case studies. Similarly, the same reviewer, considering Section 5.2, suggests that it would be helpful to better explain the content and add some a case study to help readers understand why the improvement is so significant. Another major weakness of this work is indicated by the second reviewer in the comparison with the related work. Moreover, the reviewer believes this paper is wrongly positioned in the literature, and the title is misleading too (too much focused on conversational recommendation). Please, read carefully the two reviews and update the manuscript accordingly. Reviewer Comments: Reviewer: 1 Recommendation: Major Revision Comments: The writing of this paper needs significant improvement. First, quite some notations are confusing and hard to understand. Second, multiple equations are presented without a clear explanation. It would be helpful to explain the intuition behind major equations/model so that it’s easier for readers to know why use those equations. I would suggest the authors have someone proofread the article. W is used in multiple equations/locations for different meanings. It’s better to use different notations. I can guess, however it’s better to point out r_{I,j} is 1 or 0 (some might think it is 1 or -1) Section 3.1: equation and explanation of projected linear recommendation seem weird and with wrong/confusing notations. Section 3.2: What’s p_{u,I,a} What does “takes the sum of user and item embeddings as input” and why you do that? “projection from the rationale space to the user preference space”: please explain rationale space Instead of max(k_{𝑈𝑢}, 1), the author may consider other way to smooth the value of K^U_{u}. BTW, K^U_{u} is a vector, and max over a vector and 1 (a scalar) is not a good notation, although I can guess what it means. It would be helpful to write clearly how 𝑠𝑢,𝑖,𝑎 is estimated. P8: “We optimize 𝑀RE via the linear regression:” again, please explain why you optimize this “We use a rule-based seeker with a simple prior: provided a target item and justification, it selects the most popular rationale present in the justification but not the target’s historical rationales k^𝐼_𝑖 to critique” this is confusing. Please explain in details and why the intuition/why simulate user this way? “cross-entropy loss between predicted scores and the goal item”: the score doesn’t seem to be a probability distribution, thus this statement is mathematically correct, although I can guess what you mean I found it hard to understand why/how you can separately optimize 𝑀rec, 𝑀just, and 𝑀RE, as the dependencies between those models are very tight. Please explain in more detail. It’s good to list the 3 research questions and doing some experiments to answer those. Besides numbers and figurers, it would be more helpful and insightful to explain the reasons and show some examples to illustrate what’s learned and why the proposed technique works so much better. I am very surprised to see SR@1 is improved so much over MM-VAE. It would be helpful to explain why, maybe with some case studies. Same for Section 5.2, it would be helpful to explain why and add some case studies to help readers understand why the improvement is so significant. The whole user simulation assumes user know the target item rationale thus can provide correct feedback. This may not be the case since. I understand it’s hard to simulate real users, while the paper should at least acknowledge the assumptions and limitations of the simulation. More detailed guidance for human annotator will be helpful? Did you actually tell the human annotator what’s the correct target item? It would be helpful to acknowledge the gap between MTurk users and real recommendations users, since the motivations for them are so different. Additional Questions: What is the key contribution of this paper?: This paper trains a recommender system to jointly suggest items and explain its reasoning via subjective rationales. It fine-tunes this model to incorporate iterative user feedback via self-supervised bot-play using product reviews for training. The major benefit is training the conversational recommendation model doesn’t require dialog transcripts. The work is evaluated based on simulated users and hired user subjects (Turks) and the results are very promising. a) Novelty (Comment on the novelty of the idea, and the methods and techniques proposed): Good Novelty and potential for innovation details: Training bot play using user review data seems novel. Relevance to the journal (TORS Scope Statement): Excellent Relevance to the journal: b) Soundness (Comment on the soundness of the idea, and the methods and techniques proposed.): Good soundness details: c) Experimental evaluation [Important guidelines]: (Do the authors use an appropriate methodology and suitable statistical methods? Is the choice of datasets, baseline methods and evaluation metrics in empirical comparisons well justified? Have all compared models in offline experiments been systematically tuned?): Fair Experimental evaluation details: The work is evaluated based on simulated users and hired user subjects (Turks). This is not ideal and there are some strong assumptions about user in the experiments. On the other hand, finding real user for such kind of research is hard for researchers and I believe it's ok do experimental evaluation this way. Experiments are not reproducible given the current writing. d) Theoretical background and evaluation: (Have the research questions and hypotheses been made explicit? Have all claims been substantiated through empirical evidence, theoretical considerations, or proofs?): Fair Theoretical background and evaluation details: Some hypotheses are not made explicit. Research questions are interesting and the claims are made based on experimental results with simulated or MTurk users. Comparison with previous and related works (Consider, for example: Are previous and related works adequately discussed? Are the relations of previous work to the current work clear?): Good Comparison with previous and related works details: Presentation: (Comment about the quality of English and the structure of the paper: Are there parts that should be rewritten, expanded upon, or dropped? Is the paper concise?): Fair Presentation details: Reasons to accept (Provide 1-3 concise bulleted-list items.): Reasonable approach with surprisingly good results. Thus it would be interesting to understand why Reasons to revise/reject (Provide 1-3 concise bulleted-list items.): 1) Writing needs significant improvement 2) It would be much better to explicitly state the assumptions and limitations of the proposed approach and evaluation methodologies Does this paper have potential real world significance?: Yes Real world significance comments: Reviewer: 2 Recommendation: Minor Revision Comments: The paper presents a framework for training conversational recommender systems using bot-play on historical user reviews without the need for large collections of human dialogues. Even though the framework is proposed as a solution for training conversational recommender systems, the authors address only a particular step. I recommend better placing the work in the literature and better highlighting the contribution by changing the title (if it is allowed by the journal rules). Furthermore, I recommend the authors provide more details that accompany the reader for the whole paper. Please, revise the bibliography style. Additional Questions: What is the key contribution of this paper?: The paper presents a framework for training conversational recommender systems using bot-play on historical user reviews without the need for large collections of human dialogues. The authors apply their framework to two popular recommendation models (BPR-Bot and PLRec-Bot), each showing superior or competitive performance compared to SOTA recommendation and critiquing methods. The authors demonstrate through human evaluation and user studies that models trained with their bot-play framework are more useful, informative, knowledgeable, and adaptive compared to SOTA baselines. a) Novelty (Comment on the novelty of the idea, and the methods and techniques proposed): Fair Novelty and potential for innovation details: The paper is an extended version of the RecSys 2022 conference paper entitled Self-Supervised Bot Play for Transcript-Free Conversational Recommendation with Rationales. Accordingly, the novelty is not so evident. Relevance to the journal (TORS Scope Statement): Good Relevance to the journal: The topic is fully relevant to the TORS Journal. b) Soundness (Comment on the soundness of the idea, and the methods and techniques proposed.): Good soundness details: The proposed framework is technically sound. However, since this is an extended version of a conference paper, the authors could provide more details to help the reader better understand the model. c) Experimental evaluation [Important guidelines]: (Do the authors use an appropriate methodology and suitable statistical methods? Is the choice of datasets, baseline methods and evaluation metrics in empirical comparisons well justified? Have all compared models in offline experiments been systematically tuned?): Good Experimental evaluation details: The experimental evaluation is basically the same as the RecSys paper. d) Theoretical background and evaluation: (Have the research questions and hypotheses been made explicit? Have all claims been substantiated through empirical evidence, theoretical considerations, or proofs?): Good Theoretical background and evaluation details: The research questions are clearly defined and answered. The results are well described and properly sustain the hypotheses. Comparison with previous and related works (Consider, for example: Are previous and related works adequately discussed? Are the relations of previous work to the current work clear?): Fair Comparison with previous and related works details: The main weakness of this work is the comparison with the related work. The authors claim that their framework avoids the need for expensive dialogue transcript datasets that limit the applicability of previous conversational recommender agents. However, i) they address only one step of a conversational recommendation (cor) (e.g., the critiquing); ii) they do not compare their model with any conversational recommender system. In my opinion, this paper is wrongly positioned in the literature. Perhaps, the title is misleading too (too much focused on conversational recommendation). Presentation: (Comment about the quality of English and the structure of the paper: Are there parts that should be rewritten, expanded upon, or dropped? Is the paper concise?): Fair Presentation details: In this paper, compared to the RecSys version (since more space is available), the authors could provide more details, examples, etc., to improve the readability. More specifically, the conversational recommendation is more focused on the user experience than the general accuracy with respect to a traditional recommender system. However, how the rest of the conversation (except for the critiquing step) could be managed remains obscure to the reader. In the experimental evaluation (sec. 4.1), some examples of how the rationales are extracted could be very useful. English is good. The references are often incomplete. Reasons to accept (Provide 1-3 concise bulleted-list items.): - the paper is well-written and organized - the research questions are well defined - the results are clear Reasons to revise/reject (Provide 1-3 concise bulleted-list items.): - the proposed framework is introduced as a solution for conversational recommender systems, but it addresses only a limited aspect - the motivations are limited - the presentation has room for improvement Does this paper have potential real world significance?: Maybe Real world significance comments: