Shilling Attacks on Recommender Systems

In this episode of Data Skeptic's Recommender Systems series, Kyle sits down with Aditya Chichani, a senior machine learning engineer at Walmart, to explore the darker side of recommendation algorithms. The conversation centers on shilling attacks—a form of manipulation where malicious actors create multiple fake profiles to game recommender systems, either to promote specific items or sabotage competitors. Aditya, who researched these attacks during his undergraduate studies at SPIT before completing his master's in computer science with a data science specialization at UC Berkeley, explains how these vulnerabilities emerge particularly in collaborative filtering systems. From promoting a friend's ska band on Spotify to inflating product ratings on e-commerce platforms, shilling attacks represent a significant threat in an industry where approximately 4% of reviews are fake, translating to $800 billion in annual sales in the US alone. The discussion delves deep into collaborative filtering, explaining both user-user and item-item approaches that create similarity matrices to predict user preferences. However, these systems face various shilling attacks of increasing sophistication: random attacks use minimal information with average ratings, while segmented attacks strategically target popular items (like Taylor Swift albums) to build credibility before promoting target items. Bandwagon attacks focus on highly popular items to connect with genuine users, and average attacks leverage item rating knowledge to appear authentic. User-user collaborative filtering proves particularly vulnerable, requiring as few as 500 fake profiles to impact recommendations, while item-item filtering demands significantly more resources. Aditya addresses detection through machine learning techniques that analyze behavioral patterns using methods like PCA to identify profiles with unusually high correlation and suspicious rating consistency. However, this remains an evolving challenge as attackers adapt strategies, now using large language models to generate more authentic-seeming fake reviews. His research with the MovieLens dataset tested detection algorithms against synthetic attacks, highlighting how these concerns extend to modern e-commerce systems. While companies rarely share attack and detection data publicly to avoid giving attackers advantages, academic research continues advancing both offensive and defensive strategies in recommender systems security.

Guest

Aditya Chichani: Aditya Chichani is a Senior Machine Learning Engineer at Walmart, where he designs and launches production-grade ML models and scalable systems that power search for millions of shoppers. His work focuses on ranking, retrieval, intent understanding, and bridging gaps in product attribute comprehension. Several initiatives he has led have driven significant GMV lifts and relevance gains across Walmart Search. Before Walmart, Aditya worked as a Software Engineer at Barclays, developing scalable microservices and real-time payment solutions for major clients such as Amazon. He holds a Master's degree in Electrical Engineering & Computer Sciences (EECS) from UC Berkeley, where he specialized in Data Science. He also spent time at the Berkeley Artificial Intelligence Research (BAIR) Lab, working on few-shot learning in NLP using semi-supervised methods under [Prof. Michael Mahoney](https://scholar.google.com/citations?user=QXyvv94AAAAJ&hl=en) and Francisco Utrera. Aditya has received Excellence Awards at Walmart and Barclays, as well as the Fung Excellence Scholarship at UC Berkeley. Beyond work, he mentors UC Berkeley affinity groups in AI and Data Science and stays active in the research community through conferences such as SIGIR, ICDM, CIKM, and RecSys.

Transcript

Kyle Polich : Welcome to Data Skeptic a podcast exploring the methods use cases and consequences of recommender systems

Kyle Polich : Welcome to another installment of Data skeptic recommender Systems For this episode or really for just the intro here let's put on what the cybersecurity community calls our black hats Forget about being a good guy

Kyle Polich : How can we manipulate recommender systems

Kyle Polich : What vulnerabilities do they have that we can take advantage of

Kyle Polich : Well a popular one is called a shilling attack

Kyle Polich : A shilling attack is when a malicious user probably one person really creates multiple profiles and then they start interacting with recommender systems in a way where they intend to manipulate its outputs

Kyle Polich : Usually that goal is to promote a specific item Imagine you're in a ska band you want to promote your ska band on a site like Bandcamp or something like that maybe Spotify where your first thing is to go and promote your own group you know whether it's upvoting or liking or favooring do all those things to your specific piece of content or artist or whatever but you can take that a step further Look at those items you're competing with and go and give them a down vote

Kyle Polich : Now if one person does that it kind of nets out in the noise but if a single user can puppeteer a bunch of accounts perhaps they can influence these networks

Kyle Polich : Perhaps they can influence the output of the recommender system Our guest today Aditya is a senior machine learning engineer at Walmart with a strong academic background researching these shilling attacks and other malicious strategies In this interview Aditya walks us through how collaborative filtering works As you should know by now that's one of the core algorithms common in recommender systems And as we understand that algorithm we can start to understand the different types of shilling attacks that it might be vulnerable to

Kyle Polich : Yet like anything it's a cat and mouse game just as a bad actor can try and take on these tasks Research approaches can be used to detect them We'll get into all that and more in this interview

Aditya Jaanni : I'm Aditya Jaanni I'm a senior machine learning engineer currently working at Walmart So I work for the Walmart search team I did my master from UC Berkeley and the undergrad research that we are going to be talking about today is from my undergrad which is SBIT

Kyle Polich : And can you share a few details on what you were studying in school Obviously machine learning but any specialty or focus within that

Aditya Jaanni : So

Aditya Jaanni : at Berkeley I did my master's in computer science but with a specialization in data science And in SPIT it was just you know bachelor's in Computer Science and I was you know trying to figure out OK where do my interests really lie right And I

Aditya Jaanni : This particular people that you're going to be talking about today I think it shaped you know my interests and this is one of those defining moments where I was like OK no uh this is something that I would not mind spending the rest of my you know life working on So yeah it was definitely that moment when I decided that OK you know I want to get into machine learning and search and recommendation seems to be something that I'm very interested in

Kyle Polich : What made that specifically an interesting problem worthy of study

Aditya Jaanni : Of course there are a lot of places where machine learning can be applied You could apply machine learning in something as specific as energy grid optimization right But recommendations in general it's something that each and every person regardless of whether he or she is an ML or not has interacted with right So it's very intuitive to think about OK you know

Aditya Jaanni : Now that I'm watching these movies on Netflix what other movies is Netflix going to recommend me right If I'm buying a product on let's say Walmart or Amazon you know it's seeing that OK as you buy more and more products how these companies kind of learn about you and then try to show you the most relevant products right So it's very interesting

Kyle Polich : Well I have a somewhat similar academic background to you and I was never really exposed to recommender systems I think I got a good education but I didn't have that elective or that topic What first got you exposed to it

Aditya Jaanni : My professor at SPIT she was doing research and recommendation systems So she was the one you know Kiran Gawande is the name and she was the one who kind of introduced us to uh recommendation systems and what kind of problems are there in recommendation systems And then once we started exploring

Aditya Jaanni : There's this very famous challenge called Netflix Price Challenge and recommendations right It it started in the mid 2000s and they had like a million dollars prize So I started reading about that and the more I read about you know the vast number of problems that are there in recommendations the more I got interested in it

Kyle Polich : And I know one of the big methodologies in recommender systems is collaborative filtering which uh I guess we're going to talk about as we get into the shilling attacks I I wanted to get to

Kyle Polich : But for listeners who maybe have just a passive familiarity what is collaborative filtering Right

Aditya Jaanni : so I would say collaborative filtering was one of the oldest and first methods of you know how you can do recommendations and I would say it's also one of the most intuitive ones So basically what you're doing in collaborative filtering is you know you have a user item ratings matrix right You know that OK these are the users and they have rated these items either explicitly or implicitly

Aditya Jaanni : And then what you're trying to do is OK if you're doing user user collaborative filtering what you're trying to figure out is there are these users who are probably who probably have similar tastes to let's say our source user

Aditya Jaanni : And they have seen these movies or these items right So because they are similar to each other maybe this user would also be interested in those items right So you're trying to build that understanding on let's say if two users are similar to each other and a user has interacted

Aditya Jaanni : With some items maybe the other user would also want to interact with it right So that's the core of user-user collaborative filtering And then an item item collaborative filtering it's doing the exact same thing but from item's point of view right So I would say the most

Aditya Jaanni : intuitive way to think about it is in Amazon you have if you bought this item you would also be interested in these items right Users who bought this also bought that Or let's say if you are buying a TV maybe you would also be interested in buying home theater So it's about creating this similarity matrix but between items

Kyle Polich : Well this seems like a really intuitive approach I can understand why people would pick it up and why it would give good results but what could go wrong

Aditya Jaanni : Specifically for user-based collaborative filtering right Usually what happens is I mean think about it you have millions of users but then you probably have like 100s of millions of items right So and at any given 0.1 particular user would have interacted with very few items right Think about how many items did you buy

Aditya Jaanni : On Amazon or on Walmart in the last one year and then compare it with how many total number of items would be present in their catalog right So the signal is very sparse To begin with that's a big problem that OK you don't really have a lot of data especially for user user collab filtering And then you could also expose yourself to problems such as shilling attacks that we'll get into where essentially what people would do is OK

Aditya Jaanni : Now that they know that user-user collaborative filtering is a very common method that companies could be using you could create fake profiles you could try to become rate most popular items similarly as genuine users and then you could try to either promote your own item or nuke your competitor's item And then because of this uh you know even genuine items would kind of start feeling that impact on recommendations right

Aditya Jaanni : So you also expose yourself to such problems when you start using collaborative filtering

Kyle Polich : So that's one particular type of shilling attack that promotion I guess in practical terms let's say we were doing music recommendation and I wanted to promote my friend's band I guess would I like big acts like Taylor Swift and the Beatles and then like my friend's band and have enough accounts that people would start to associate those three together Is that the basic idea

Aditya Jaanni : What you

Aditya Jaanni : have talked about is I would say segmented attack where essentially what you're trying to do is

Aditya Jaanni : OK you know that these particular items are very popular These particular music albums are very popular and they are similar to the genre of your friend's music right So essentially first you listen you know you create profiles and listen to these musics so that you know you kind of have co-occurrence with a lot of genuine users And then you also start to you know highly rate your friend's music uh listen to it more give it a lot of more places so that

Aditya Jaanni : It shows up as a top recommendation and then it will start getting recommended to people who have also listened to Taylor Swift for example

Kyle Polich : Well in a case like that you're assuming my friend's band is not particularly popular it's a real David and Goliath you have a giant with millions or more downloads and someone with maybe less than 1000 Could there be a better strategy to it Should I maybe find like a regionally popular band but not too popular Would I have more success shilling through a strategic approach like that

Aditya Jaanni : So there's two things right There's selected items and there's target items So selected items are the ones that are already hugely popular and you know you are listening to them or rating them highly so that essentially you kind of have some connections with genuine users and then you know you kind of try to push your own target item I'm assuming your point was what if your target item is also popular

Kyle Polich : Right I can make a limited number of fake accounts let's say 500 uh maybe that's a drop in the bucket for a top 40 artists or something but it could sway a ratio if it's a smaller group

Aditya Jaanni : So for selected items right let's say if Taylor Swift's album already has like millions of listeners right Your goal by creating those 500 profiles and listening to Taylor Swift's music

Aditya Jaanni : Music is not so that you can boost Taylor Swift's music right It's to kind of create connections with other millions of listeners who are already actively listening to Taylor Swift's music and actually like it right And now you start promoting your own or your friend's music which is of course you know uh naturally you would only have incentive to do that but you know to use a shilling profile when

Aditya Jaanni : Your French music is already not so popular right And you're trying to get it more viral right So your French music doesn't have so many listeners but now that these profiles have listened to Taylor Swift's music and they feel like genuine profiles these fake profiles will start listening to your French music as well And the idea is that now because of this association

Aditya Jaanni : Of the recommendation system would think that OK you know these profiles also listen to friends uh you know your friend's music And because they are similar to these other genuine profiles let me also start recommending this uh you know friends's music to all these genuine profiles right So that's the end goal

Kyle Polich : So I think that's what we call the segmented attack Are there other vulnerabilities here we should be worried about

Aditya Jaanni : Yeah

Aditya Jaanni : absolutely So all of these chilling attacks essentially depend on how much information do you have about the system right So let's say if you know nothing absolutely nothing right So let's say there are no explicit ratings you know the way you have for IMDb or something of that sort

Aditya Jaanni : And all you know is that OK on an average the rating for all of the items is let's say 3 on a scale of 1 to 5 right

Aditya Jaanni : So if you have absolutely no knowledge apart from you know overall average distribution what you could do is just when you create these attacker profiles you also have pillar items where essentially what you're trying to do is create rate these pillar items to kind of you know go undetected in the system

Aditya Jaanni : For these filler items right all you could do is just give a random rating between 1 to 5 Let's say if 3 is the average you give it 3 right So this could be a random attack So this is the least I would say powerful chilling attack but then because you know you don't really have a lot of information about the system

Aditya Jaanni : Then you have average attack where essentially let's say you know that for each filler item what is the average rating that that particular filler item gets And then you particular you try to give that average rating So for example let's say you're trying to attack a movie recommendation system You already know that OK you know Harry Potter let's say has IMDb of 8 or whatever right

Aditya Jaanni : And it's a popular movie So you give it a higher rating when you are saying that you like the movie from 1 to 5 you give it a higher rating So it's a better attack than random but it also requires you to have the average rating of each pillar item And then an interesting attack is bandwagon attack where instead of you know targeting a specific segment

Aditya Jaanni : You select items which are very popular right You already know that OK these are popular items they have been uh you know interacted with by a lot of genuine users So you start rating those popular items highly right Because you know they are popular by default know that they must be highly rated So you just start rating them highly to kind of you know get into to create connections with genuine users So these are like the different kind of shilling attacks

Kyle Polich : Do you have a sense of how many fake profiles one would need to make an impact

Aditya Jaanni : I would say it depends on what kind of system it is For user user collaborative filtering right you don't really need to create too many profiles right It could be a very small set because of how user-user collaborative filtering works

Aditya Jaanni : For one particular user you don't have a lot of other users who would have seen or you know interacted with the exact same items as this particular user has right So it's a very small subset

Aditya Jaanni : And because of that let's say you create like 500 shilling profiles or even lesser right depending on how big your system is it's very easy to kind of break into that subset And that's why user user collaborative filtering is essentially a lot more prone to these attacks But instead of that if you think about item item collaborative filtering right

Aditya Jaanni : At any given point let's say a user maybe interacts with 10 items but if you think from the item's point of view so many users would interact with that particular item so the signal is a lot stronger and because of that what happens is you would essentially have to you know

Aditya Jaanni : Create a lot more shilling profiles You would have to create a lot more items to kind of create the same impact that you would want let's say on a user you guys here So it would be a lot more expensive to you know attack a system which is based on item-based collaborative filtering for example

Kyle Polich : We've talked about how this can be done against music systems and movie systems and probably just about any system that uses user user collaborative filtering with this risk out there it begs the question can you detect this

Aditya Jaanni : Yeah absolutely As we have talked about the shelling attacks right one thing that you would have noticed is how shilling attackers behave So essentially what they are trying to do is they are trying to first highly rate popular items or let's say you know give average rating to a lot of items

Aditya Jaanni : And then specifically they try to either push their own target item or nuke let's say you know arrivals item right So this behavior is kind of different from what a genuine user would do right

Aditya Jaanni : So essentially you kind of bank on this One way to do it is to figure out OK how many other user profiles does this particular profile have a high correlation with right How similar is this profile to a lot of other profiles Because essentially that would be the first goal of a shilling attacker right To become similar to a lot of profiles to kind of amplify the impact of their attack

Aditya Jaanni : And then you look for behavior where you know let's say for this particular target item your other genuine users would rate it between a scale of 1 to 5 with no particular focus on let's say one particular rating right

Aditya Jaanni : But your shilling attacker would either always rate that particular set of items highly or very poorly So that is kind of what you bank on So uh some way that you kind of do shilling attack detection is either you use PC where you're trying to kind of you know represent all of these users into a smaller you know lower dimensional latent space

Aditya Jaanni : Usually genuine users kind of fall in the same cluster but users with shilling attacker profiles would have a different distribution because they have different goals So because of that they kind of tend to fall out of this cluster So that's one way which is used for attack detection

Aditya Jaanni : And of course it's a cat and mouse game right The more you improve how to detect schilling attacker profiles the more shilling attackers kind of try to improvise Although we are talking about collaborative filtering as of now collaborative filtering is a pretty old method Now you have these

Aditya Jaanni : Much more advanced recommendation systems So why are we talking about chilling attacks now So chilling attacks have also evolved in that way So let's say earlier you would just have fake ratings Recommendation systems kind of evolved to also include other signals such as reviews right

Aditya Jaanni : But shilling attackers now also create fake reviews to kind of promote or nuke the item So it's also that where the recommendation systems are trying to use very different kinds of signals to create a more robust recommendation

Aditya Jaanni : But the ensuing attackers are also trying to you know kind of mimic that approach so that they are harder to detect And as long as there is some benefit for the malicious attackers to gain from it right they would always try to do it So if you think about it I think that I was reading a recent report by World Economic Forum where they mentioned that around 4%

Aditya Jaanni : Of reviews today are fake and you would think that 4% is a small number but it kind of translates to around $800 billion annual sales just in the US alone So it's a very big market And because it's so beneficial right Let's say you think about Yelp or you think about Amazon giving fake reviews good reviews for your own item

Aditya Jaanni : would make such a big monetary difference right So there's a lot of advantage to gain for the shilling attackers So it's kind of OK for them to put a lot more effort and investment into it Before let's say if they just try to predict average ratings now they are using LLMs to generate fake reviews So they have also evolved to kind of you know ensure that they still

Aditya Jaanni : End up getting that monetary gain and that benefit from doing these shilling attacks

Kyle Polich : Delete me makes it easy quick and safe to remove your personal data online at a time when surveillance and data breaches are common enough to make everyone vulnerable

Kyle Polich : Want an easier way to deal with data breaches Get Delete me The fact is we're all at risk How many times have you gotten an email or a letter saying your data has been breached It's unsettling But the good news is DeleteM can help DeleteM does all the hard work of wiping you and your family's personal information from data broker websites

Kyle Polich : As someone with an active online presence privacy is really important to me I've been shocked at how much of my personal information was floating around on data broker sites Since using DeleteM I've received detailed reports showing exactly what they've found and removed giving me peace of mind knowing my digital footprint is being minimized Take control of your data

Kyle Polich : and keep your private life private by signing up for Delete Me Now at a special discount for our listeners Get 20% off your Delete Me plan when you go to joindeleteme.com/data and use the promo code data at checkout The only way you get the 20% off is to go to joindeleteme.com/data Enter code data at checkout That's joindeleteme.com/datacodedata

Kyle Polich : Well I'm thinking about your methodology and how you're describing detecting them and one thing that occurred to me is you often have people with similar tastes like let's stick to the movie domain people that like campy B movie horror films right it's sort of a small genre but loyal fans that love that kind of stuff How do you tell the difference between those people who might all see all of the main movies and someone doing a shilling attack

Aditya Jaanni : Yeah absolutely So essentially what you're talking about is we don't want to have a lot of false positives when we are detecting these shilling profiles A lot of companies would think that it's kind of OK to maybe end up having a few shilling attackers in their system then you know weeding out genuine users and you know creating that mistrust for that company right

Aditya Jaanni : So the overall idea is to kind of avoid these false positives So essentially what you do in those cases is you try to come up with heuristics on how you know you are going to assume whether a profile is a shilling attack or not And I'll give you more context on that So essentially let's say we have decided

Aditya Jaanni : That this is the correlation threshold above which we are going to consider that these two users are similar and this is the profile number of profiles threshold where we will consider that OK let's say if this particular user is similar to 100 other users so this

Aditya Jaanni : 100 number this number of profiles threshold is another metric on the basis of which you will decide whether the profile is a shilling attacker or not So you tune these metrics right You tune these parameters on at what point at what correlation threshold do you assume whether these profiles are similar

Aditya Jaanni : And at what number of profiles threshold do you start assuming that this profile is a potential shilling attacker profile right So that's one way And the other way is it can be a tiered approach right Essentially you shouldn't just do that OK you know any profile which is uh meeting these particular thresholds is by default a shilling attacker profile and will you know directly remove this account right

Aditya Jaanni : Usually what companies do is they have a tiered approach where first they kind of get a potential subset of profiles uh which are kind of tagged as suspicious profiles and then you kind of pass it through multiple either different models where your models in the latter stage could be more advanced right Because essentially you know how it works in search and recommendation systems where you have a retrieval stage and then you have a ranking stage

Aditya Jaanni : Where first you are going to get a sim use simple heuristics and try to get a big potential subset and then within that you use more modified and you know more advanced models to finally decide whether this is actually a positive or not whether it is a shilling attacker profile or not So you can use that tiered approach maybe have a human in the loop then decide finally whether it is a shilling attacker profile or not

Kyle Polich : Well one very popular data set you could look at for benchmarking some of your techniques is the Movie lens database Could you comment on uh your experience with that or any other ways that you looked at it

Aditya Jaanni : Yeah so you know when I started I was first thinking about OK maybe we should use the Netflix price tagging data set but it's a pretty big data set And at that time the idea was to kind of first focus on OK how do we design recommendation systems how do we attack these systems and then

Aditya Jaanni : How do we detect these attacks right So it was not mainly on the data set size but of course the numbers would get more and more reliable as you kind of increase the data set size So Group lens MovieLens is I would say one of the most popular data sets for recommendation problems

Aditya Jaanni : Group lens has multiple tiers of data sets So the smallest data set I believe is the 100K ratings data set which is what we use right So it has around 1000 users and around 1700 items and overall around 100K ratings for this user cross items right So this is the data set that we used and I I believe it's something that was

Aditya Jaanni : Created by University of Minnesota So a huge thank you to them for creating this So that was the data set that we used and it's very clean and it has a lot more information than just user item ratings right So it also has information about demographics and all of that And because it was uh as far as I remember it was created

Aditya Jaanni : Purely from for academia and by users who volunteered to you know be part of that uh data set collection So you don't have those privacy concerns or came up with Netflix right where users didn't really choose to be part of that data set choose to be part of that challenge right That's I would say also a positive point for the data set

Kyle Polich : Well given that policy of data collection I wouldn't expect there to be any shilling attacks in the MovieLenss data set so what is there to detect

Aditya Jaanni : Yeah that's a good point That's why when we created the data set we injected shilling profiles first So we created a synthetic data set where essentially what we tried to do was depending on what each kind of attack would be right So let's say if you are thinking about average attack

Aditya Jaanni : You would know the average rating of those filler items right So we synthetically generated essentially OK we know that for these 10 items these are the average ratings and we would as a shilling attacker we would start rating these items that way

Aditya Jaanni : So based on what each kind of attack was we tried to mimic that particular attack into group lens uh into our mobile lens data set and then we created a separate attack detection model where we were trying to detect each of these attacks where of course you know you wouldn't know what kind of attack was done on the model It's a black box for you and then you're just trying to detect attacker profiles

Aditya Jaanni : I don't think a lot of companies would be willing to provide this data right because it's so sensitive right No company would willingly provide this data on OK you know how many fake profiles were created for our company how did we detect them because if you tell how you are detecting them it could give an edge to the attackers

Aditya Jaanni : Right The attackers would try to then game that So it is also very sensitive field because of which a lot of times you would see that all the papers on shilling attacks are mostly from academia right Because companies for natural reasons do not want to kind of release their data publicly

Kyle Polich : If you're injecting some shilling attacks if you put just one in maybe that would go under the radar Could you talk a little bit about the degree to which you have to synthesize this for it to be detectable and also how you detect it

Aditya Jaanni : Yeah absolutely So we did one of the parameters apart from your profile threshold which is the number of attacker profiles and correlation threshold We also had this parameter on attack size right in terms of OK for the same number of shilling profiles

Aditya Jaanni : How many let's say uh popular items would you have to rate in order to you know kind of create impact How many items otherwise let's say if you're creating a segmented attack how many other items that belong to the same segment would you have to kind of rate in order to create this attack detection

Aditya Jaanni : And then we kind of plotted a graph on OK as you increase the attack threshold what happens to your overall detection So it is true that as you kind of increase the attack size it also becomes easier to detect these attacks but it's balanced If you just create one profile you are not going to have any impact on your items right

Aditya Jaanni : Just creating one fake profile and highly rating your target item is not going to make any dent at all So just creating one profile is not very useful for the attacker and they kind of try to create this balance where you want to have as many attacker profiles and as big of an attack as you can while flying under the radar So it's a pretty hard job but that's the whole cat and mouse game again

Kyle Polich : As you go through that principal components approach we discussed how I guess separated does the data become Is it obvious that you have shilling attacks or could it just be one user or a few users that have very similar tastes

Aditya Jaanni : From what we saw it's more obvious when let's say the attacks are let's say random attack or average attack because the attackers don't really have a lot of information

Aditya Jaanni : And then for a lot of those items they are rating it exactly the same way Let's say for random attacks all the filler items would have somewhat the average distribution rating right So in such cases when you apply PCA it's pretty easy to detect such attacks

Aditya Jaanni : But then for attacks with higher knowledge right it gets difficult to reliably find such profiles because a lot of your genuine profiles would also behave in the same way And then there are other ways in which attackers kind of try to go undetected

Aditya Jaanni : Example the whole idea of this attack detection and PCA is to find attacker profiles by deciding that OK how are they behaving differently from our genuine users So that is the part that then shilling attackers play on that maybe let's say

Aditya Jaanni : Instead of always highly rating your target item let's say you're always giving it the perfect rating of 5 maybe you vary between 3 to 5 for your target item So you're trying to mimic a genuine user's way to kind of go undetected Maybe when you want to nuke an item you don't always give it the worst rating possible you just give it a lower rating So there are these other ways or maybe you add noise in your filler items Maybe you don't always give it an average rating

Aditya Jaanni : You give it an average rating with some standard deviation So there are these other ways to kind of try to behave as genuinely as possible to kind of go and detect it

Kyle Polich : Is there anywhere listeners could follow you online as well

Aditya Jaanni : Yeah absolutely They could uh reach out to me on LinkedIn Uh my handle is just my first and last name Ad Teachanni Uh if at any point there are people who would kind of want to connect with me maybe in person to talk about the search and recommendations problem or

Aditya Jaanni : Selling attacks in general then I would be coming to ICDM conference this year I'm organizing a multimodal source and recommendations workshop which is something that you know I have been doing you know for some time I have organized workshops at CITM and CIR So yeah if they want to kind of talk to me about these things they could come there and meet me in person

Kyle Polich : And is there anything you can comment on about how this may or may not relate to your uh day job and uh whether or not you get to do fun stuff like this at uh your uh corporate gig

Aditya Jaanni : Absolutely So like I said right I work at Walmart so so I work on ranking which is related to recommendations but it's a slightly different problem where essentially for recommendations you don't have an explicit query But for search ranking your customer has explicitly searched for something

Aditya Jaanni : And then you are trying to figure out OK what items to show to that customer So it's still an extremely interesting problem The stakes are even higher because for recommendations for example if you don't show the most relevant items to the user the user won't take offense But let's say if the user explicitly searches for something and you still don't show relevant products to the customer then it would be a problem So that's something that I currently work on but my current focus of work is not

Aditya Jaanni : on these attacks per se or you know chilling attacks or something of that sort It's more on this is the customer's query how do we ensure that we are showing the most relevant products to the customer something that they would like to purchase and so on and so forth

Kyle Polich : Thank you so much for taking the time to come on and share your work

Aditya Jaanni : Absolutely Yeah it was great talking to you

Shilling Attacks on Recommender Systems