Publications – Analytics of Software, GAmes And Repository Data (ASGAARD) Lab

1.

Markos Viggiato; Dale Paas; Cor-Paul Bezemer

Prioritizing Natural Language Test Cases Based on Highly-Used Game Features Inproceedings

Proceedings of the 31st Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering (ESEC/FSE), pp. 1–12, 2023.

Files:

Abstract | BibTeX | Tags: Computer games, Game development, Natural language processing, Testing

@inproceedings{ViggiatoFSE2023,

title = {Prioritizing Natural Language Test Cases Based on Highly-Used Game Features},

author = {Markos Viggiato and Dale Paas and Cor-Paul Bezemer },

year  = {2023},

date = {2023-12-01},

urldate = {2023-12-01},

booktitle = {Proceedings of the 31st Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering (ESEC/FSE)},

pages = {1--12},

abstract = {Software testing is still a manual activity in many industries, such

as the gaming industry. But manually executing tests becomes im-

practical as the system grows and resources are restricted, mainly

in a scenario with short release cycles. Test case prioritization is a

commonly used technique to optimize the test execution. However,

most prioritization approaches do not work for manual test cases

as they require source code information or test execution history,

which is often not available in a manual testing scenario. In this

paper, we propose a prioritization approach for manual test cases

written in natural language based on the tested application features

(in particular, highly-used application features). Our approach con-

sists of (1) identifying the tested features from natural language test

cases (with zero-shot classification techniques) and (2) prioritizing

test cases based on the features that they test. We leveraged the

NSGA-II genetic algorithm for the multi-objective optimization of

the test case ordering to maximize the coverage of highly-used

features while minimizing the cumulative execution time. Our find-

ings show that we can successfully identify the application features

covered by test cases using an ensemble of pre-trained models

with strong zero-shot capabilities (an F-score of 76.1%). Also, our

prioritization approaches can find test case orderings that cover

highly-used application features early in the test execution while

keeping the time required to execute test cases short. QA engineers

can use our approach to focus the test execution on test cases that

cover features that are relevant to users.},

keywords = {Computer games, Game development, Natural language processing, Testing},

pubstate = {published},

tppubtype = {inproceedings}

}

Close

2.

Markos Viggiato

Leveraging Natural Language Processing Techniques to Improve Manual Game Testing PhD Thesis

2023.

Files:

Abstract | BibTeX | Tags: Computer games, Game development, Natural language processing, Testing

@phdthesis{ViggiatoPhD,

title = {Leveraging Natural Language Processing Techniques to Improve Manual Game Testing},

author = {Markos Viggiato },

year = {2023},

date = {2023-01-17},

urldate = {2023-01-17},

abstract = {The gaming industry has experienced a sharp growth in recent years, surpassing other popular entertainment segments, such as the film industry. With the ever-increasing scale of the gaming industry and the fact that players are extremely difficult to satisfy, it has become extremely challenging to develop a successful game. In this context, the quality of games has become a critical issue. Game testing is a widely-performed activity to ensure that games meet the desired quality criteria. However, despite recent advancements in test automation, manual game testing is still prevalent in the gaming industry, with test cases often described in natural language only and consisting of one or more test steps that must be manually performed by the Quality Assurance (QA) engineer (i.e., the tester). This makes game testing challenging and costly. Issues such as redundancy (i.e., when different test cases have the same testing objective) and incompleteness (i.e., when test cases miss one or more steps) become a bigger concern in a manual game testing scenario. In addition, as games become bigger and the number of required test cases increases, it becomes impractical to execute all test cases in a scenario with short game release cycles, for example.

Prior work proposed several approaches to analyze and improve test cases with associated source code. However, there is little research on improving manual game testing. Having higher-quality test cases and optimizing test execution help to reduce wasted developer time and allow testers to use testing resources more effectively, which makes game testing more efficient and effective. In addition, even though players are extremely difficult to satisfy, their priorities are not considered during game testing. In this thesis, we investigate how to improve manual game testing from different perspectives.

In the first part of the thesis, we investigated how we can reduce redundancy in the test suite by identifying similar natural language test cases. We evaluated several unsupervised approaches using text embedding, text similarity, and cluster-

ing techniques and showed that we can successfully identify similar test cases with a high performance. We also investigated how we can improve test case descriptions to reduce the number of unclear, ambiguous, and incomplete test cases. We proposed and evaluated an automated framework that leverages statistical and neural language models and (1) provides recommendations to improve test case descriptions, (2) recommends potentially missing steps, and (3) suggests existing similar test cases.

In the second part of the thesis, we investigated how player priorities can be included in the game testing process. We first proposed an approach to prioritize test cases that cover the game features that players use the most, which helps to avoid bugs that could affect a very large number of players. Our approach (1) identifies the game features covered by test cases using an ensemble of zero-shot techniques with a high performance and (2) optimizes the test execution based on highly-used game features covered by test cases. Finally, we investigated how sentiment classifiers perform on game reviews and what issues affect those classifiers. High-performing classifiers can be used to obtain players' sentiments about games and guide testing based on the game features that players like or dislike. We show that, while traditional sentiment classifiers do not perform well, a modern classifier (the OPT-175B Large Language Model) presents a (far) better performance. The research work presented in this thesis provides deep insights, actionable recommendations, and effective and thoroughly evaluated approaches to support QA engineers and developers to improve manual game testing.},

keywords = {Computer games, Game development, Natural language processing, Testing},

pubstate = {published},

tppubtype = {phdthesis}

}

Close

The gaming industry has experienced a sharp growth in recent years, surpassing other popular entertainment segments, such as the film industry. With the ever-increasing scale of the gaming industry and the fact that players are extremely difficult to satisfy, it has become extremely challenging to develop a successful game. In this context, the quality of games has become a critical issue. Game testing is a widely-performed activity to ensure that games meet the desired quality criteria. However, despite recent advancements in test automation, manual game testing is still prevalent in the gaming industry, with test cases often described in natural language only and consisting of one or more test steps that must be manually performed by the Quality Assurance (QA) engineer (i.e., the tester). This makes game testing challenging and costly. Issues such as redundancy (i.e., when different test cases have the same testing objective) and incompleteness (i.e., when test cases miss one or more steps) become a bigger concern in a manual game testing scenario. In addition, as games become bigger and the number of required test cases increases, it becomes impractical to execute all test cases in a scenario with short game release cycles, for example.

Prior work proposed several approaches to analyze and improve test cases with associated source code. However, there is little research on improving manual game testing. Having higher-quality test cases and optimizing test execution help to reduce wasted developer time and allow testers to use testing resources more effectively, which makes game testing more efficient and effective. In addition, even though players are extremely difficult to satisfy, their priorities are not considered during game testing. In this thesis, we investigate how to improve manual game testing from different perspectives.

In the first part of the thesis, we investigated how we can reduce redundancy in the test suite by identifying similar natural language test cases. We evaluated several unsupervised approaches using text embedding, text similarity, and cluster-
ing techniques and showed that we can successfully identify similar test cases with a high performance. We also investigated how we can improve test case descriptions to reduce the number of unclear, ambiguous, and incomplete test cases. We proposed and evaluated an automated framework that leverages statistical and neural language models and (1) provides recommendations to improve test case descriptions, (2) recommends potentially missing steps, and (3) suggests existing similar test cases.

In the second part of the thesis, we investigated how player priorities can be included in the game testing process. We first proposed an approach to prioritize test cases that cover the game features that players use the most, which helps to avoid bugs that could affect a very large number of players. Our approach (1) identifies the game features covered by test cases using an ensemble of zero-shot techniques with a high performance and (2) optimizes the test execution based on highly-used game features covered by test cases. Finally, we investigated how sentiment classifiers perform on game reviews and what issues affect those classifiers. High-performing classifiers can be used to obtain players' sentiments about games and guide testing based on the game features that players like or dislike. We show that, while traditional sentiment classifiers do not perform well, a modern classifier (the OPT-175B Large Language Model) presents a (far) better performance. The research work presented in this thesis provides deep insights, actionable recommendations, and effective and thoroughly evaluated approaches to support QA engineers and developers to improve manual game testing.

Close

3.

Markos Viggiato; Dayi Lin; Abram Hindle; Cor-Paul Bezemer

What Causes Wrong Sentiment Classifications of Game Reviews? Journal Article

IEEE Transactions on Games, pp. 1–14, 2021.

Files:

Abstract | BibTeX | Tags: Computer games, Natural language processing, Sentiment analysis, Steam

@article{markos2021sentiment,

title = {What Causes Wrong Sentiment Classifications of Game Reviews?},

author = {Markos Viggiato and Dayi Lin and Abram Hindle and Cor-Paul Bezemer},

year  = {2021},

date = {2021-04-05},

urldate = {2021-04-05},

journal = {IEEE Transactions on Games},

pages = {1--14},

institution = {University of Alberta},

abstract = {Sentiment analysis is a popular technique to identify the sentiment of a piece of text. Several different domains have been targeted by sentiment analysis research, such as Twitter, movie reviews, and mobile app reviews. Although several techniques have been proposed, the performance of current sentiment analysis techniques are still far from acceptable, mainly when applied in domains on which they were not trained. In addition, the causes of wrong classifications are not clear. In this paper, we study how sentiment analysis performs on game reviews. We first report the results of a large scale empirical study on the performance of widely-used sentiment classifiers on game reviews. Then, we investigate the root causes for the wrong classifications and quantify the impact of each cause on the overall performance. We study three existing classifiers: Stanford CoreNLP, NLTK, and SentiStrength. Our results show that most classifiers do not perform well on game reviews, with the best one being NLTK (with an AUC of 0.70). We also identified four main causes for wrong classifications, such as reviews that point out advantages and disadvantages of the game, which might confuse the classifier. The identified causes are not trivial to be resolved and we call upon sentiment analysis and game researchers and developers to prioritize a research agenda that investigates how the performance of sentiment analysis of game reviews can be improved, for instance by developing techniques that can automatically deal with specific game-related issues of reviews (e.g., reviews with advantages and disadvantages). Finally, we show that training sentiment classifiers on reviews that are stratified by the game genre is effective.},

keywords = {Computer games, Natural language processing, Sentiment analysis, Steam},

pubstate = {published},

tppubtype = {article}

}

Close