<ici-import>
 <journal 	issn="2538-421X"/>
 <issue number="1" volume="21" year="2024" publicationDate="2024-06-01" numberOfArticles="10">
			<article externalId="A-10-2025-2">
			<type>OTHERS_CITABLE</type>
			
					<languageVersion language="fa">
						<title>یک الگوریتم خوشه‌بندی جدید چندگامه مبتنی بر تاخیر تکرار برای بهبود کیفیت سرویس اینترنت اشیاء</title>
						<abstract>اینترنت اشیاء به شبکه&#8204;ای از اشیاء فیزیکی اطلاق می&#8204;شود که در آن اشیاء دارای شناسه&#8204;ی منحصر، قادر هستند با یکدیگر و یا با کاربر نهایی از طریق اینترنت ارتباط برقرار کنند. به دلیل محدود بودن برد رادیویی اشیاء و همچنین کاهش انرژی مصرفی، انتقال اطلاعات از طریق اشیاء واسط انجام می&#8204;شود که لزوم مسیریابی را مشخص می&#8204;کند. یک الگوریتم مسیریابی بطور مستقیم بر قابلیت اطمینان، تاخیر، انرژی مصرفی، گذردهی شبکه، استفاده&#8204;ی موثر از پهنای باند و طول عمر شبکه تاثیر می&#8204;گذارد. در این مقاله یک رویکرد مسیریابی جدید مبتنی بر خوشه&#8204;بندی توزیع&#8204;شده و تاخیر تکراری برای بهبود کیفیت سرویس اینترنت اشیاء پیشنهاد می&#8204;شود که اشیاء شبکه به تعدادی خوشه مجزا از هم تقسیم می&#8204;شود. خوشه&#8204;بندی بر اساس حالت&#8204;های مختلف اشیاء همسایه انجام می&#8204;شود و برای ارسال داده&#8204;های سرخوشه&#8204;ها به ایستگاه پایه نیز از روش تاخیر تکرار استفاده می&#8204;شود. در رویکرد پیشنهادی برای انتخاب سرخوشه&#8204;ی واسط مناسب، تاخیر انتقال از شیء سرخوشه تا شیء واسط، تاخیر انتقال از شیء واسط تا چاهک و انرژی باقیمانده&#8204;ی شیء واسط در نظر گرفته می&#8204;شود. نتایج شبیه&#8204;سازی&#8204;های انجام شده در نرم&#8204;افزار Cooja نشان می&#8204;دهد که رویکرد پیشنهادی بطور متوسط در مقایسه با الگوریتم&#8204;های LEACH، LEACH-E، NCACM، Distributed Clustering از نظر میزان مصرف انرژی 33% ، مرگ اولین گره 14%، تعداد اشیاء مرده 12% و نرخ دریافت صحیح بسته&#8204;ها 9% عملکرد بهتری دارد.</abstract>
						<pdfFileUrl>http://jsdp.rcisp.ac.ir/article-1-1279-fa.pdf</pdfFileUrl>
						<publicationDate>2024-08-03</publicationDate>
						<pageFrom>3</pageFrom>
						<pageTo>14</pageTo>
				
							<doi>10.61186/jsdp.21.1.3</doi>
						<keywords>
<keyword>اینترنت اشیاء</keyword>
<keyword>خوشه‌بندی توزیع‌شده</keyword>
<keyword>تاخیر تکرار</keyword>
<keyword>مسیریابی</keyword>
<keyword>کیفیت سرویس خدمات</keyword>
</keywords>
				</languageVersion>
				

					<languageVersion language="en">
						<title>A new multi-hop clustering algorithm based on iterative delay to enhance QoS for Internet of Things</title>
						<abstract>In general, Internet of Things (IoT) as a new technology refers to a network of physical things in which objects have a unique identity and are able to communicate with each other or with the end user via the Internet. &#160;The Internet of Things refers to a collection of sensor-embedded devices, processing ability, software, and other technologies that connect and exchange data with other devices and systems over the Internet or other communications networks. Due to the limited radio range of objects and also the reduction of energy consumption, information transmission is carried out through the intermediate objects, which highlights the necessity for routing. Routing algorithms can be classified into static and dynamic techniques as well as source initiated and destination initiated approaches. In general, routing algorithms can be classified into data centric, hierarchical, geographical, and quality of service-based mechanisms. A routing algorithm directly affects reliability, transmission latency, power consumption, network throughput, bandwidth utilization, and network lifetime. This paper proposes a new routing method based on distributed clustering and iterative latency to improve the Quality of Service (QoS) of IoT, which divides network things into a number of separate clusters. The proposed method consists of four stages, i.e. network clustering, steady state, multi-hop transmission based on delay estimation, and investigation of adjacent headers. Clustering is performed based on the different states of neighbors, and the iterative delay mechanism is used between the cluster heads. The simulation results conducted through Cooja tool indicate that the proposed method outperforms the LEACH, LEACH-E, NCACM, and distributed clustering techniques in terms of energy consumption and packet delivery ratio by 33% and 9%. Furthermore, simulation results illustrate that the proposed method outperforms in terms of the first node death time and the number of dead objects in scattered and dense networks by 14% and 12%, respectively.</abstract>
						<pdfFileUrl>http://jsdp.rcisp.ac.ir/article-1-1279-en.pdf</pdfFileUrl>
						<publicationDate>2024-08-03</publicationDate>
						<pageFrom>3</pageFrom>
						<pageTo>14</pageTo>
				
							<doi>10.61186/jsdp.21.1.3</doi>
						<keywords>
<keyword>Internet of Things (IoT)</keyword>
<keyword>Distributed clustering</keyword>
<keyword>Iterative delay</keyword>
<keyword>Routing</keyword>
<keyword>Quality of Service (QoS)</keyword>
</keywords>
				</languageVersion>
				


	<authors>
	<author>
	<name>Farnaz</name>
	<surname>Rasekhi</surname>
	<email>farnaz.rasekhi@gmail.com</email>
	     <order>1</order>
        <instituteAffiliation>Islamic Azad University, Tabriz Branch</instituteAffiliation>  
	    <role>AUTHOR</role>
	 </author>
	<author>
	<name>Shahram</name>
	<surname>Babaie</surname>
	<email>sh.babaie@iaut.ac.ir</email>
	     <order>2</order>
        <instituteAffiliation>Islamic Azad University, Tabriz Branch</instituteAffiliation>  
	    <role>AUTHOR</role>
	 </author>
	</authors>


	</article>



			<article externalId="A-10-2451-1">
			<type>OTHERS_CITABLE</type>
			
					<languageVersion language="fa">
						<title>استخراج و ترکیب ویژگیهای کارآمد از توالی پروتئین به منظور دسته‌بندی پروتئین بر اساس جنگل چرخش</title>
						<abstract>چکیده
پیش&#8204;&#8206;بینی عملکرد پروتئین یکی از چالش&#8204;های اصلی در بیوانفورماتیک می&#8204;باشد که کاربردهای زیادی دارد. در سال&#173;های اخیر در تحقیقات بسیاری از روش&#8204;های یادگیری ماشین در این زمینه استفاده شده&#8204;اند. در این روش&#8204;ها ابتدا باید از توالی پروتئین ویژگی&#8204;های مختلف استخراج شود و بر اساس ویژگی&#8204;های استخراج شده عمل دسته&#8204;بندی انجام شود. غالبا روش&#8204;های استخراج ویژگی بر اساس خصوصیات فیزیکی و شیمایی توالی پروئتین می&#8204;باشد. بنابراین استخراج ویژگی&#8204;هایی مناسب از توالی پروتئین باعث افزایش و بهبود عملکرد روش&#8204;های یادگیری ماشین می&#8204;شود. در این مقاله، یک مجموعه جدید از ویژگی&#173;ها بر اساس روش&#173;های PSSM،&#160; PsePSSM، K-gram ، AAC و روش نوین TFCRF که تا کنون در این کاربرد استفاده نشده برای استخراج ویژگی&#8204;های مناسب پیشنهاد شده است. ویژگی&#8204;های استخراج شده با استفاده از این روش قدرت تمایز کنندگی خوبی بین داده&#8204;ها در دسته&#8204;ها، به مدل&#8204;های یادگیری ماشین می&#8204;دهد. در روش TFCRF وزن&#8204;دهی ویژگی&#8204;ها علاوه بر توجه به چگونگی توزیع آنها در توالی&#8204;های مختلف به چگونگی توزیع آنها در طبقات مختلف نیز توجه می&#8204;شود. در مرحله بعد با استفاده از ویژگی&#8204;های استخراج شده با استفاده از روش جنگل چرخ عمل دسته&#8204;بندی انجام می&#8204;شود. روش پیشنهادی با دسته&#8204;بند&#8204;های مختلف و روش&#8204;های متفاوت مقایسه شده است. نتایج حاصل نشان دهنده کارایی مناسب روش پیشنهادی نسبت به سایر روش&#8204;های نوین در این کاربرد &#160;می&#8204;باشد.</abstract>
						<pdfFileUrl>http://jsdp.rcisp.ac.ir/article-1-1387-fa.pdf</pdfFileUrl>
						<publicationDate>2024-08-03</publicationDate>
						<pageFrom>15</pageFrom>
						<pageTo>26</pageTo>
				
							<doi>10.61186/jsdp.21.1.15</doi>
						<keywords>
<keyword>توالی پروتئین</keyword>
<keyword>استخراج ویژگی</keyword>
<keyword>TFCRF</keyword>
<keyword>جنگل چرخش</keyword>
<keyword>فاکتور ارتباط</keyword>
</keywords>
				</languageVersion>
				

					<languageVersion language="en">
						<title>Extracting and combination efficient feature from protein sequence for classify protein based on rotation forest</title>
						<abstract>Abstract
Protein function prediction is one of the main challenges in bioinformatics, which has many applications. In recent years, many researches in this field have been used machine learning methods. In these methods, First, different features should be extracted from the protein sequence and classification should be done based on the extracted features. The feature extraction methods are based on the physical and chemical properties of the protein sequence. Therefore, extracting suitable features from protein sequence increases and improves the performance of machine learning methods. In this paper, usage of a new set of features based on Position-Specific Scoring Matrix (PSSM), Pseudo-Position Specific Scoring Matrix (PsePSSM), K-gram, Amino Acid Composition (AAC) and the new Term Frequency and Category Relevancy Factor (TFCRF) method, which has not been used in this application so far, is proposed to extract suitable features.
In the PSSM method for protein BLAST searches, a scoring matrix is used, in which amino acid substitution scores are given separately for each position in a multi-sequence protein alignment. The PsePSSM feature is described by considering different ranking correlation factors along a protein sequenc to preserve information about the amino acid sequence. The normalized occurrence frequency of a certain number of amino acids in the protein is calculated by the ACC method. An K-gram is a set of K successive items in a protein that&#160; include amino acid. 
In the TFCRF weighting method, in addition to paying attention to how these are distributed in different sequences, how these are distributed in different classes is also paid attention to.The features extracted using this method give machine learning models a good discriminating power between data in classes. In the next step, classification is done using the extracted features using the rotation forest method. This classifier is a successful ensemble method for a wide range of data mining applications. In this method, the feature space is changed through Principal Component Analysis (PCA), which increases the power of this classifier. The proposed method has been compared to different classifiers. The results show that the efficiency of the proposed method is much better than other state-of&#8211;the-art methods in this application.</abstract>
						<pdfFileUrl>http://jsdp.rcisp.ac.ir/article-1-1387-en.pdf</pdfFileUrl>
						<publicationDate>2024-08-03</publicationDate>
						<pageFrom>15</pageFrom>
						<pageTo>26</pageTo>
				
							<doi>10.61186/jsdp.21.1.15</doi>
						<keywords>
<keyword>Protein sequence</keyword>
<keyword>feature extraction</keyword>
<keyword>TFCRF</keyword>
<keyword>rotation forest</keyword>
<keyword>relevancy factor</keyword>
</keywords>
				</languageVersion>
				


	<authors>
	<author>
	<name>jamshid</name>
	<surname>pirgazi</surname>
	<email>j.pirgazi@mazust.ac.ir</email>
	     <order>1</order>
        <instituteAffiliation>Faculty of Electrical and Computer Engineering, University of Science and Technology of Mazandaran, Behshahr, Iran</instituteAffiliation>  
	    <role>AUTHOR</role>
	 </author>
	<author>
	<name>Ali</name>
	<surname>Ghanbari sorkhi</surname>
	<email>ali.ghanbari@mazust.ac.ir</email>
	     <order>2</order>
        <instituteAffiliation>Faculty of Electrical and Computer Engineering, University of Science and Technology of Mazandaran, Behshahr, Iran</instituteAffiliation>  
	    <role>AUTHOR</role>
	 </author>
	<author>
	<name>majid</name>
	<surname>Iranpour Mobarakeh</surname>
	<email>iranpour@pnu.ac.ir</email>
	     <order>3</order>
        <instituteAffiliation>Department of Computer Engineering and IT, Payam Noor University, Tehran, Iran</instituteAffiliation>  
	    <role>AUTHOR</role>
	 </author>
	</authors>


	</article>



			<article externalId="A-10-2375-1">
			<type>OTHERS_CITABLE</type>
			
					<languageVersion language="fa">
						<title>تشخیص حالت غیر نرمال ماشین های دوار با داده کاوی در پارامترهای حفاظتی</title>
						<abstract>برای محافظت از ماشین&#173; های دوار و جلوگیری از کارکرد آنها در حالت&#173;های غیر عادی به &#173;صورت سنتی از سیستم&#173;های کنترل حفاظتی و داده&#173; های فرایندی بهره&#173; گیری می&#173; شود. در این مقاله روشی پیشنهاد شده است که بتوان از تاثیرات غیر مستقیم حالت&#173;های کارکرد غیر عادی با استفاده از شیوه&#173; های داده &#173;کاوی حالت غیر طبیعی کارکرد ماشین &#173;های دوار را تشخیص داد. یکی از حالت &#173;های خطرناک کارکرد غیر عادی در کمپرسورها به&#173; عنوان یکی از ماشین&#173; های دوار با اهمیت در صنایع، وضعیت سرج می&#173; باشد. دراین مقاله، با استفاده از داده های واقعی ذخیره شده در طول سه سال متوالی یک کمپرسور سه مرحله ای واحد سرمایش یک پالایشگاه گاز ارتباط میان وضعیت سرج کمپرسور و میزان لرزش نقاط مختلف آن بررسی شده است. با شیوه &#173;های داده&#173; کاوی اثبات شده است که ارتباط مستقیمی بین حالت سرج و میزان لرزش وجود دارد. همچنین نقاط حساس&#173; تر به لرزش در زمان&#173;های سرج شناسایی شده است و اثبات شده است که از طریق اندازه&#173; گیری این نقاط می&#173;توان سرج را تشخیص داد. بنابراین &#160;علاوه بر شیوه&#173; های موجود و سنتی قبلی که از داده&#173; های فرایندی استفاده می&#173; کنند، می&#173;توان از میزان لرزش نقاط به&#173; عنوان یک سیستم حفاظتی افزونه برای تشخیص سرج &#160;بهره گرفت و از این طریق حفاظت بیشتری از کمپرسور در برابر وضعیت سرج &#160;بعمل آورد. در این مطالعه، ارزیابی شیوه &#173;های مختلف داده&#173; کاوی نیز صورت گرفته است که نتایج روش نزدیکترین همسایه با تعداد همسایه دو دارای بهترین کارایی بوده است و همچنین اثرات تعداد رکورد موجود در مجموعه داده روی کیفیت و دقت نتایج بررسی شده است.</abstract>
						<pdfFileUrl>http://jsdp.rcisp.ac.ir/article-1-1351-fa.pdf</pdfFileUrl>
						<publicationDate>2024-08-03</publicationDate>
						<pageFrom>27</pageFrom>
						<pageTo>38</pageTo>
				
							<doi>10.61186/jsdp.21.1.27</doi>
						<keywords>
<keyword>ماشین های دوار</keyword>
<keyword>داده کاوی</keyword>
<keyword>تشخیص سرج</keyword>
<keyword>کمپرسور</keyword>
<keyword>پارامترهای حفاظتی</keyword>
</keywords>
				</languageVersion>
				

					<languageVersion language="en">
						<title>Exploring on rotating machines abnormal state with data mining in protective parameters</title>
						<abstract>In order to protect rotating machines and prevent their operation in unusual situations, protective control systems and process data are traditionally used. In this article, a method has been proposed to detect the indirect effects of abnormal operating modes using data mining methods. One of the dangerous conditions of abnormal operation in compressors, as one of the important rotating machines in industries, is the surge condition. In this article, the real data stored during three years of a three-stage refrigerant compressor in a gas refinery are used. the relationship between the surge state of the compressor and the amount of vibration in its different parts has been investigated. It has been proven with data mining methods that there is a direct relationship between the state of surge and the amount of vibration. Also, more sensitive points to vibration during the surges have been identified and it has been proven that by measuring these points, surges can be detected. Therefore, in addition to the existing and previous traditional methods that use process data, it is possible to use the amount of vibration of the points as an extension protection system for surge detection. in this way, more protection of the compressor against the state of surge can be achieved. In this study, various data mining methods have been evaluated, and the results of the nearest neighbor method with the number of neighbors of two have the best performance, and the effects of the number of records in the data set on the quality and accuracy of the results have been investigated.</abstract>
						<pdfFileUrl>http://jsdp.rcisp.ac.ir/article-1-1351-en.pdf</pdfFileUrl>
						<publicationDate>2024-08-03</publicationDate>
						<pageFrom>27</pageFrom>
						<pageTo>38</pageTo>
				
							<doi>10.61186/jsdp.21.1.27</doi>
						<keywords>
<keyword>Rotating machine</keyword>
<keyword>Datamining</keyword>
<keyword>Surge detection</keyword>
<keyword>Compressor</keyword>
<keyword>Protection parameters</keyword>
</keywords>
				</languageVersion>
				


	<authors>
	<author>
	<name>Elham</name>
	<surname>Parvinnia</surname>
	<email>eparvinnia@gmail.com</email>
	     <order>1</order>
        <instituteAffiliation>Computer engineering department, Shiraz branch. Islamic Azad university</instituteAffiliation>  
	    <role>AUTHOR</role>
	 </author>
	<author>
	<name>mohammad</name>
	<surname>Safari</surname>
	<email>Safari.md@gmail.com</email>
	     <order>2</order>
        <instituteAffiliation>Computer engineering department, Shiraz branch. Islamic Azad university</instituteAffiliation>  
	    <role>AUTHOR</role>
	 </author>
	<author>
	<name>seyedalireza</name>
	<surname>khayami</surname>
	<email>khayami.alr@gmail.com</email>
	     <order>3</order>
        <instituteAffiliation>Shiraz university</instituteAffiliation>  
	    <role>AUTHOR</role>
	 </author>
	</authors>


	</article>



			<article externalId="A-10-815-7">
			<type>OTHERS_CITABLE</type>
			
					<languageVersion language="fa">
						<title>ارائه روشی جدید برای خوشه بندی داده های مخلوط بر مبنای تعداد ویژگی مشابه</title>
						<abstract>خوشه &#173;بندی عملیاتی است که در آن مجموعه&#173;ای از نمونه داده&#8204;ها، نسبت به میزان شباهت، دسته&#173;بندی می&#173;شوند. نمونه داده&#173;های خوشه&#173;بندی، عددی یا مخلوطی از عددی و غیرعددی (اسمی) هستند. یافتن میزان شباهت و اندازه&#8204;گیری فاصله، از چالش&#173;های خوشه&#173;بندی داده &#173;های مخلوط است. در این مقاله سعی شده است در محاسبه میزان شباهت و تعیین فاصله، به پارامتر &#34;تعداد ویژگی&#8204;های مشابه&#34; توجه شود. در نسبت دادن هر نمونه به خوشه در مواردی که فاصله&#8204;ها برابر یا نزدیک باشد، تعداد ویژگی&#8204;های مشترک نمونه&#8204;ها تعیین کننده خوشه مناسب خواهد بود. برای محاسبه فاصله در الگوریتم مورد نظر از تفاضل عددی نرمالسازی شده برای ویژگی&#8204;های عددی و از فاصله همینگ برای ویژگی&#8204;های غیرعددی استفاده شده است. تعیین مرکز خوشه اولیه نیز مانند بسیاری از روش&#8204;ها بصورت تصادفی انجام شده است و در تکرارهای بعدی الگوریتم، نمونه مناسب&#8204;تر به عنوان مرکز خوشه انتخاب می&#8204;شود. الگوریتم مورد نظر با 5 الگوریتم دیگر در 5 مجموعه&#8204; داده مقایسه شده است. در بررسی نتایج، از سه معیارAccuracy ، RI، F-Measure &#160;استفاده شده است. طبق نتایج آزمایشات، در سه مجموعه&#8204;داده، الگوریتم موردنظر حداقل دو درصد بهتر از دو الگوریتم و یک درصد بهتر از یکی دیگر از الگوریتم&#8204;ها عمل کرده است. در یکی دیگر از مجموعه&#8204;داده&#8204;ها الگوریتم موردنظر نتایج برابر یا نزدیک به یک درصد دقت بهتر نسبت به الگوریتم برتر داشت. در مجموعه&#8204;داده آخر نیز الگوریتم مورد نظر در رتبه دوم از بین 5 الگوریتم قرار داشت.</abstract>
						<pdfFileUrl>http://jsdp.rcisp.ac.ir/article-1-1329-fa.pdf</pdfFileUrl>
						<publicationDate>2024-08-03</publicationDate>
						<pageFrom>39</pageFrom>
						<pageTo>52</pageTo>
				
							<doi>10.61186/jsdp.21.1.39</doi>
						<keywords>
<keyword>خوشه‌بندی</keyword>
<keyword>داده مخلوط</keyword>
<keyword>فاصله مقادیر</keyword>
<keyword>تشابه مقادیر</keyword>
<keyword>مرکز خوشه.</keyword>
</keywords>
				</languageVersion>
				

					<languageVersion language="en">
						<title>Presenting a new method for mixed data clustering based on the number of similar features</title>
						<abstract>Clustering is an operation in which a set of data samples is categorized according to the degree of similarity. Examples of clustering data are numerical or a mixture of numerical and non-numerical (nominal) data. Finding similarities and measuring distances is one of the challenges of mixed data clustering. In the related works, to detect the degree of similarity and obtain the distance value, only the parameter of the distance value was considered and the cluster was selected based on its value. Clustering in this way, especially for mixed data, has not had very accurate results.
In this paper, we have tried to pay attention to the parameter &#34;number of similar features&#34; in calculating the degree of similarity and determining the distance. In assigning each sample to a cluster in cases where the distances are equal or close, the number of common features of the samples will determine the appropriate cluster. That is, we will pay attention to the &#34;number of similar features&#34; in addition to the distance to select the cluster. This idea believes that in cases where the distance of the cluster centers is close to the data object, it is better to choose the cluster center that has more features similar to the data object. Logically and also according to the proposed algorithm, the amount of similarity should be in a larger number of features, not just a few limited features but with high similarity.
The parameter of the &#34;number of similar features&#34; has a specific definition and is obtained with a suitable threshold. If the distance value of two features is less than the threshold, those two features are considered as similar features.
To calculate the distance in the algorithm, the normalized numerical difference for numerical properties and the Hamming distance for non-numerical properties are used. Determining the initial cluster centers, like many methods, is done randomly, and in subsequent iterations of the algorithm, more appropriate samples are selected as the cluster centers. The algorithm is compared with 5 other algorithms in 5 datasets. 
In examining the results, three criteria of Accuracy, RI and F-Measure have been used. According to the test results, in the mixed and integer datasets, the algorithm performs at least two percent better than the two algorithms and one percent better than the other algorithm. In another data set, the proposed algorithm had results equal to or close to one percent better accuracy than the superior algorithm. In the last data set, the proposed algorithm was ranked second among 5 algorithms. In general, the proposed algorithm won the top rank in most of the results, and in the rest of the cases, it won the second rank out of the five tested algorithms.</abstract>
						<pdfFileUrl>http://jsdp.rcisp.ac.ir/article-1-1329-en.pdf</pdfFileUrl>
						<publicationDate>2024-08-03</publicationDate>
						<pageFrom>39</pageFrom>
						<pageTo>52</pageTo>
				
							<doi>10.61186/jsdp.21.1.39</doi>
						<keywords>
<keyword>Clustering</keyword>
<keyword>Mixed data</keyword>
<keyword>Distance of values</keyword>
<keyword>Similarity of values</keyword>
<keyword>Cluster Center.</keyword>
</keywords>
				</languageVersion>
				


	<authors>
	<author>
	<name></name>
	<surname></surname>
	<email>hid.rezaei@nit.ac.ir</email>
	     <order>1</order>
        <instituteAffiliation></instituteAffiliation>  
	    <role>AUTHOR</role>
	 </author>
	<author>
	<name>Negin</name>
	<surname>Daneshpour</surname>
	<email>ndaneshpour@sru.ac.ir</email>
	     <order>2</order>
        <instituteAffiliation></instituteAffiliation>  
	    <role>AUTHOR</role>
	 </author>
	</authors>


	</article>



			<article externalId="A-10-2408-1">
			<type>OTHERS_CITABLE</type>
			
					<languageVersion language="fa">
						<title>یک چارچوب توزیعی دو مرحله‌ای مبتنی بر خوشه‌بندی برای شناسایی چهره در مقیاس بالا</title>
						<abstract>در زمینه&#8204;ی شناسایی چهره، چالش&#8204;های افت دقت، افزایش نیازمندی به حافظه، و افزایش پیچیدگی زمانی از مشکلات مهمی به شمار می&#8204;آیند. به منظور حل این مسائل، این تحقیق یک رویکرد دومرحله&#8204;ای سه&#8204;واحدی معرفی کرده است: واحد زیرشبکه&#8204;ها، واحد خوشه&#8204;یاب، و واحد تصمیم&#8204;گیر نهایی. در مقابل روش&#8204;های مبتنی بر توزیع تصادفی، روش ارائه شده، از خوشه&#8204;بندی به عنوان روش توزیع مسئله به زیرشبکه&#8204;ها استفاده می&#8204;کند. هر زیرشبکه&#8204;، یک شبکه عصبی عمیق نظارتی است که با داده&#8204;های آموزشی مربوط به دسته&#8204;های خود آموزش می&#8204;بیند. واحد خوشه&#8204;یاب، شباهت بردارهای ویژگی داده&#8204;های آزمون را با میانگین بردارهای ویژگی دسته&#8204;ها مقایسه می&#8204;کند و بهترین خوشه را پیدا می&#8204;کند. در نهایت، واحد تصمیم&#8204;گیر نهایی با ترکیب نتایج دو واحد قبلی، بهترین دسته را انتخاب می&#8204;کند. نتایج نشان می&#8204;دهد که روش پیشنهادی، در مقایسه با روش&#8204;های مشابه، از نظر صحت، بازخوانی، و امتیاز F1 عملکرد بهتری دارد. این روش ضمن سریع&#8204;تر بودن، دارای دقت بالاتری نسبت به روش&#8204;های بدون توزیع می&#8204;باشد و در مقایسه با روش&#8204;های توزیعی تصادفی، سرعتی برابر و دقتی بالاتر دارد. آزمایش&#8204;ها بر روی مجموعه&#8204;دادگان VGGFace2 و &#160;MS-Celeb-1M و Glint360K اجرا شده و نشان می&#8204;دهد که این روش علاوه بر عملکرد بهتر، مقیاس&#8204;پذیری بالاتری را در بازشناسی چهره دارد.</abstract>
						<pdfFileUrl>http://jsdp.rcisp.ac.ir/article-1-1362-fa.pdf</pdfFileUrl>
						<publicationDate>2024-08-03</publicationDate>
						<pageFrom>53</pageFrom>
						<pageTo>70</pageTo>
				
							<doi>10.61186/jsdp.21.1.53</doi>
						<keywords>
<keyword>بازشناسی چهره</keyword>
<keyword>شناسایی چهره</keyword>
<keyword>خوشه‌بندی</keyword>
<keyword>یادگیری عمیق</keyword>
<keyword>یادگیری توزیعی</keyword>
</keywords>
				</languageVersion>
				

					<languageVersion language="en">
						<title>A two-stage clustering-based distributed framework for large-scale face identification</title>
						<abstract>Face recognition poses challenges in accuracy, memory efficiency, and computational complexity. This study proposes a two-stage, three-module approach: Subnetwork modules, Cluster-Finder unit, and Final-Decision module. Unlike random distribution methods, our approach employs clustering for distribution. Each subnetwork, a supervised deep neural network, is trained with cluster-specific data. The Cluster-Finder unit compares test data similarity with each subnetwork&#8217;s representative. The Final-Decision module selects the best class. Results indicate superior accuracy, recall, and F1 score compared to competitive methods. The approach is faster and more accurate than non-distribution methods, with comparable speed and higher accuracy than random distribution methods. Experiments on VGGFace2, MS-Celeb-1M, and Glint360K datasets confirm both superior performance and scalability. The proposed method, using KMeans for distribution, outperforms Softmax Dissection and Dynamic Active Class Selection. It simplifies training without additional manipulations, offering efficiency over methodologies like Softmax Dissection and ArcFace parallelization. In conclusion, this study focuses on pre-processing and post-processing without added training complexity. A divide-and-conquer approach addresses accuracy and efficiency challenges. In this study, various sources leading to errors in face recognition systems have been examined. These sources include: imprecise features, overfitting, challenging classes, distribution issues, and decision-making complexities. Various classification scenarios are explored, including non-distributed and models with random and intelligent distributions. Inaccurate features uniformly impact all scenarios, with overfitting posing the greatest challenge in non-distributed scenarios. Challenging classes are better distinguished in intelligent distribution scenarios. Inappropriate distribution has less impact in intelligent scenarios, and decision-making challenges exist in both distributions</abstract>
						<pdfFileUrl>http://jsdp.rcisp.ac.ir/article-1-1362-en.pdf</pdfFileUrl>
						<publicationDate>2024-08-03</publicationDate>
						<pageFrom>53</pageFrom>
						<pageTo>70</pageTo>
				
							<doi>10.61186/jsdp.21.1.53</doi>
						<keywords>
<keyword>face recognition</keyword>
<keyword>face identification</keyword>
<keyword>clustering</keyword>
<keyword>deep learning</keyword>
<keyword>distributed learning.</keyword>
</keywords>
				</languageVersion>
				


	<authors>
	<author>
	<name>Sayed Mohammad</name>
	<surname>Ahmadi</surname>
	<email>sm.ahmadi@stu.qom.ac.ir</email>
	     <order>1</order>
        <instituteAffiliation>University of Qom</instituteAffiliation>  
	    <role>AUTHOR</role>
	 </author>
	<author>
	<name>Rouhollah</name>
	<surname>Dianat</surname>
	<email>rdianat@qom.ac.ir</email>
	     <order>2</order>
        <instituteAffiliation>University of Qom</instituteAffiliation>  
	    <role>AUTHOR</role>
	 </author>
	</authors>


	</article>



			<article externalId="A-10-2170-1">
			<type>OTHERS_CITABLE</type>
			
					<languageVersion language="fa">
						<title>الگوریتمی مبتنی بر گراف برای خوشه‌بندی سوره‌های قرآن کریم</title>
						<abstract>قرآن کتاب نازل&#8204;شده از طرف خداست و تا به امروز اندیشمندان و پژوهش&#8204;گران مختلفی در جهت شناخت قرآن و فهم آن تلاش نموده&#8204;اند. در دسترس بودن سیستم&#8204;های رایانه&#8204;ای فرصت مغتنمی است که با افزایش سرعت پژوهش&#8204;گران در پیمودن مسیر، آن&#8204;ها را در رسیدن به قله&#8204;های بلندتری یاری کند. خوشه&#8204;بندی یکی از روش&#8204;هایی است که برای فهم ساختار داده به کار می&#8204;رود. در این مقاله به خوشه&#8204;بندی سوره&#8204;های قرآن کریم بر اساس هم&#8204;وقوعی کلمات در آن پرداخته&#8204; و برای دست&#8204;یابی به این هدف از یک رویکرد موجود مبتنی بر گراف استفاده نموده&#8204;ایم. در پژوهش جاری ابتدا هر سوره را به صورت یک گراف غیرجهت&#8204;دار و وزن&#8204;دار بازنمایی کرده&#8204;، سپس بردار هر سوره را بر اساس گراف سوره&#8204; تشکیل داده&#8204;ایم و پس از آن سوره&#8204;ها را خوشه&#8204;بندی نموده&#8204;ایم. برای ارزیابی کیفیت خوشه&#8204;بندی از معیار نیم&#8204;رخ استفاده کرده&#8204;ایم. بر اساس این معیار در بهترین خوشه&#8204;بندی در بین اجراهای مختلف مقدار نیم&#8204;رخ ۰/۹۱ به دست&#8204;آمده است. این پژوهش زیرساخت ساختاری مناسبی برای توصیف لایه معنایی سوره&#8204;&#8204;ها و آیات قران پیش روی پژوهش&#8204;گران حوزه زبان&#8204;شناسی محاسباتی در دامنه علوم قرآنی فراهم می&#8204;سازد.</abstract>
						<pdfFileUrl>http://jsdp.rcisp.ac.ir/article-1-1220-fa.pdf</pdfFileUrl>
						<publicationDate>2024-08-03</publicationDate>
						<pageFrom>71</pageFrom>
						<pageTo>88</pageTo>
				
							<doi>10.61186/jsdp.21.1.71</doi>
						<keywords>
<keyword>خوشه‌بندی متن</keyword>
<keyword>بازنمایی شبکه‌ای متن</keyword>
<keyword>گراف متن</keyword>
<keyword>زیرگراف‌ پرتکرار</keyword>
<keyword>قرآن‌کاوی رایانشی</keyword>
</keywords>
				</languageVersion>
				

					<languageVersion language="en">
						<title>A Graph-based Algorithm for Clustering Qur’anic Surahs</title>
						<abstract>The Holy Qur&#39;an is revealed from God Almighty. Up to now many scholars and researchers have tried to understand the Holy Qur&#39;an and comprehend it. The availability of computer systems is a great opportunity to help researchers reach higher peaks by speeding them up in their way. Clustering is one of the methods has been used to understand the structure of the data. In clustering, we want to divide samples of data into groups so that the members of each cluster are similar together and are different from the members of the other clusters. Clustering of Qur&#39;anic surahs has been the subject of some computer studies on the Qur&#39;an. In these studies, different approaches have been considered to vectorizing the surahs. In a study, Thabet formed vectors of each surah by considering some stems of Qur&#39;anic words as features and the normalized probability of their occurrences in the surah as feature values and clustered just 24 surahs due to the sparseness of the obtained data matrix. With a similar approach in vectorizing the surahs, Moisl calculated the minimum surah length threshold per feature in order to solve the problem of shorter surahs by using some concepts of statistical sampling theory, and could cluster more surahs. Instead of using words as features, Sharaf considered 13 features including existence of referring to the story of Adam and Ebliys, number of the phrase &#171;یا أَیُّهَا الَّذینَ آمَنُوا&#187; (O you who believe), and determined the method of measuring each feature. Then, he formed data matrix and clustered the Qur&#39;anic surahs. In another study, Sufi et al. considered the topics identified for each verse in the Tafsir Rahnama as features and constructed a binary data matrix based on the presence or absence of that topic in the Tafsir of that surah and applied clustering. In this article, we have clustered the surahs of the Holy Qur&#39;an based on the co-occurrence of words in it. To achieve this goal, we have used an existing graph-based approach. In the present study, we first represent each surah as a weighted undirected graph. Then we form the vector of each surah by considering closed frequent sub-graphs as features and relative occurrence of them in each surah as feature values, and eventually cluster the surahs. We used the Silhouette score to evaluate the quality of clustering. Based on this criterion, in the best clustering among different runs, the Silhouette score of 0.91 was obtained. This research provide a proper structural infrastructure for specifying the semantic layer of Holy Qur&#39;an surahs for computational linguistics researchers in the domain of Qur&#39;anic studies.</abstract>
						<pdfFileUrl>http://jsdp.rcisp.ac.ir/article-1-1220-en.pdf</pdfFileUrl>
						<publicationDate>2024-08-03</publicationDate>
						<pageFrom>71</pageFrom>
						<pageTo>88</pageTo>
				
							<doi>10.61186/jsdp.21.1.71</doi>
						<keywords>
<keyword>Document Clustering</keyword>
<keyword>Text Graph</keyword>
<keyword>Frequent subgraph</keyword>
<keyword>Computational Qur'an mining</keyword>
</keywords>
				</languageVersion>
				


	<authors>
	<author>
	<name>Maryam Sadat</name>
	<surname>Mottaghi</surname>
	<email>m.motaghi88@chmail.ir</email>
	     <order>1</order>
        <instituteAffiliation>Shahid Beheshti university</instituteAffiliation>  
	    <role>AUTHOR</role>
	 </author>
	<author>
	<name>Behrouz</name>
	<surname>Minaei-Bidgoli</surname>
	<email>b_minaei@iust.ac.ir</email>
	     <order>2</order>
        <instituteAffiliation>University of Science and Technology of Iran</instituteAffiliation>  
	    <role>AUTHOR</role>
	 </author>
	</authors>


	</article>



			<article externalId="A-10-2042-1">
			<type>OTHERS_CITABLE</type>
			
					<languageVersion language="fa">
						<title>طراحی شبکه بهینه‌ساز خودکار الهام گرفته از الگوریتم بهینه‌سازی BFGS با حافظه محدود</title>
						<abstract>امروزه علی&#8204;رغم توسعه مدل&#8204;های یادگیری ماشین برای استخراج ویژگی&#8204;ها به صورت خودکار، هنوز الگوریتم&#8204;های بهینه&#8204;سازی به صورت دستی طراحی می&#8204;شوند. یکی از اهداف فرایادگیری (meta-learning)، خودکار کردن فرایند بهینه&#8204;سازی است. الگوریتم&#8204;های بهینه&#8204;سازی دستی مبتنی بر بردار گرادیان تنها براساس عملیات ضرب داخلی، ضرب اسکالر و جمع برداری بر روی بردارهای ورودی نوشته می شوند. بنابراین می توان گفت که این الگوریتمها در فضای هیلبرت بعد مساله بهینه&#8204;سازی اجرا می شوند. ما نیز قصد داریم با ایده گرفتن از این مطلب، فضایی برای یادگیری ورودی&#8204;ها ایجاد کنیم که مستقل از ابعاد ورودی باشد. بدین منظور با ایده گرفتن از الگوریتم BFGS با حافظه محدود (L-BFGS) و همچنین سلول LSTM یک ساختار جدید با نام Hilbert LSTM (HLSTM) معرفی می&#8204;کنیم که فرایند یادگیری در آن مستقل از ابعاد ورودی انجام می&#8204;شود. به عبارتی الگوریتم یادگیری در فضای هیلبرت مساله بهینه&#8204;سازی اجرا می&#8204;شود. برای رسیدن به این هدف از لایه ضرایب خطی استفاده می&#8204;کنیم که ترکیب خطی بردارهای ورودی را محاسبه می&#8204;کند و ضرایب این ترکیب خطی، با کمک ضرب داخلی بردارهای ورودی بدست می&#8204;آید. آزمایش&#8204;های ما نشان می&#8204;دهند که نتایج به&#8204;دست آمده توسط بهینه&#8204;ساز ارائه شده، به مراتب بهتر از نتایج الگوریتم&#8204;های بهینه&#8204;سازی دستی است.</abstract>
						<pdfFileUrl>http://jsdp.rcisp.ac.ir/article-1-1142-fa.pdf</pdfFileUrl>
						<publicationDate>2024-08-03</publicationDate>
						<pageFrom>89</pageFrom>
						<pageTo>100</pageTo>
				
							<doi>10.61186/jsdp.21.1.89</doi>
						<keywords>
<keyword>Hilbert LSTM</keyword>
<keyword>LSTM</keyword>
<keyword>L-BFGS</keyword>
<keyword>فرایادگیری</keyword>
<keyword>بهینه‌سازی خودکار</keyword>
</keywords>
				</languageVersion>
				

					<languageVersion language="en">
						<title>Designing L-BFGS inspired automatic optimizer network</title>
						<abstract>Nowadays using features learned by machines is common and these types of features have excellent quality in comparison with hand-designed features. While many machine learning models are developed to extract features automatically, however, the optimizing algorithms are still designed manually. In this paper, we propose a method to cast the optimizing algorithm as a machine learning problem. This is a branch of machine learning which is named meta-learning or learning to learn.
Gradient-based optimization algorithms (e.g. gradient descent and BFGS) receive the gradient vector in each step and, by using the information of the previous points and gradients, estimate the update vector at the current point. The inputs and outputs of these algorithms are vectors whose dimension is the same as the optimization problem. These algorithms are written solely based on vector addition, scalar-product, and inner-product operations. Therefore, we can say that these algorithms are executed in a Hilbert space whose dimension is determined by the optimization problem. In this paper, we propose a novel method for learning to optimize over a Hilbert space of unknown dimensionality.
We introduce a new neural network module named Hilbert LSTM (HLSTM) which is based on a novel LSTM cell whose learning process is independent of the input data dimension. This independency is the result of restricting the network to the operations on a Hilbert space, prohibiting the network to work directly with the entries within a vector. To achieve this goal, we use a linear coefficients layer that linearly combines the input vectors based on coefficients computed by their inner products. Training the network based on the inner product between vectors leads to learning an optimization algorithm that is independent of the data dimension. Our experiments show that the proposed optimizer achieves better results in comparison with hand-designed algorithms.</abstract>
						<pdfFileUrl>http://jsdp.rcisp.ac.ir/article-1-1142-en.pdf</pdfFileUrl>
						<publicationDate>2024-08-03</publicationDate>
						<pageFrom>89</pageFrom>
						<pageTo>100</pageTo>
				
							<doi>10.61186/jsdp.21.1.89</doi>
						<keywords>
<keyword>Hilbert LSTM</keyword>
<keyword>LSTM</keyword>
<keyword>L-BFGS</keyword>
<keyword>meta-learning</keyword>
<keyword>automatic optimization</keyword>
</keywords>
				</languageVersion>
				


	<authors>
	<author>
	<name>Mohammad</name>
	<surname>Etesam</surname>
	<email>etesam@mail.um.ac.ir</email>
	     <order>1</order>
        <instituteAffiliation>Ferdowsi University of Mashhad</instituteAffiliation>  
	    <role>AUTHOR</role>
	 </author>
	<author>
	<name>Ashkan</name>
	<surname>Sadeghi-Lotfabadi</surname>
	<email>sadeghia@mail.um.ac.ir</email>
	     <order>2</order>
        <instituteAffiliation>Ferdowsi University of Mashhad</instituteAffiliation>  
	    <role>AUTHOR</role>
	 </author>
	<author>
	<name>Kamaledin</name>
	<surname>Ghiasi-Shirazi</surname>
	<email>k.ghiasi@um.ac.ir</email>
	     <order>3</order>
        <instituteAffiliation>Ferdowsi University of Mashhad</instituteAffiliation>  
	    <role>AUTHOR</role>
	 </author>
	</authors>


	</article>



			<article externalId="A-10-92-1">
			<type>OTHERS_CITABLE</type>
			
					<languageVersion language="fa">
						<title>نور-استم نسخه 1. یک مجموعه داده معیار برای ارزیابی ریشه‌یاب‌های عربی</title>
						<abstract>ریشه&#8204;یابی مرحله اصلی چندین فرایند پردازشی مانند متن&#8204;کاوی، بازیابی اطلاعات و پردازش زبان طبیعی است. ابزارهای تشخیص میانوند کلمات عربی با چالش&#8204;های زیادی روبرو هستند که بیشتر ناشی از ماهیت پیچیده کلمات این زبان و سبک&#173;های نوشتاری متفاوت آن&#173;ها است. تا جایی که ما می&#173;دانیم، هیچ مجموعه داده&#173;ی ریشه&#173;یابیِ معیاری وجود ندارد که طیف گسترده&#173;ای از چالش&#173;های ریشه&#173;یابی را پوشش دهد. بنابراین، ما توسعه یک مجموعه داده برای ارزیابی پایداری ریشه&#8204;یاب&#8204;ها را در چنین موقعیت&#173;های چالش برانگیزی ارزشمند می&#173;دانیم. این مقاله، نور-استم، یک مجموعه داده معیار با سبک&#8204;های نوشتاری مختلف را برای ارزیابی ابزارهای تشخیص میانوند (استم) عربی معرفی &#8204;می&#8204;کند. جهت تایید عملکرد این دادگان، عملکرد سه ریشه&#8204;یاب&#8204; عربی (نور ۱۰، NLTK و تاشفین) مورد ارزیابی قرار گرفته است. نتایج نشان می&#173;دهد که سنجه&#173;ی اف در ریشه&#173;یاب تاشفین بهتر از سایر ریشه&#8204;یاب&#8204;ها است که این موضوع در تحقیقات مرتبط نیز مشاهده شده است.</abstract>
						<pdfFileUrl>http://jsdp.rcisp.ac.ir/article-1-1346-fa.pdf</pdfFileUrl>
						<publicationDate>2024-08-03</publicationDate>
						<pageFrom>101</pageFrom>
						<pageTo>112</pageTo>
				
							<doi>10.61186/jsdp.21.1.101</doi>
						<keywords>
<keyword>دادگان معیار</keyword>
<keyword>ریشه‌یاب</keyword>
<keyword>نور-استم</keyword>
<keyword>میانوند</keyword>
<keyword>استخراج اطلاعات</keyword>
</keywords>
				</languageVersion>
				

					<languageVersion language="en">
						<title>Noor-stem v.1 A Benchmark Dataset for Evaluating the Arabic Stemmers</title>
						<abstract>The main task of the tokenization is to divide the sentences of the text into its constituent units and remove punctuation marks (dots, commas, etc.). Each unit is a continuous lexical or grammatical writing chain that is an independent semantic unit. Tokenization occurs at the word level and the extracted units can be used as input to other components such as stemmer. Stemming is the main step of several processing tasks such as text mining, information retrieval, and natural language processing.Arabic stemmers face many challenges, mostly caused by the complex nature of Arabic words and their different writing styles. To our knowledge, there is no gold stemming dataset, which contains a wide variety of different possible stemming challenges, so that, stemmers face numerous and different possible real-world challenges to stem the words. Thus, we find it valuable to develop a dataset for evaluating the sustainability of stemmers in such a variety of challenging situations. In this paper, we introduce Noor-Stem, a benchmark dataset with various writing styles for the evaluation of Arabic stemmers. We use two thousand Arabic words in this dataset. We choose the words from different sources such as holy Quran as well as the Arabic websites and assign them to two groups of human experts to determine the correct stem for each word. The first chosen collection of words includes non-repetitive words of the Quran according to their morphological structure. This collection, with more than 16,000 words, is completely by its Quranic usage, labeling only the words stems. The necessity of morphological analysis in Quranic texts as an example of the index of classical Arabic texts has given rise to this evaluation. The second word collection includes 10 thousand words from the non-repetitive words of the text data in general classic Arabic texts. Out of more than 2,600,000 non-repetitive words, considering that the dataset is going to be gold and each stem must be labeled/ensured by a couple of experts, 10,000 words are chosen, regarding the comprehensive and unique patterns to fully measure the length. The variety of patterns can face each stemmer with a serious challenge to demonstrate its performance in various processes. We evaluate the performance of three Arabic stemmers (Light 10, NLTK and Tashaphyne) on this dataset. The results show that the F-measure of Tashaphyne is better than the other stemmers, which re-proves the superiority of this stemmer in this type of problem, as well.</abstract>
						<pdfFileUrl>http://jsdp.rcisp.ac.ir/article-1-1346-en.pdf</pdfFileUrl>
						<publicationDate>2024-08-03</publicationDate>
						<pageFrom>101</pageFrom>
						<pageTo>112</pageTo>
				
							<doi>10.61186/jsdp.21.1.101</doi>
						<keywords>
<keyword>Benchmark Dataset</keyword>
<keyword>Stemmer</keyword>
<keyword>Noor-Stem</keyword>
<keyword>Infix</keyword>
<keyword>Information Retrieval</keyword>
</keywords>
				</languageVersion>
				


	<authors>
	<author>
	<name>Azal</name>
	<surname>Al-Aswad</surname>
	<email>azal.alamery2@gmail.com</email>
	     <order>1</order>
        <instituteAffiliation></instituteAffiliation>  
	    <role>AUTHOR</role>
	 </author>
	<author>
	<name>Behrouz</name>
	<surname>Minaei-Bidgoli</surname>
	<email>b_minaei@iust.ac.ir</email>
	     <order>2</order>
        <instituteAffiliation></instituteAffiliation>  
	    <role>AUTHOR</role>
	 </author>
	<author>
	<name>Mohammad-Ebrahim</name>
	<surname>Shenassa</surname>
	<email>me.shenasa@iau-tnb.ac.ir</email>
	     <order>3</order>
        <instituteAffiliation></instituteAffiliation>  
	    <role>AUTHOR</role>
	 </author>
	<author>
	<name>Sayyed-Ali</name>
	<surname>Hossayni</surname>
	<email>sayyed.hossayni@yandex.com</email>
	     <order>4</order>
        <instituteAffiliation></instituteAffiliation>  
	    <role>AUTHOR</role>
	 </author>
	<author>
	<name>Habib</name>
	<surname>Seryani</surname>
	<email>hseryani@noornet.net</email>
	     <order>5</order>
        <instituteAffiliation></instituteAffiliation>  
	    <role>AUTHOR</role>
	 </author>
	</authors>


	</article>



			<article externalId="A-10-1041-2">
			<type>OTHERS_CITABLE</type>
			
					<languageVersion language="fa">
						<title>بهبود پالایش مشارکتی در سیستم‌های توصیه‌گر با کمک خوشه‌بندی فازی C– میانگین مرتب‌شده و الگوریتم ازدحام ذرات تطبیقی – آشوبی</title>
						<abstract>سیستم&#8204;های توصیه&#8204;گر زیرمجموعه&#8204;ای از سیستم&#8204;های هوشمند پالایش اطلاعات هستند که در فضای اینترنت علایق کاربر را شناسایی نموده و توصیه&#8204;های مرتبط با سلیقه&#8204;ی کاربر را ارائه می&#8204;دهند. پالایش مشارکتی مبتنی بر کاربر، از مهم&#8204;ترین انواع سیستم&#8204;های توصیه&#8204;گر است. از مهم&#8204;ترین چالش&#8204;ها در این سیستم&#8204;ها پراکندگی و حجم زیاد داده&#8204;ها است که بر کارایی آن&#8204;ها اثرگذار است. در روش پیشنهادی، برای اولین بار از الگوریتم خوشه&#8204;بندی فازی C-میانگین مرتب&#8204;شده و الگوریتم&#8204; تکاملی ازدحام ذرات تطبیقی آشوبی برای خوشه&#8204;بندی کاربران استفاده&#8204; شده &#8204;است. هدف روش پیشنهادی بهبود میزان خطای پیش&#8204;بینی در مجموعه داده&#8204;های حجیم با پراکندگی زیاد و کاهش تأثیر داده های پرت و نویز است. به منظور ارزیابی و اثبات کارایی روش پیشنهادی، آزمایش&#8204;هایی روی پایگاه داده&#8204;های واقعی اجرا شده&#8204; است. نتایج آزمایش&#8204;ها نشان&#8204;دهنده&#8204;ی برتری روش پیشنهادی نسبت به روش&#8204;های مرز دانش بر اساس معیارهای میانگین خطای مطلق، جذر میانگین مربعات خطا، نرخ صحت و زمان محاسباتی است.
&#160;</abstract>
						<pdfFileUrl>http://jsdp.rcisp.ac.ir/article-1-1129-fa.pdf</pdfFileUrl>
						<publicationDate>2024-08-03</publicationDate>
						<pageFrom>113</pageFrom>
						<pageTo>124</pageTo>
				
							<doi>10.61186/jsdp.21.1.113</doi>
						<keywords>
<keyword>سیستم توصیه‌گر</keyword>
<keyword>پالایش مشارکتی</keyword>
<keyword>خوشه‌بندی فازی</keyword>
<keyword>الگوریتم تکاملی</keyword>
<keyword>الگوریتم ازدحام ذرات تطبیقی آشوبی</keyword>
</keywords>
				</languageVersion>
				

					<languageVersion language="en">
						<title>Improving Collaborative Recommender Systems by Integrating Fuzzy C-Ordered Means Clustering and Chaotic Self-Adaptive Particle Swarm Optimization Algorithm</title>
						<abstract>Recommender systems are a subset of intelligent information filtering systems that discovers user interests and provide user-friendly recommendations. User-based collaborative filtering recommender systems is one of the most important types of recommender systems. However, they are faced with voluminous data and sparsity problems that have negative effects on the performance of the systems. In the proposed method, fuzzy C-ordered means clustering algorithm is integrated with a chaotic self-adaptive particle swarm evolutionary algorithm for clustering users. The proposed method aims to improve the rating prediction in large sparse datasets and reduce the negative impact of outliers and noisy data. Experiments have been conducted on real-world datasets to evaluate and prove the efficiency of the proposed method. Experimental results show the superiority of the proposed method that the state-of-the-art methods based on prediction error criteria, accuracy rates, and the computational time.</abstract>
						<pdfFileUrl>http://jsdp.rcisp.ac.ir/article-1-1129-en.pdf</pdfFileUrl>
						<publicationDate>2024-08-03</publicationDate>
						<pageFrom>113</pageFrom>
						<pageTo>124</pageTo>
				
							<doi>10.61186/jsdp.21.1.113</doi>
						<keywords>
<keyword>Recommender systems</keyword>
<keyword>Collaborative filtering</keyword>
<keyword>Fuzzy clustering</keyword>
<keyword>Evolutionary algorithm</keyword>
<keyword>Chaotic self-adaptive particle swarm optimization algorithm.</keyword>
</keywords>
				</languageVersion>
				


	<authors>
	<author>
	<name>Javad</name>
	<surname>Hamidzadeh</surname>
	<email>J_Hamidzadeh@sadjad.ac.ir</email>
	     <order>1</order>
        <instituteAffiliation>Sadjad University of Technology</instituteAffiliation>  
	    <role>AUTHOR</role>
	 </author>
	<author>
	<name>Mona</name>
	<surname>Moradi</surname>
	<email>Mmoradi@semnan.ac.ir</email>
	     <order>2</order>
        <instituteAffiliation>Semnan University</instituteAffiliation>  
	    <role>AUTHOR</role>
	 </author>
	</authors>


	</article>



			<article externalId="A-10-2377-1">
			<type>OTHERS_CITABLE</type>
			
					<languageVersion language="fa">
						<title>بهبود قدرت تعمیم مدل‌های تشخیص کلام نفرت‌انگیز مبتنی بر تطبیق دامنه</title>
						<abstract>امروزه با رشد فعالیت در شبکه&#8204;های اجتماعی شاهد افزایش کلام نفرت&#8204;انگیز به صورت برخط هستیم و به&#8204;همین منظور مسئلۀ تشخیص نفرت در فضای مجازی دارای اهمیت است. همچنین تطبیق دامنه نیز در این مسئله و به&#8204;طورکلی در حوزۀ پردازش زبان طبیعی، یکی از چالش&#8204;های مهم است. در بسیاری از مسائل، ضمن تغییر دامنه با افت عملکرد مواجهیم که این موضوع در مسئلۀ نفرت نیز صادق است. در این پژوهش با استفاده از روش&#8204;های تطبیق دامنه سعی در افزایش قدرت تعمیم&#8204;پذیری مدل&#8204;های تشخیص نفرت خواهیم داشت. برای این منظور روش&#8204;های مبتنی بر ترنسفورمر شامل آموزش خصمانۀ دامنه و ترکیب متخصصان را به کار می&#8204;گیریم و همچنین از آموزش چند منبعی استفاده می&#8204;کنیم. آزمایش&#8204;ها با استفاده از چهار مجموعه&#8204;داده در حوزۀ نفرت انجام می&#8204;شوند. در ابتدا مد&#8204;ل&#8204;ها را به&#8204;صورت درون&#8204; دامنه&#8204;ای و تک منبعی ارزیابی می&#8204;کنیم. در مرحلۀ بعد با اضافه کردن دامنه&#8204;های دیگر به بخش آموزش، شاهد افت نتایج و انتقال منفی هستیم. سپس آزمایش&#8204;های برون دامنه&#8204;ای را ابتدا به&#8204;صورت تک منبعی با مدل DistilBERT انجام می&#8204;دهیم که با تغییر دامنه نتایج به طور قابل توجهی کاهش می&#8204;یابند. به&#8204;منظور افزایش قدرت تطبیق دامنۀ مدل&#8204; در بخش برون دامنه&#8204;ای، روی چند منبع آموزش را انجام می&#8204;دهیم که حدوداً در نیمی از موارد سبب بهبود نتایج می&#8204;شود که نتیجۀ معناداری نیست. در ادامه با استفاده از روش&#8204;های مبتنی بر ترنسفورمر شامل آموزش خصمانۀ دامنه و ترکیب متخصصان سعی در افزایش قدرت تطبیق دامنۀ مدل&#8204;ها خواهیم داشت که در 87% از آزمایش&#8204;های برون دامنه&#8204;ای چند منبعی شاهد افزایش عملکرد هستیم. البته این روش&#8204;ها در عملکرد آزمایش&#8204;های درون دامنه&#8204;ای هم مؤثر هستند. مسئلۀ مهمی که گاهی موجب افت&#8204;وخیز چشمگیر نتایج می&#8204;شود، مجموعه&#8204;داده&#8204;ها هستند. شباهت داده&#8204;ها و تشابه توزیع بعضی دامنه&#8204;ها باعث افزایش قدرت تطبیق دامنۀ مدل می&#8204;شوند.
&#160;</abstract>
						<pdfFileUrl>http://jsdp.rcisp.ac.ir/article-1-1341-fa.pdf</pdfFileUrl>
						<publicationDate>2024-08-03</publicationDate>
						<pageFrom>125</pageFrom>
						<pageTo>142</pageTo>
				
							<doi>10.61186/jsdp.21.1.125</doi>
						<keywords>
<keyword>کلام نفرت‌انگیز</keyword>
<keyword>تطبیق دامنه</keyword>
<keyword>تعمیم</keyword>
<keyword>طبقه‌بندی</keyword>
<keyword>ترنسفورمر</keyword>
</keywords>
				</languageVersion>
				

					<languageVersion language="en">
						<title>Domain adaptation-based method for improving generalization of hate speech detection models</title>
						<abstract>Today, with the growth of activity in social media, we see an increase in hate speech online and for this reason, the issue of recognizing hate in cyberspace is important. Also, domain adaptation is one of the important challenges in this task and in general in the field of natural language processing. In many issues, while changing the domain, we face a drop in performance, which is also true in the task hate speech. In this research, we try to increase the generalizability of hate detection models by using domain adaptation methods. For this purpose, we use Transformer-based methods, including domain adversarial training and mixture of experts, and we also use multi-source training. Experiments are conducted using four datasets in the domain of hate. At first, we evaluate the models in an in-domain and single-source manner. In the next step, by adding other domains to the education section, we see a drop in results and a negative transfer. Then we perform the out-of-domain tests first as a single source with the DistilBERT model, which significantly reduces the results by changing the domain. In order to increase the power of domain adaptation of the model in the out-of-domain part, we perform the training on several sources, leads to improve the results in about half of the cases, which is not significant. In the following, we try to increase the domain adaptation power of the models, using transformer-based methods including domain adversarial training and the mixture of experts, which leads to increase in performance in 87% of multi-source out-of-domain tests. Of course, these methods are also effective in the performance of in-domain tests. An important issue that sometimes causes a significant drop in results is datasets. The similarity of the data and the similarity of the distribution of some domains increase the power of domain adaptation of the model and on the contrary.</abstract>
						<pdfFileUrl>http://jsdp.rcisp.ac.ir/article-1-1341-en.pdf</pdfFileUrl>
						<publicationDate>2024-08-03</publicationDate>
						<pageFrom>125</pageFrom>
						<pageTo>142</pageTo>
				
							<doi>10.61186/jsdp.21.1.125</doi>
						<keywords>
<keyword>hate speech</keyword>
<keyword>classification</keyword>
<keyword>transformer</keyword>
<keyword>domain adaptation</keyword>
<keyword>generalization</keyword>
</keywords>
				</languageVersion>
				


	<authors>
	<author>
	<name>Seyedeh Fatemeh</name>
	<surname>Nourollahi</surname>
	<email>sfn1373@gmail.com</email>
	     <order>1</order>
        <instituteAffiliation>Qom University</instituteAffiliation>  
	    <role>AUTHOR</role>
	 </author>
	<author>
	<name>Razieh</name>
	<surname>Baradaran</surname>
	<email>r.baradaran@stu.qom.ac.ir</email>
	     <order>2</order>
        <instituteAffiliation>Qom University</instituteAffiliation>  
	    <role>AUTHOR</role>
	 </author>
	<author>
	<name>Hossein</name>
	<surname>Amirkhani</surname>
	<email>amirkhani@qom.ac.ir</email>
	     <order>3</order>
        <instituteAffiliation>Qom University</instituteAffiliation>  
	    <role>AUTHOR</role>
	 </author>
	</authors>


	</article>


	</issue>
 </ici-import>
 
  
  
  
  
 