Updating in doctor's office times: computer prototype aims to reduce breast cancer research search times
| Jaime Andrés Hurtado Giraldo, Master's student in Engineering with Emphasis in Systems Engineering. Credit: Édgar Bejarano, Communications Office, Faculty of Engineering. |
The production of scientific literature related to breast cancer, both in the identification of treatments and potential cures, is significantly extensive around the world. However, the pressing pace with which physicians coexist during their shifts in clinics and hospitals prevents a thorough review of these advances. One research project aims to use artificial intelligence and complex network systems to reduce search times and provide a more schematic understanding of current publications, nationally and internationally.
Lea el artículo en español aquí.
Too much information at hand, too little time to take ownership of it
The vast production of scientific research in areas of interest, such as breast cancer, brings with it a difficult paradox that often confronts those who are called to be at the forefront of the best treatments. It is common that doctors do not have enough time, immersed in the dynamics of clinics and hospitals, to evaluate the results of such research. This situation implies a lag in praxis, which can impede the use of novel treatments and improvements in the lives of patients undergoing surgery.
Aware of this paradox, which accuses health professionals in their access to available information, chemical engineer Jaime Andrés Hurtado Giraldo, taking advantage of his experience in automation of industrial processes, decided to direct his research project, within the framework of the Master's Degree in Engineering with Emphasis in Systems Engineering, towards an innovative solution.
"The problem is that there is a lot of scientific literature and it is very difficult to be constantly updated with new publications, due to the large volume that is handled. We try to prioritize analytical methods in order to find solutions to this problem," says researcher Hurtado Giraldo.
His research project, which is directed by professor and researcher Oswaldo Solarte Pabón and co-directed by professor and researcher Víctor Andrés Bucheli Guerrero, from the same academic unit, seeks to offer an automated tool that operates through principles of computing and local or cloud storage, in addition to artificial intelligence and complex network systems, so that doctors and researchers can perform more dynamic searches, whose impact can be evidenced when offering specific treatments for each case, in line with what is being done worldwide.
Research: prototype information search and complex networks
Researcher Jaime Andrés Hurtado Giraldo conducted an exploration to find scientific literature aggregators with access via application programming interface (API) and related software libraries, through which cancer publications can be accessed. The aggregators chosen were PubMed, ArXiv and Core. In the case of PubMed, the international reference database for biomedical literature, the link was achieved through the MetaPub library, a Python library designed to facilitate interaction between researchers, system developers and PubMed. After this, the research focused on automating the downloading of scientific publications, access to metadata and the construction of networks of authors, institutions, countries and keywords.
To achieve this purpose, the development of the information search prototype was started, whose main characteristics are to be open source and to work through microservices. These microservices operate independently, as if they were several computers within the same device, which allows the application to be much more efficient and scalable. "The microservices architecture has allowed us to group various software implementations into a single system, facilitating the integration of their functionalities and allowing both the optimization of resources and the possibility of installing the prototype on the user's personal computer," explains researcher Hurtado Giraldo.
“Based on keywords, a search is conducted for scientific publications that are relevant to each researcher's topic of interest. We took certain keywords extracted from pathological records related to breast cancer in the city of Cali. In a previous work, we reviewed some pathological records of several medical institutions in the city, collecting words related to drug names, tumor names, alternatives to what can be called breast cancer and some findings made by doctors. Subsequently, using a microservices-based architecture, the scientific literature was accessed, processed and analyzed in an automated manner," adds the researcher. For the subject breast cancer, the keyword list included about 232 records, which served as a starting point for the article search.
The developed prototype can be deployed both in a local environment (for example, on the doctor's computer) and in the cloud. The system is configured to perform continuous integration and deployment in production using the Google Cloud, allowing it to be accessed from any location. For use, each microservice can be queried independently, and additionally a web interface has been developed that integrates these microservices, depending on the user's needs. "For example, there is a microservice that is responsible for connecting to databases and downloading scientific articles, another that processes the information by accessing the texts to search for authors, institutions, locations and keywords. Another microservice stores the information obtained in databases, there is one that is responsible for building networks and their metrics with the extracted information, and finally, there is the microservice that provides the web interface for configuration, consultation and visualization," says the researcher.
The visualization in the web interface of the data obtained helps to locate connections that go beyond publications, giving the user the opportunity to find communicating vessels between researchers, the institutions and countries to which they belong, highlighting data of interest thanks to the use of complex network metrics. "There are mathematical formulas which allow us to calculate the interactions that each node has (formed by the results of a publication found in the aggregators). The idea is that, as there are so many connections in the networks, one can visualize nodes in a certain place in the interface and highlight them depending on their importance, based on complex network metrics," explains researcher Jaime Andrés Hurtado Giraldo.
With these complex network metrics, the interface user can visualize the level of interaction of the nodes, depending on each case. "Less dense networks indicate that authors do not have as much research. Apart from the authors, we also present the graphs of institutions and countries, which are created from the information of the e-mail domains and locations detected in the text," adds researcher Hurtado Giraldo.
![]() |
| Information nodes which link connections between authors for each published research. Credit: researcher's courtesy. |
Patient health outcomes and impact
The research currently allows us to see practical results through the web interface developed. "We have an implementation that allows us to deploy the application in the cloud or locally, and this, in an automated way, from the keywords that are already ready, connects to the databases, opens each PDF that downloads (where the publication is contained, if it is open access) and searches within the PDF emails, to create networks of researchers through these emails. A similar process is done with institutions, locations and keywords. We have initial tests that have allowed us to obtain networks with around 1,500 authors, and in addition to these authors, we have also detected 800 organizations related to the subject of breast cancer", explains the researcher.
![]() |
| Information related to the author's institutions, derived from each article found. Credit: researcher's courtesy. |
The implementation of this interface seeks to minimize search time for researchers in the health field, specifically in the area of breast cancer. Thanks to its microservices technology, it is open so that other scientific literature aggregators can be added, as deemed appropriate, to enrich the scope of the research within the prototype.
The search prototype developed will soon undergo a verification process by medical personnel in a health care institution in the city of Cali, so that this will be the laboratory that will test the benefits of this technology applied to the area of breast cancer and its patients. "We are used to doctors keeping very busy, and they, based on their experience and knowledge, try to formulate medications and treatments. It is hoped that soon an automated tool will help them to visualize new alternatives," concludes the researcher, and clarifies that this is a tool whose potential goes beyond the area of medicine:
"It's software that would also be useful for other types of research. For anyone who doesn't have time to do extensive research, this interface could save a lot of time. I believe that computer tools have come to help us with the reduction of work time in certain areas."
If interested in being in touch with the Master's student or any further information about the investigation, please write the Faculty of Engineering Communications Office: comunicaingenieria@correounivalle.edu.co.


Comentarios
Publicar un comentario