Description
DESCRIPTION
At SerdatIA, we are looking for a Junior Data \& AI Engineer to join our Data \& AI team.
This Pull Request represents the onboarding of a new collaborator who will help us design, build, and evolve scalable data platforms and intelligent solutions.
If you enjoy solving complex problems, working with modern technologies, and transforming data into real business value, this PR may be for you.
Pull Request Description
Why are we opening this PR?
Our team continues to grow, and we are seeking an engineer capable of contributing to real-world Big Data and Artificial Intelligence projects.
You will work alongside engineers, architects, and data specialists to build solutions based on cloud architectures, Big Data processing, and Generative AI.
Included Changes
Data Engineering
Building data pipelines from extraction and ingestion through processing and availability
Developing ETL / ELT processes
Integrating data from multiple structured, semi-structured, and unstructured sources
Hands-on experience with Batch and Streaming or Near Real-Time (NRT) processing
Big Data processing with real-world projects involving large volumes of data
Optimizing processes and Spark jobs
Fundamental principles of Lakehouse architectures
Modeling and transforming data for various business use cases
Practical application of best practices for data quality, governance, and traceability
Monitoring and resolving incidents of various natures
Cloud Engineering
Developing cloud solutions with hands-on production experience
Designing architectures using cloud services from any major cloud provider (Azure, AWS, GCP)
Automating production deployment processes using DevOps tools and CI/CD
Some experience in FinOps tasks for performance and cost optimization
Configuring and optimizing storage, processing, and compute services
Monitoring and troubleshooting
Certifications at Engineer level or higher in Azure, AWS, and/or GCP are highly valued.
Artificial Intelligence
Designing and implementing LLM-based solutions
Developing business-oriented Generative AI applications
Building RAG architectures
Integrating language models with enterprise knowledge sources
Working with vector databases
Integrating AI APIs and Cloud AI services
Designing efficient prompts
Implementing solutions using frameworks such as LangChain, LangGraph, Pipecat, etc.
Participating in industrialization and deployment of AI solutions in production environments
Applying best practices for AI security, governance, and responsible usage
Repository Structure
bigdata\-ai\-engineer/
programming\-languages/
Java
Python
Scala
data\-platform/
pipelines
spark\-jobs
databricks
delta\-lake
sql
artificial\-intelligence/
llm\-applications
rag\-solutions
ai\-agents
vector\-search
prompt\-engineering
cloud/
microsoft\-azure
amazon\-web\-services
google\-cloud\-platform
engineering/
git
docker
ci\-cd
azure\-devops
clean\-code
testing
documentation
Acceptance Criteria
This PR will be approved if:
Technical Skills
* \*\*Programming fundamentals\*\*: Strong understanding of OOP and structured/functional paradigms, data structures and algorithms (lists, dictionaries, recursion, Big O), collaborative Git usage (commits, branches, PRs), and clean code (DRY, short functions).
* \*\*AI-powered coding assistants\*\*: Daily and efficient use of tools like GitHub Copilot, Claude Code, or Gemini CLI to support development, testing, and documentation—while always retaining independent technical judgment.
* \*\*Application and API development\*\*: Consuming and designing simple REST endpoints (HTTP verbs, JSON, token/API key authentication), relational databases (basic SQL: SELECT, JOIN, INSERT/UPDATE), basic NoSQL concepts, unit testing, and debugging via hypothesis formulation.
* \*\*Architecture and deployment (introductory level)\*\*: Conceptual understanding of microservices, using Docker and docker-compose for local environments, understanding and modifying CI/CD pipelines, and awareness of message queues (Kafka, RabbitMQ).
* \*\*Data handling and hygiene\*\*: Manipulating CSV, JSON, and Parquet data in Python (pandas); detecting nulls, duplicates, and inconsistencies; understanding data quality principles (\*garbage in, garbage out\*), privacy (PII), and basic descriptive statistics.
* \*\*Machine Learning and Deep Learning (conceptual foundation)\*\*: Understanding of supervised/unsupervised learning, overfitting, intuition about neural networks and Transformers; introductory experience with libraries such as scikit-learn, PyTorch, or TensorFlow.
* \*\*Consumption and integration of Generative AI\*\*: Calling and integrating LLM APIs (OpenAI, Anthropic, Azure AI) from Python/JavaScript, managing context (\*context engineering\*), tokens, temperature, and securely handling API keys via environment variables.
* \*\*Generative AI architectures\*\*: Practical understanding of RAG (\*Retrieval-Augmented Generation\*) patterns, generating and using embeddings, vector databases, agents with tool calling, and familiarity with the MCP (\*Model Context Protocol\*) standard.
Data \& AI Mindset
* \*\*Critical thinking toward AI\*\*: Actively verifying outputs, detecting hallucinations, and understanding risks and permissions when allowing agents to execute actions.
* \*\*Data culture and quality\*\*: Awareness that clean, well-governed data forms the foundation of any model’s or AI solution’s performance.
* \*\*Ethics and privacy\*\*: Rigor in protecting sensitive data and adhering to security best practices.
* \*\*Technical judgment regarding tools\*\*: Sensitivity to differences in capabilities, latency, and costs among various AI providers and models.
Team Collaboration
* \*\*Curiosity and continuous self-learning\*\*: Proactivity in exploring new tools and libraries within the fast-evolving AI ecosystem and sharing knowledge.
* \*\*Communication and transparency\*\*: Ability to ask timely questions, document solutions, and give/receive constructive feedback.
* \*\*Team integration\*\*: Ease in closely collaborating with senior profiles and data specialists, learning from their experience and contributing meaningfully to the team.
What can we offer you?
Full-time position
Permanent contract
Professional career development, aligned with your contributions
100% remote position
Additionally, if you join Grupo Seresco, you’ll enjoy benefits such as:
22 working days of vacation plus 2 flexible days and either Christmas Eve or New Year’s Eve off.
Flexible working hours—we enjoy three months of continuous morning shifts.
Exclusive discounts across numerous services and products (travel, automotive, fashion, leisure, sports, culture, home, technology, gastronomy‿ via the Corporate Benefits platform.
At SERESCO, you can configure how to allocate your gross salary (Serflex program), using part of it flexibly. You may choose among the following: Medical Insurance (Sanitas, IMQ, or Caser), including family members; Gourmet Voucher/Card; Transport Card; and Daycare Voucher.
We also offer a highly attractive incentive plan if you help us find professionals like yourself.
Since we assist in identifying/inventing new projects, there is also a business acquisition reward program.
Study support fund for formal education.
In-person and online training, enabling you to stay current at your own pace.
REQUIREMENTS
What are we looking for?
* Degree/Master’s in Computer Engineering, Telecommunications, Mathematics, Data Science, or equivalent technical education.
* Solid programming foundation in Python (and/or JavaScript/TypeScript) and data structures.
* Regular use of Git for version control and team collaboration.
* Experience consuming language model APIs (OpenAI, Anthropic, Azure OpenAI, etc.) from your own code.
* Conceptual and practical understanding of RAG architectures, embeddings, and AI agents.
* Knowledge of SQL and data manipulation using libraries such as Pandas.
* Familiarity with Docker and containers for local development.
* Technical judgment and critical thinking to evaluate and validate AI model outputs.
Preferred qualifications:
* Hands-on experience or personal/academic projects with AI frameworks (LangChain, LlamaIndex, LangGraph, etc.).
* Knowledge and usage of vector databases (Chroma, Pinecone, Qdrant, pgvector, etc.).
* Familiarity with emerging standards such as Model Context Protocol (MCP).
* Knowledge or certifications in cloud platforms (Azure, AWS, GCP).
* Experience with CI/CD pipelines (Azure DevOps, GitHub Actions).