Skip to content
15% Off Your Second Order · Minimum Order £50 15% Off Second Order · Minimum £50

Final Year Projects for Computer Science and Engineering Students: 17 Scoped Ideas

Seventeen final year projects for computer science and engineering students, each with a public dataset or an open specification, a baseline to beat, a stack, a stretch goal and a source we opened. Twelve are built on datasets and reference implementations; five are anchored to a dated standard or regulation, from post-quantum cryptography to WCAG 2.2.

This is a list of final year projects for computer science and engineering students, and every idea on it comes with the thing a supervisor asks for first: a dataset you can download or a specification you can implement, a baseline to compare against, a stack, a stretch goal and a source. There are seventeen. Twelve are built on public datasets and open reference implementations; five, in the last section, are tied to a standard or a regulation with a date attached, so the reason the work matters can be cited in your introduction.

The final year project is the piece of work an interviewer will ask you about, so choose something you can explain in two minutes as well as build. Treat each entry below as a scope statement rather than a specification: a project with four features delivered and evaluated beats one with eight half-built. The next two sections set out how to pick an idea and how to narrow it; the lists follow.

Which Final Year Project Is Best for Computer Science?

The one you can finish, demonstrate and explain. A working system with a modest scope and a measured evaluation scores better than an ambitious idea that runs out of time. Choose a problem you can get data for, a method your supervisor can assess, and a scale that fits the weeks you actually have.

That rule rules out more ideas than it sounds like it does. "Quantum cryptography" is a research area, not a project; "implement and benchmark ML-KEM key exchange in an existing chat application" is a project. The difference is that the second one has a deliverable, a measurement and an end.

How Do You Choose a Final Year Project Topic?

Start from three constraints. What data or hardware can you actually get hold of? What can your supervisor supervise? How many weeks do you have once teaching and exams are removed? Pick a topic that satisfies all three and still interests you, then write a one-page scope with a minimum deliverable and a stretch goal.

How to turn an idea into a final year project scope

  1. Find the data or hardware first A public dataset with a license and a citation beats a private one you hope to be given.
  2. Name the evaluation and the baseline Decide how you will show the thing works, and against what, before you write any code.
  3. Narrow the problem to a live area Something currently changing gives you recent literature to cite and a motivation paragraph that writes itself.
  4. Check the ethics route Human participants, personal data or scraping need approval, and approval takes weeks.
  5. Write a one-page scope A minimum deliverable you can demonstrate, and a stretch goal you add only if the weeks allow.
Do the steps in this order; a project that starts with the code and looks for the data later is the one that runs out of weeks.
  • Data first. If the project needs data, find the dataset before you commit to the idea. A public dataset with a license and a citation beats a private one you hope to be given.
  • Name the evaluation on day one. Decide how you will show the thing works, and against what baseline, before you write any code. Reports lose marks in the evaluation chapter more often than in the implementation chapter.
  • Prefer a narrow problem in a live area. Working on something that is currently changing, such as post-quantum migration or AI transparency, gives you recent literature to cite and a clear motivation paragraph.
  • Check the ethics route early. Anything involving human participants, personal data or scraping needs approval, and that approval takes weeks.

Our computer science assignment help page explains how we work on projects and reports, and there are worked examples in the computer science assignment samples archive. If your project is dissertation-shaped, read how to write a dissertation proposal first. If your project is engineering rather than software, the shear force and bending moment diagram worked example shows how we set out a calculation.

Projects with a Public Dataset and a Baseline

Each of these eight names the data, the number to beat and the smallest version worth demonstrating. The dataset links are the pages we opened; the license and the citation each one asks for are on those pages, and your report should quote them. Where a published figure is given, it is the baseline your evaluation chapter compares against.

  1. Credit card fraud detection on an imbalanced dataset.
    • Objective: Detect the 492 fraudulent transactions among 284,807 card payments, and show what the class imbalance does to every metric you report.
    • Data and baseline: The ULB Machine Learning Group's dataset on Kaggle holds two days of September 2013 transactions by European cardholders, with 28 principal components plus Time and Amount; frauds are 0.172% of the rows, and the dataset page recommends area under the precision-recall curve rather than accuracy (Credit Card Fraud Detection, Kaggle). Baseline: logistic regression on the raw features with a time-ordered split, reported as AUPRC.
    • Stack: Python, pandas, scikit-learn, imbalanced-learn.
    • Stretch goal: Compare class weights, undersampling and an isolation forest on the same split, and report precision at a fixed recall for each, so a reader can see the trade a bank would actually make.
  2. Sentiment classification on the IMDB reviews.
    • Objective: Classify film reviews as positive or negative and beat a published number.
    • Data and baseline: The Large Movie Review Dataset from Stanford has 25,000 reviews for training and 25,000 for testing, labeled only when strongly positive (7 out of 10 or above) or strongly negative (4 or below) (Large Movie Review Dataset). The paper that introduced it reported 88.89% accuracy, and that is your number to beat (Maas et al., 2011).
    • Stack: Python, scikit-learn for a TF-IDF and logistic regression baseline, PyTorch or Hugging Face Transformers for the model that beats it.
    • Stretch goal: Fine-tune a small pretrained transformer, then report the accuracy gained against the training time and memory it cost, and test both models on a handful of sarcastic reviews you write yourself.
  3. Intrusion detection on CIC-IDS2017.
    • Objective: Train a classifier that separates attack flows from benign traffic, then test whether it recognizes an attack family it was not trained on. For the same problem inside a car, our report on connected-vehicle security reviews a study that trains classifiers to detect communication attacks on connected and autonomous vehicles.
    • Data and baseline: The Canadian Institute for Cybersecurity's dataset covers five days, 3 to 7 July 2017: Monday benign traffic only, Tuesday FTP and SSH brute force, Wednesday denial of service and Heartbleed, Thursday web attacks and infiltration, Friday botnet, port scanning and DDoS, with more than 80 flow features extracted by CICFlowMeter (CIC-IDS2017). The dataset asks to be cited through Sharafaldin, Habibi Lashkari and Ghorbani (2018). Baseline: a random forest on the flow features, trained and tested within one day.
    • Stack: Python, pandas, scikit-learn; Wireshark or CICFlowMeter if you want to capture a flow of your own for the demonstration.
    • Stretch goal: Train on Wednesday's denial-of-service traffic and test on Friday's DDoS. The drop in recall, and your explanation of it, is the contribution.
  4. Remaining useful life on the NASA turbofan data.
    • Objective: Predict how many operating cycles an engine has left from its sensor channels, and say how far ahead the prediction can be trusted.
    • Data and baseline: NASA's Prognostics Data Repository publishes the C-MAPSS turbofan degradation simulation: four sets run under different combinations of operating conditions and fault modes, each recording several sensor channels as the fault develops (NASA Prognostics Data Repository; Saxena and Goebel, 2008). Baseline: a gradient-boosted regressor on windowed sensor features, reported as root mean squared error on the simplest of the four sets.
    • Stack: Python, pandas, scikit-learn or LightGBM; PyTorch for the sequence model.
    • Stretch goal: An LSTM over the same windows, and a second comparison on the set with several fault modes, which is where the two models are likely to separate.
  5. Activity recognition from smartphone sensors.
    • Objective: Recognize six activities (walking, walking upstairs, walking downstairs, sitting, standing, lying) from accelerometer and gyroscope signals.
    • Data and baseline: The UCI Human Activity Recognition dataset has 10,299 instances with 561 features from 30 volunteers aged 19 to 48, recorded at 50 Hz on a waist-mounted Samsung Galaxy S II; 70% of the volunteers form the training set and 30% the test set, so the split is by person, not by row (UCI Machine Learning Repository; CC BY 4.0; Anguita et al., 2013). Baseline: a linear support vector machine on the 561 published features.
    • Stack: Python, scikit-learn; PyTorch for a one-dimensional convolutional model on the raw signals.
    • Stretch goal: Record a few minutes of the same six activities from three consenting classmates on a current phone at 50 Hz and test whether the trained model transfers. This needs the ethics approval mentioned above, so start the form in week one.
  6. Real-time object detection with an edge deployment.
    • Objective: Run a detector on live video from a webcam or a small board, and report the accuracy given up for each step of speed gained.
    • Data and baseline: Microsoft COCO, introduced in 2014 with 328,000 images and 2.5 million labeled instances across 91 object types (Lin et al., 2014; cocodataset.org). Baseline: a pretrained detector's reported COCO mean average precision, plus your own measurement of its frames per second on the target device.
    • Stack: PyTorch, a pretrained YOLO-family or Detectron2 model, OpenCV; a Raspberry Pi or Jetson board as the edge target.
    • Stretch goal: Quantize and prune the model, then plot mean average precision against frames per second for every variant on the same hardware. That curve is the deliverable.
  7. Traffic forecasting on a sensor network.
    • Objective: Predict traffic speed at each sensor 15, 30 and 60 minutes ahead from the network's recent history. Our report on smart city transportation shows how six UK cities use traffic and sensor data, which gives the introduction a real-world setting.
    • Data and baseline: The DCRNN repository (Li, Yu, Shahabi and Liu, 2018) publishes the Los Angeles (metr-la.h5) and Bay Area (pems-bay.h5) files at five-minute steps together with its own results: a mean absolute error of 2.67, 3.08 and 3.56 on METR-LA at the three horizons (DCRNN on GitHub). Baseline: a historical average per sensor and time of day, then a gradient-boosted model per sensor, both scored with the same metric.
    • Stack: Python, pandas, PyTorch; PyTorch Geometric if the stretch goal is attempted.
    • Stretch goal: Reproduce the published DCRNN figures, or account for the gap between yours and theirs. Either is a result.
  8. Federated learning on a realistic non-IID split.
    • Objective: Train a model across simulated clients without centralizing their data, and measure what that costs in accuracy and communication rounds.
    • Data and baseline: Federated averaging was introduced by McMahan and colleagues in 2016 and remains the baseline every later method is compared against (Communication-Efficient Learning of Deep Networks from Decentralized Data). For the data, the UCI activity dataset above splits naturally by volunteer, so thirty people become thirty clients with genuinely different distributions. Baseline: the same model trained centrally on the pooled data.
    • Stack: Python and PyTorch with a hand-written federated averaging loop; the loop is short and teaches more than a framework does.
    • Stretch goal: Accuracy and communication cost against the central model as the clients per round and the local epochs change, with the data deliberately uneven across clients.

Projects Built on a Specification or a Reference Implementation

These four have no dataset. The baseline is a published specification, a test suite or a reference implementation; the evaluation is whether your build meets it and what it costs when it does.

  1. A Raft key-value store.
    • Objective: Build a replicated key-value service that keeps working when a server, the leader included, is killed.
    • Specification and baseline: Raft is described by its authors as a consensus algorithm designed to be easy to understand, equivalent to Paxos in fault tolerance and performance (raft.github.io; Ongaro and Ousterhout, 2014). MIT's 6.5840 course publishes its Spring 2026 Raft lab in four parts, leader election, log replication, persistence and log compaction, each with a test suite you can run (MIT 6.5840, Lab 3). Passing those tests is the baseline.
    • Stack: Go.
    • Stretch goal: A key-value service on top of the log with snapshots, and a measurement of throughput and recovery time while the leader is repeatedly killed.
  2. A collaborative editor on CRDTs.
    • Objective: Let several people edit one document at once, offline included, without a central server deciding the order of edits.
    • Specification and baseline: Yjs is an MIT-licensed CRDT framework that supports offline editing, undo and redo and shared cursors, with y-websocket and y-webrtc as its connection providers (Yjs on GitHub). A working editor on Yjs is the baseline and can be built in a week; the project is what you measure on it.
    • Stack: TypeScript, Yjs, y-websocket, a small Node.js server.
    • Stretch goal: Implement a minimal sequence CRDT of your own, replay the same edit traces through it and through Yjs, and compare convergence time, message size and memory.
  3. Verifiable credentials for a campus.
    • Objective: Issue a credential such as proof of module completion, hold it in a wallet and verify it, with the three roles kept separate. No blockchain is needed, and the report should say why.
    • Specification and baseline: The W3C Verifiable Credentials Data Model v2.0 became a Recommendation on 15 May 2025. It defines a verifiable credential as a specific way to express a set of claims made by an issuer, in a three-party ecosystem of issuers, holders and verifiers, secured either by embedded proofs (Data Integrity) or enveloping proofs (JOSE and COSE) (W3C, Verifiable Credentials Data Model v2.0). Conformance to the data model, checked against the specification's own examples, is the baseline.
    • Stack: Node.js or Python, a JSON signing library, a browser-based wallet.
    • Stretch goal: Selective disclosure, so the holder proves the module was completed without revealing the grade, and a threat model of what the verifier learns.
  4. An air-quality monitor with a calibration study.
    • Objective: Build a low-cost sensor node, then find out how wrong it is and correct it.
    • Data and baseline: Before touching hardware, use the UCI Air Quality dataset: 9,358 hourly readings from five metal-oxide sensors co-located with a certified reference analyzer at road level in an Italian city, March 2004 to February 2005 (UCI Machine Learning Repository; CC BY 4.0; De Vito et al., 2008). Fitting a calibration model on it is the baseline exercise. Then compare your own node against the nearest public monitor through OpenAQ, which aggregates air-quality data from hundreds of sources behind an open API (OpenAQ).
    • Stack: An ESP32 or Raspberry Pi Pico W with a particulate sensor and a temperature and humidity sensor, MQTT, a small time-series database and a dashboard.
    • Stretch goal: A drift model that re-calibrates the node against the reference over several weeks, with the error before and after as the result.

2026 Projects Anchored to a Standard or a Regulation Date

These five are newer than the lists above and each one is anchored to a document you can cite in your introduction: a standard, a specification or a regulation with a date attached. That matters because the first question in a project viva is usually why the work is worth doing, and a published deadline answers it better than an opinion.

  1. Post-quantum migration for an existing application.
    • Objective: Replace the classical key exchange or signature scheme in a working application with a post-quantum algorithm, then measure the cost.
    • Why now: NIST published three post-quantum standards in August 2024, FIPS 203 (ML-KEM, key encapsulation), FIPS 204 (ML-DSA, digital signatures) and FIPS 205 (SLH-DSA, hash-based signatures), and has selected HQC as a further key-encapsulation algorithm for standardization (NIST Post-Quantum Cryptography project).
    • Deliverable and evaluation: A hybrid handshake, plus measurements of handshake size, latency and CPU cost against the classical baseline. The numbers are the contribution.
  2. Content provenance and AI labeling tool.
    • Objective: Attach and verify tamper-evident provenance metadata on images or audio, and detect when it has been stripped.
    • Why now: Most remaining provisions of the EU AI Act apply from 2 August 2026, and providers of systems generating synthetic audio, image, video or text that were already on the market face transparency duties by 2 December 2026 (European Commission; AI Act implementation timeline). The C2PA specification gives you an open format to implement against (C2PA specification 2.1).
    • Deliverable and evaluation: A signing and verification pipeline, plus a test set showing which common edits preserve or break the credential.
  3. A Model Context Protocol server for a campus system.
    • Objective: Expose a real data source, such as a library catalog or a timetable, to an AI assistant through a standard protocol rather than a bespoke integration.
    • Why now: The Model Context Protocol is an open protocol with a published specification (revision 2025-06-18) that defines resources, prompts and tools as the units a server offers, and sets out consent and tool-safety requirements (MCP specification).
    • Deliverable and evaluation: A working server, an authorization model, and an evaluation of how a client behaves when the tools are misdescribed. The security section is the interesting part.
  4. An accessibility audit tool for WCAG 2.2.
    • Objective: Automate the checks that can be automated and produce a triaged report for the ones that cannot.
    • Why now: WCAG 2.2 became a W3C Recommendation on 5 October 2023 and was republished as a revised Recommendation on 12 December 2024 (W3C publication history). It adds nine success criteria over WCAG 2.1, including Target Size (Minimum), Dragging Movements, Consistent Help, Redundant Entry and Accessible Authentication (W3C, WCAG 2.2). Several of the new criteria are poorly covered by existing tooling.
    • Deliverable and evaluation: A checker plus a benchmark of its precision and recall against pages audited by hand.
  5. Customer segmentation and recommendation on a public retail dataset.
    • Objective: Cluster customers on recency, frequency and monetary features, then build and evaluate a recommender for each segment.
    • Why now: The Online Retail dataset gives you 541,909 real transactions from a UK online retailer with a customer ID, under a permissive license and with a DOI you can cite (UCI Machine Learning Repository). Having a licensed, citable dataset removes the commonest reason these projects stall.
    • Deliverable and evaluation: A time-based train and test split, precision and recall at k, and a most-popular-item baseline. See our data mining for e-commerce personalization dissertation sample for a full worked version.

Also read how to select a topic for your dissertation and the computer science dissertation sample.

One practical note to close on. A project is marked on what you can demonstrate and explain, not on how ambitious the title sounds. Pick the idea whose data you can obtain, whose method your supervisor can assess, and whose smallest useful version you could finish in half the time you have. The rest is scope you can add if the weeks allow.

Need help scoping, building or writing up a final year computer science project? Message us on WhatsApp with the project brief, the module and the submission date.

Sources

  • ULB Machine Learning Group, Credit Card Fraud Detection [Dataset]. Kaggle. kaggle.com. 284,807 transactions, 492 frauds, September 2013.
  • Maas, A. L., Daly, R. E., Pham, P. T., Huang, D., Ng, A. Y. and Potts, C. (2011) Learning Word Vectors for Sentiment Analysis, Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics. Dataset page: ai.stanford.edu.
  • Sharafaldin, I., Habibi Lashkari, A. and Ghorbani, A. A. (2018) Toward Generating a New Intrusion Detection Dataset and Intrusion Traffic Characterization, 4th International Conference on Information Systems Security and Privacy (ICISSP). Dataset page: unb.ca/cic.
  • Saxena, A. and Goebel, K. (2008) Turbofan Engine Degradation Simulation Data Set, NASA Prognostics Data Repository, NASA Ames Research Center. nasa.gov.
  • Anguita, D., Ghio, A., Oneto, L., Parra, X. and Reyes-Ortiz, J. L. (2013) A Public Domain Dataset for Human Activity Recognition using Smartphones, European Symposium on Artificial Neural Networks. Dataset: UCI Machine Learning Repository, archive.ics.uci.edu, DOI 10.24432/C54S4K.
  • Lin, T.-Y., Maire, M., Belongie, S., Bourdev, L., Girshick, R., Hays, J., Perona, P., Ramanan, D., Zitnick, C. L. and Dollár, P. (2014) Microsoft COCO: Common Objects in Context, arXiv:1405.0312. arxiv.org; dataset site cocodataset.org.
  • Li, Y., Yu, R., Shahabi, C. and Liu, Y. (2018) Diffusion Convolutional Recurrent Neural Network: Data-Driven Traffic Forecasting, ICLR 2018. Repository and results: github.com/liyaguang/DCRNN.
  • McMahan, H. B. et al. (2016) Communication-Efficient Learning of Deep Networks from Decentralized Data, arXiv:1602.05629. arxiv.org.
  • Ongaro, D. and Ousterhout, J. (2014) In Search of an Understandable Consensus Algorithm (Extended Version). raft.github.io.
  • MIT 6.5840 Distributed Systems (Spring 2026) Lab 3: Raft. pdos.csail.mit.edu.
  • Yjs, A CRDT framework with a powerful abstraction of shared data. github.com/yjs/yjs. MIT license.
  • W3C (2025) Verifiable Credentials Data Model v2.0, W3C Recommendation, 15 May 2025. w3.org/TR/vc-data-model-2.0.
  • De Vito, S., Massera, E., Piga, M., Martinotto, L. and Di Francia, G. (2008) On field calibration of an electronic nose for benzene estimation in an urban pollution monitoring scenario, Sensors and Actuators B: Chemical. Dataset: UCI Machine Learning Repository, archive.ics.uci.edu, DOI 10.24432/C59K5F.
  • OpenAQ, Open air quality data. openaq.org.
  • NIST, Post-Quantum Cryptography project. csrc.nist.gov. FIPS 203, 204 and 205, published August 2024, and the selection of HQC for further standardization.
  • European Commission, Regulatory framework for artificial intelligence. digital-strategy.ec.europa.eu, and the AI Act implementation timeline. Application dates of 2 August 2026 and 2 December 2026.
  • C2PA, Content Credentials specification, version 2.1. c2pa.org.
  • Model Context Protocol, Specification, revision 2025-06-18. modelcontextprotocol.io.
  • W3C (2023) Web Content Accessibility Guidelines (WCAG) 2.2. First published as a W3C Recommendation on 5 October 2023 and revised on 12 December 2024. Current version: w3.org/TR/WCAG22; dates from the W3C publication history.
  • Chen, D. (2015) Online Retail [Dataset]. UCI Machine Learning Repository. archive.ics.uci.edu, DOI 10.24432/C5BW33.

Frequently Asked Questions

Which final year project is best for computer science?

Choose the project you can finish, show working and explain in a viva. Scope beats ambition: a small system with a measured evaluation against a named baseline earns more credit than a large idea that runs out of weeks. Pick a problem with data you can obtain, a method your supervisor can assess, and a size that fits one or two terms.

How do you choose a final year project topic?

Start from three constraints: what data or hardware you can actually get, what your supervisor can supervise, and how many weeks you have. Then pick a topic that satisfies all three and still interests you. Write a one-page scope with a minimum deliverable and a stretch goal before you commit.

What are good final year project ideas for computer science in 2026?

Projects tied to a published standard or a regulation date age well. Examples with citable sources include migrating an application to NIST's post-quantum algorithms, building content-provenance labeling ahead of the EU AI Act's transparency dates, an accessibility audit tool for WCAG 2.2, and a Model Context Protocol server for a campus system.

How long should a final year project take?

Most undergraduate projects run across one or two terms alongside taught modules, which in practice means ten to twenty weeks of part-time work. Reserve the last quarter of that for writing and the demonstration. A project that is still being coded in the final fortnight rarely gets written up well.

What should a final year project report include?

An abstract, an introduction with aims and objectives, a literature or technology review, the design and implementation, testing and evaluation against the stated objectives, a discussion of limitations, and a conclusion. The evaluation chapter is where most marks are won and lost, so plan the measurements before you build.

WhatsApp